Applying Probability and Data with R to All of Us Datasets

 Applying Probability and Data with R to All of Us Datasets

Issued on 04 May 2026 by

University of Pittsburgh

This badge represents completion of an HSLS Self-Paced Learning module, published by the Health Sciences Library System at the University of Pittsburgh. Participants who earn this badge should be able to: - Recall simple functions in R for exploring any dataset.  - Define the five main data domains in the All of Us database. - Classify AoU variables as categorical or numerical. - Calculate summary statistics for both categorical and numerical variables.  - Apply data cleaning procedures to datasets in R, including filtering, removing missing values (NA), and managing duplicates.  - Derive new variables from existing datasets.  - Generate a new dataframe by merging various datasets.  - Use random sampling techniques to accurately represent target distributions.  - Describe the data preparation processes involved in the 'All of Us' code templates developed by HSLS. 

Issuer

University of Pittsburgh

teaching@pitt.edu

Criteria

Recipients of this badge:

  • Viewed all content pages in this module.