MDI Biological Laboratory
Archive

Reproducible and FAIR Bioinformatics Analysis of Omics Data 2025

A training course for graduate students, post-doctoral trainees, and others who would like to incorporate bioinformatics into their biomedical research.

Apply Now
Application Deadline June 13

Overview

This course is an updated and extended introduction to our previous Applied Bioinformatics course. The renewed focus is on FAIR data – that is data that are Findable, Accessible, Interoperable and Reusable. This addresses a key initiative of the NIH and will prepare participants to benefit from the vast amount of publicly available biomedical data. We have maintained our emphasis on teaching students how to analyze gene expression data, because the skills required to analyze large transcriptomic data sets are rapidly transferable to proteomics and metabolomics.

The course begins with a complete introduction to the R statistical programming environment, and is designed throughout to be comfortable for participants who are new to R, bioinformatics and biostatistics. At the same time, the course is designed to be rewarding for participants with substantial experience in these areas, because each learning module includes exercises appropriate for beginner, intermediate and advanced students. A substantial amount of the course is dedicated to independent work on assigned problems. We have found that this approach leads to much higher levels of confidence and better retention of key concepts as long as challenges are appropriate to a specific student and students have plenty of access to knowledgeable teaching assistants. This class will have at least one teaching assistant for every six attendees.

The two week format of Reproducible and FAIR Bioinformatics Analysis of Omics Data enables students to build confidence in diverse areas including the following:

  • Planning Omics Experiments
  • Accessing the UNIX Environment
  • Identifying Differentially Expressed Genes
  • Pathway Analysis of Gene Expression Data
  • Applying Machine-Learning and Data-Driven Approaches to Gene Expression Data
  • Taking Advantage of Publicly Available Data
  • Ensuring Rigor and Reproducibility
  • Creating Publication Quality Visualizations of Complex Data
  • Sharing Code and Data
  • Analyzing Single-Cell RNA-seq Experiments
  • Analyzing Microbiome Data
  • Documenting Statistical Approach in a Publication
  • Developing a Data Management Plan

Course Directors

Course Faculty

Schedule

Reproducible and FAIR Bioinformatics
Analysis of Omics Data 2025


Day 1 – Thursday, June 26

 11:00
12:00            LUNCH
13:00            Course Introduction
14:00            Break
14:15             Defining your RNA-seq strategy
15:15             Break
15:30            Introduction to high-throughput data analysis
16:30            Break
16:45            Package Loading
18:00           DINNER

Day 2 – Friday, June 27

 11:00             RCR
12:00            LUNCH
13:00            Introduction to R Studio
14:00            Break
14:15             Introduction to R data types
15:15             Break
15:30            R logic loops and functions
16:30            Break
16:45            Promises and challenges of next-generation sequencing in contemporary biology
18:00           DINNER

Day 3 – Saturday, June 28

 11:00            UNIX Files and Directories
12:00            LUNCH
13:00            UNIX Jobs and Processes
14:00            Break
14:15             Pre-processing RNA-seq data with fastp
15:15             Break
15:30            Quantification with salmon
16:30            Break
16:45            Collaborative open research: lessons learned from working reproducibly with others
18:00           DINNER

Day 4 – Sunday, June 29

11:00            Gene ID Conversion in R
12:00            LUNCH
13:00            Exploratory data analysis and normalization of transcriptomic data
14:00            Break
14:15             edgeR and differential gene expression
15:15             Break
15:30            Over-representation analysis
16:30            Break
16:45            R Markdown and R Notebook
18:00           DINNER

Day 5 – Monday, June 30

 11:00            GSEA and pathway activation analysis
12:00            LUNCH
13:00            Online tools for gene set and pathway analysis
14:00            Break
14:15             Hands-on machine learning in R
15:15             Break
15:30            Choosing a machine learning approach to match your research question
16:30            Break
16:45            Rigor and reproducibility in machine learning
18:00           DINNER

Day 6 – Tuesday, July 1

 11:00            Group Project – RNA-seq data reanalysis
12:00            LUNCH
13:00            Group Project – RNA-seq data reanalysis
14:00            Break
14:15             Group Project – RNA-seq data reanalysis
15:15             Break
15:30            Group Project – RNA-seq data reanalysis
16:30            DINNER
16:45            Group Project – Presentations
18:00

Day 7 – Wednesday, July 2

11:00            Overview of rigor and reproducibility
12:00            LUNCH
13:00            Tools to access publicly available transcriptomic databases
14:00            Break
14:15             Data Cleaning
15:15             Break
15:30            Philosophy of Statistics
16:30            Break
16:45            Workflows for the integration of Human and Mouse Multiomics Data
18:00           DINNER

Day 8 – Thursday, July 3

 11:00            Using R for Basic Laboratory Statistics and Graphs
12:00            LUNCH
13:00            Correlation Analysis
14:00            Break
14:15             Models and Scientific Inquiry (lm)
15:15             Break
15:30            Analysis of Cytokine Data in R
16:30            Break
16:45            Testing for Associations in Count Data in R
18:00           DINNER

Day 9 – Friday, July 4 – HOLIDAY

Day 10 – Saturday, July 5

11:00            PCR Analysis in R
12:00            LUNCH
13:00            Model-Based Normalization in R
14:00            Break
14:15             Introduction to ggplot geometries and statistics
15:15             Break
15:30            Analysis of Cytokine Data in R
16:30            Break
16:45            FREE
18:00           DINNER

Day 11 – Sunday, July 6

11:00            Beyond the Bar – Dendextend and ComplexHeatmap
12:00            LUNCH
13:00            Beyond the Bar – FactoExtra
14:00            Break
14:15             Pirate plot and corrplots
15:15             Break
15:30            Chartering your way to better R Code – LLMs as Copilots
16:30            Break
16:45            Using LLM to Replicate Published Analyses
18:00           DINNER

Day 12 – Monday, July 7

 11:00            Writing a data management plan
12:00            LUNCH
13:00            Reanalysis of publicly available data on a shiny web server
14:00            Break
14:15             Public repositories for ‘omics data
15:15             Break
15:30            Sharing metadata – how to annotate your experience
16:30            Break
16:45            The human microbiome in health and disease
18:00           DINNER

Day 13A – Tuesday, July 8

 11:00            Introduction to single-cell RNA-seq
12:00            LUNCH
13:00            Cell QC and filtering
14:00            Break
14:15             Feature selection and clustering
15:15             Break
15:30            Normalization, dimension reduction, and visualization
16:30            Break
16:45            The human microbiome in health and disease
18:00           DINNER

Day 13B – Tuesday, July 8

 11:00            Introduction to microbiome analysis and experiment design
12:00            LUNCH
13:00            Denoising and QC Filtering with Cutadapt and DADA2
14:00            Break
14:15             Microbiome Data Prep: Trimming fastq files, merging paired fastq files with DADA2
15:15             Break
15:30            Microbiome Data Prep: Silva taxonomic reference and building ASV counts data frame with DADA2, looking at metadata/sample info data frame
16:30            Break
16:45            FREE
18:00           DINNER

Day 14A – Wednesday, July 9

 11:00            Clustering, cluster markers, and cell identity prediction
12:00            LUNCH
13:00            Exploring a pre-processed single-cell RNA-seq dataset from a publication
14:00            Break
14:15             TBD
15:15             Break
15:30            Sample integration and differential expression analysis
16:30            Break
16:45            FREE
18:00           LOBSTER BAKE

Day 14B – Wednesday, July 9

 11:00            Microbiome Data Analysis: Building physloeq object. Applying ggplot language/tools to look at alpha diversity and beta diversity.
12:00            LUNCH
13:00            Microbiome Data Analysis: Normalizing for Relative abundance, log2 fold change, making relative abundance plots
14:00            Break
14:15             Microbiome Data Analysis: Statistics for each type of plot we’ve made (alpha, beta, relative abundance)
15:15             Break
15:30            Microbiome Data Analysis: Training a random forest model from gut microbiome data
16:30            Break
16:45            FREE
18:00           LOBSTER BAKE

Day 15– Thursday, July 10
Graded Presentations

 

 

Tuition

$2,750 USD

The tuition includes all meals and housing.

 

CEUs

Students currently enrolled in the Molecular and Cellular Biology (MCB) Graduate Program at the Geisel School of Medicine at Dartmouth College may receive 1 full credit for completing this as an elective.

Dartmouth MCB

Funding

This research training opportunity is supported by a research education grant from the National Human Genome Research Institute of the National Institutes of Health under grant number R25 HG011447.

nhgri logo