MDI Biological Laboratory
Bioinformatics

A Q&A with Joel Graber, Ph.D., on the Comparative Genomics and Data Science Core

  • April 30, 2025

Basic scientific research is anything but—the modern science used to unravel the body’s biggest mysteries is increasingly complex. The brilliant minds working to answer the big questions of human aging need powerful tools, technology and expertise working alongside them. 

The Comparative Genomics and Data Science Core (CGDS) is one of MDI Bio Lab’s three Core Facilities. Experts in using powerful computational methods to process and analyze the massive data sets produced by modern experimentation, the CGDS collaborate with faculty to dig deeper into what the data holds.

We sat down with core Director and Senior Staff Scientist, Joel Graber, Ph.D., to find out about the core.

1. Can you explain what the Comparative Genomics and Data Science (CGDS) Core does?

Like most of our world, modern biology is very data rich. A single experiment can include tens or hundreds of millions of data points. The CGDS Core is made up of myself and three analysts who have the computational and statistical skills needed for this work. Our primary job is to manage, analyze, and visualize these data sets, helping our onsite research groups (as well as MDI Bio Lab’s INBRE partners) interpret their data to move their research forward.

It’s worth noting that maintaining an institutional core team provides two key benefits: first, it removes the necessity for each individual research group to hire (or train from within) their own data science specialist, and second it facilitates persistent skills and knowledge for working with genome-scale data sets. In my view, it’s really the best way to provide cost-effective expert knowledge institution-wide for MDI Bio Lab.

2. How does the CGDS Core support MDI Bio Lab faculty and help enhance their research? 

We work with MDI Bio Lab’s research groups to help get the most rigorous and complete answers out of the data sets that come from their experiments. Most molecular biologists and geneticists who work at the bench don’t generally receive in-depth training in the computational skills needed to work with these massive data sets, so we take on that role. Put bluntly, it’s our job to keep up with the state-of-the-art in bioinformatics and data science and apply that knowledge for the benefit of the whole Laboratory.

The CGDS team also undertakes a variety of training and educational activities, available to both our on-site researchers and also to the broader Maine INBRE community. We have hosted workshops, courses, and one-on-one consulting. These educational efforts serve several functions — instructing participants on high-level concepts such as cloud-computing and workflow systems, and also providing researchers with a hands-on experience and skill-upgrades to explore their data. While the core team will continue to carry out the large-scale computations, our goal is to get the results into the hands of the researchers, while also giving them tools and skills to better understand and investigate their data on their own.

3. What does the collaboration process look like when you work with a faculty member?

As you might guess, communication is probably the most important part of our job, since our research collaborators know the questions they need answered, but we have the skills and tools to get those answers from the data. There are many computational and statistical approaches available for analyzing “big” biological and biomedical data. To select the most appropriate of these approaches, it is critical that my team and I understand the research questions. I meet with all of MDI Bio Lab’s faculty on a regular basis to discuss projects, but, more importantly, to maintain a conversation about the overall direction and goals of their research. My team also spends a lot of time working one-on-one with research assistants, students, and post-docs, sitting at the computer working together to explore the data. Refining analyses is an iterative process but it is the best and most efficient way to answer their questions and help support other activities, such as writing publications and grant proposals. (See below for an example of this kind of collaboration.)

4. What do you think might come next for this field?

I think the biggest situation looming in our field is the combining of artificial intelligence and machine learning (AI/ML) technologies. There are already many initiatives pushing the forefront in this area, both within scientific research and the wider world. Siri, Alexa and the “for you” recommendations on your subscription TV service are all examples of AI/ML that you encounter daily.

Many biological and biotechnical projects can also be framed as problems of pattern recognition and classification, which puts them squarely in the realm of AI/ML approaches. The huge challenge for us is to understand both the strengths and weaknesses of these approaches, so that we can apply them effectively, while also ensuring that the results are meaningful. The CGDS Core’s role in this process is two-fold: First is the need to acquire and implement the algorithms associated with this AI/ML approach. Secondly, but I’d argue much, much more importantly, is building a deep enough understanding of AI/ML in the scientific research environment to assist our lab groups to use them correctly and safely. By that, I mean avoiding over- or inaccurate predictions or even hallucinations (in AI speak a hallucination is an instance where a system generates false information due to an error in its “learning,” often appearing as a plausible but incorrect output). The critical goal will be to pair our AI/ML analysis with testable hypotheses that can be verified by bench experiment.

Reproducibility, rigor and truth are the basis of all scientific endeavor, so although the field is moving in the direction of utilizing AI/ML, we are doing it with the utmost caution.


CGDS/Faculty Collaboration

The two figures below come from collaboration with the Coffman Lab, as published in Drepanos et al Scientific Reports volume 13, Article number: 12239 (2023).

The Coffman Lab studies how chronic stress, traumatic experiences and other toxic environmental exposures in early life can have persistent developmental effects that impact how the body responds to stressors later in life.

This first figure shows the mapping of the Coffman group’s data from bulk RNA-seq* of zebrafish embryos onto a previously published single-cell atlas**. The 2nd and 3rd red plots from the top suggested increased expression of liver cells. One way in which that could have happened is with larger livers.  The second figure shows the confirmation of this information, experimentally.

Figure 1

Mapping of the Coffman Lab data from bulk RNA-seq of zebrafish embryos onto a previously published single-cell atlas

Figure 2

Graph shows the confirmation of increased expression of liver cells experimentally

 

*  RNA-seq is a technique used to identify and measure RNA molecules, helping researchers understand gene activity.

**  A single-cell atlas is a comprehensive collection of gene expression patterns in individual cells, that allows for deeper understanding of cell diversity, organization and function.