When Arjun Krishnan was nine years old, in a small town in South India, he watched Jurassic Park for the first time.
“And just like everyone else who watched it,” Krishnan recalled, “my mind was completely blown.”
“I thought: Oh my god! Wouldn't it be awesome to do something like this—I don't even know what—but something like this, you know, some time in my life?”
Although he has not (yet) realized his dreams of spawning dinosaurs, he surely never lost sight of his passion for biology. Today, Arjun Krishnan, PhD, is an associate professor of biomedical informatics at the University of Colorado Anschutz (CU Anschutz). Krishnan also works as the co-director of three graduate training programs: the Human Medical Genetics and Genomics PhD program, the Computational Biosciences PhD program, and the Colorado Biomedical Informatics NLM T15 Training Program. Across these roles, Krishnan works at the intersection of biology, engineering and technology.
Specifically, Krishnan develops machine learning and AI methods that leverage massive amounts of data to derive insights about the biology of complex diseases, focusing on the role of the social network of genes in various kinds of cells in the body.
Diseases are highly complex.
Although each disease is typically grouped under a single label, one disease may present differently in different groups—and even in different individuals.
Additionally, different individuals with the same disease may have different mutations in their genome that contribute to this disease.
Krishnan and his lab, the Krishnan Lab at CU Anschutz, explore these differences in their research, seeking to understand what Krishnan calls “the basis of heterogeneity of complex diseases.”
In other words, they investigate why the same disease can differ from person to person, including differences in symptoms, genetic mutations and biological processes that contribute to and relate to a specific disease.
Heterogeneity — The presence of biological differences among individuals, cells or conditions that share the same general classification, such as differences in traits, disease symptoms, genetic mutations or underlying biological mechanisms.
“Many diseases—the ones that we are super familiar with like type 1 diabetes, Alzheimer's, whatever it is—they are all just labels that have been assigned for convenience to large groups of people so that we can do something meaningful about that large group in some way, either in terms of diagnosis or treatment,” explained Krishnan.
“But there is nothing special about that label. If you take all the people with a specific disease label, they are all so different from each other. So, we really need to understand, what is the basis of all this heterogeneity?”
The Krishnan Lab approaches this heterogeneity at the functional, cellular and phenotypic levels, and they break their research into three main areas—or pillars—of research, which seek to explore broad, fundamental questions about disease biology. These pillars are:
They accomplish this by leveraging large datasets that are largely publicly available to look at diseases such as autism, gut dysfunction, and multiple sclerosis (to name a few). Instead of running clinical trials to collect data themselves, they find, collect, synthesize and repurpose data that already exists in the world. Then, they work with experimental collaborators who further test their computational predictions.
The datasets the Krishnan Lab uses can include data from genetic sequencing and Electronic Health Records (EHRs); they can also include repositories of data from papers that have been published over the course of decades. Examples of the datasets they use are large repositories like the UK Biobank, the Gene Expression Omnibus (GEO), or the Sequence Read Archive (SRA). In this way, the Krishnan Lab extends the use of that data and the money that was invested to collect it and uncovers connections between datasets that may not have been discovered otherwise.
Read on to learn more about how the Krishnan Lab approaches this research and how the three pillars of their research combine to help derive insights from large amounts of data.
The first pillar of research in the Krishnan Lab is understanding the intricacies of diseases which may be overshadowed by broad labels. Heterogeneity can be broken into multiple types:
“The first question,” Krishnan explained, “is based on underlying mechanisms: can we identify subgroups of individuals who have the same disease, but for completely different mechanistic reasons?”
Krishnan further illuminated this idea with an example: “Another way to think about heterogeneity is that if you think about people who have rheumatoid arthritis: some subsets of people who have this particular condition will respond to a particular drug in the market, and others won't. Why is that?”
“Maybe there is a specific pathway inside relevant cell types that are somehow perturbed in some patients, but not perturbed in other patients, and maybe the drug targets that pathway—and therefore it is effective in some patients and not others. So, this is the kind of analysis that we are thinking about, or the perspective that we are thinking about.”
In this way, Krishnan and his lab work to identify subsets of people who have the same disease but for different mechanistic reasons, which can have positive implications for future research, diagnosis and treatment.
The second pillar of research in the Krishnan Lab relates to exploring and understanding understudied diseases and disease contexts.
“So, what that means is if you think about a single disease, there are many fundamental factors that influence it—fundamental biological factors like sex, age, which tissue is involved, and how common is it in the population?” said Krishnan.
Although it might be assumed that we would have a strong understanding of fundamental biological factors like sex and age, these disease contexts are actually highly understudied.
“It turns out that we have extremely poor understanding of even basic things like why do men and women present a particular disease differently?
Why are some drugs more effective in men than women?
Or age specificity: what is a specific biology that is going on during adolescence that is not present in earlier development or later life?
What is specific about old age that has not happened yet throughout the rest of the lifespan and so on?
If we can use AI to understand what is biologically specific to each sex and each stage of life, we put ourselves in a position to ask bigger questions, like what lies behind exceptional longevity… a long-term goal that builds on the lab's work to better understand how sex and age shape biology.”
— Arjun Krishnan, PhD
Krishnan explained that most public samples lack sex and age labels. “Roughly 60–70% are missing,” said Krishnan. “So, we build machine learning models that predict sex and age from gene expression, which lets us measure how unevenly research covers each and learn patterns shared across related contexts.”
Although understudied contexts like sex and age may exist due to factors like research bias, understudied contexts can also arise for other reasons. For example, the tissues associated with a particular disease are understudied because they often cannot be safely or conveniently studied in living humans.
“Whenever we get data from a live person, it is always from clinically accessible tissues that you can easily access. You can't drill a hole and take a sample of the brain or get a sample of the kidney. You always get a sample of blood,” explained Krishnan. “What about the inaccessible tissues, which are actually relevant for the disease? How would you say anything about them?”
Clinically Accessible Tissues — A tissue is a group of cells that works together to accomplish a function in an organism. For example, blood, skin and muscle are all types of tissue. Clinically accessible tissues are tissues that can be collected from a patient relatively safely and easily, such as through a blood draw or a small biopsy. Studying these tissues can give researchers insight into biological processes occurring elsewhere in the body, including in organs that are much more difficult to sample directly.
To address this understudied context, the Krishnan Lab uses large amounts of data and machine learning/AI approaches to better understand the connections between clinically accessible tissues and tissues that are not as accessible.
Disease contexts may also be understudied simply due to the rarity of the disease. Some diseases are much rarer; consequently, these diseases have much less research surrounding them. The Krishnan Lab also explores these disease contexts.
The third pillar of the Krishnan Lab’s research looks at model organisms, which are organisms that researchers use as stand-ins, or ‘models,’ for humans. Model organisms help researchers develop a deeper understanding of biology and cultivate theories about diseases and the effects of different interventions like drugs or behavioral changes.
But modeling a complex human disease does not necessarily mean trying to recreate the entire disease in another organism. Instead, researchers can develop more nuanced models that mimic a specific aspect of a disease—such as a particular biological mechanism, symptom or genetic pathway—and study that piece in greater detail.
The goal is to close what Krishnan describes as the “human-model-human loop”: starting with data from people to identify the most relevant organism, traits and genes to study, conducting experiments in that model, and then determining which findings can meaningfully translate back to humans.
Yet, how does a researcher decide which model organism is the best fit for a particular context or disease?
How does a researcher know which genes to examine to better understand exactly what’s going on in the disease?
How does a researcher know how well the findings from a given model organism might apply to humans?
“Model organisms—like a mouse, fly, worm, or yeast—all those things are super useful for studying mechanisms and getting insights about human biology. But it is very hard to answer the following question: for a given human disease, how do you figure out the best organism, the best phenotype and the best genes to study so that you can recapitulate the human biology?”
“And vice versa. If somebody were to take a mouse model or a worm model... and then do a large genetic screen or a drug screen... once you've done the experiment, how do you know which parts of your results will translate back to a human somewhat?”
The Krishnan Lab employs powerful computational machine learning models to discover the connections between model organisms and human biology.
A key part of how the Krishnan Lab turns massive datasets into insight is the network. Drawing on decades of public data, the lab builds genome-wide maps of how genes connect to one another in humans and in model organisms. The lab's machine learning models then learn the patterns in these maps and use them to predict which other genes are likely to be involved in a particular biological process or disease.
One example is GenePlexus, a tool developed by the Krishnan Lab. Researchers enter a list of genes they already know are linked to a disease or process, and GenePlexus uses the network to suggest additional genes that may be involved. The tool is openly available to researchers around the world at GenePlexus.net.
Together, these connections are the "networks of understanding" at the heart of the lab's approach: each new piece of data makes the map more complete, and each prediction points scientists toward where to look next.
Molecular Interaction Network — A map of how genes and the proteins they produce relate to one another. Each gene is a point on the map, and lines connect genes known to work together, such as genes whose proteins physically interact or genes that are active in the same cells. Because genes that work together tend to be connected, a gene's neighbors on the map offer clues about what that gene does.
“Our lab's mission is to answer all of these questions and help other scientists answer questions like this using publicly available data,” said Krishnan.
This publicly available data helps the Krishnan Lab address the pillars of research listed above and make new biological discoveries.
One of the advantages of using this data is that it is largely publicly available. “This is not data that we get because we collaborate with a specific lab or we are part of some consortium where we get early access to newly-generated data,” said Krishnan. “It's nothing of that sort. We work with data that is available to all of us right now.”
Another advantage of using this data is that it ensures that valuable data that has been generated by labs across the world does not just sit on the shelf to collect dust. In this way, the use of public data supports the research community. In 2023, the lab received the NIH DataWorks! Significant Achievement Award for Data Reuse for methods and tools that help researchers make better use of public data.
Using these datasets does not come without challenges, though. The size of the data, as well as the heterogeneous nature of the data, makes it more challenging to work with.
“The scale of this data is massive because people have been generating high dimensional data and depositing them in public databases for about 30 to 50 years,” explained Krishnan.
“And these are individual datasets generated by individual labs spread across the world—and generating them using very different technologies as technologies and platforms evolve over time. So these datasets are massive. They are extremely noisy, they are heterogeneous, but they have very good coverage in terms of the entirety of all genomes, all genes, all proteins, all metabolites, and so on. They are almost like unbiased scans, or unbiased photographs of what's happening inside cells of the body at different scales of molecular biology. So those are the advantages and disadvantages.”
The lab’s commitment to openness extends beyond the public data it uses. Its publications are released with reproducible code, reusable software or interactive tools so other researchers can build on the work. One example is GenePlexus, an open web server that makes the lab’s network-based gene prediction methods available to the broader research community.
These datasets are massive, complex, noisy and heterogeneous, so the Krishnan Lab needs robust systems to work with them to explore the pillars of their research.
“We sit right in the middle—between these data and these questions—and say, ‘How do you design machine learning algorithms and AI algorithms to use and isolate signals from these massive datasets to answer those biological questions?’”
Working with these large datasets often requires many layers of fetching, cleaning, characterizing, categorizing and analyzing these data.
For example, many times the labeling of specific tissues or diseases is not consistent across datasets. To address this, the Krishnan Lab uses Natural Language Processing (NLP) tools to annotate the datasets in a standardized way. These annotations then allow researchers to find the data relevant to the question of interest. One example of this work is Txt2Onto, a tool that uses unstructured metadata to help annotate public datasets and studies.
“I think it's lovely to be able to understand these systems that are so complex,” reflected Krishnan. “That's why I think biology is so different from, you know, physics and chemistry where you can get to a strong set of underlying rules. Biology is like: Okay—random processes everywhere. Anything goes; everything has exceptions.”
“It's so complicated, you know? Everything has gone through a process of millions of years of evolution. There are many, many ways of getting to the same solution. I think it's wonderfully complex.”