Skip to content
Brant Peterson — Brant Peterson on Valo Health's Patient-First Drug Discovery
Data in BiotechEpisode 64

Brant Peterson on Valo Health's Patient-First Drug Discovery

Brant Peterson of Valo Health discusses how patient data, genetics, and causal reasoning shape drug discovery decisions before a molecule reaches the clinic.

52:33Full transcript below
BP

Brant Peterson

Vice President & Fellow, Data Science at Valo Health

Overview

The conventional drug discovery pipeline struggles with high failure rates and immense costs, often by focusing on molecules or models before deeply understanding patient outcomes. Valo Health confronts this challenge with a “start with the end” philosophy, prioritizing the patient experience and clinical readouts from the outset. In this episode, host Ross Katz talks with Brant Peterson, VP & Fellow, Data Science at Valo Health, about how his team integrates diverse real-world data and advanced causal reasoning to de-risk drug development significantly earlier.

Peterson, with a background in genetics and molecular biology from Novartis, highlights that effective drug discovery isn’t about applying machine learning to massive datasets blindly. Instead, it demands a profound, upfront understanding of data generation processes, inherent biases, and causal structures before a single model is built. This approach allows Valo to pinpoint critical patient subgroups, understand disease progression, and validate biological hypotheses with unprecedented clarity.

The conversation delves into Valo’s unique data landscape, from drawing on extensive electronic health records (EHRs) and genomic insights through Mendelian randomization to integrating wet lab experiments. Peterson illustrates how this multi-modal strategy, applied to complex neurodegenerative diseases like Parkinson’s and Alzheimer’s, identifies druggable mechanisms and defines the patient populations most likely to benefit, ultimately accelerating the delivery of targeted therapies.

Key Takeaways

Causal Understanding Must Precede Machine Learning for Meaningful Insights

Blindly applying sophisticated machine learning models to vast electronic health record (EHR) data without understanding its generative process and biases leads to trivial or misleading conclusions. Valo Health discovered that deep, upfront causal reasoning – identifying what the data truly represents and its inherent confounders – is essential before any model development. This diagnostic step ensures that complex analytics yield actionable biological and clinical insights, not just superficial patterns.

Integrate Diverse Data Streams to Build Strong Causal Graphs

Effective drug discovery relies on triangulating evidence from multiple sources. Valo Health bridges real-world patient data from EHRs with wet lab experiments and genomic causal models, such as Mendelian randomization. This multi-modal integration allows them to construct detailed directed acyclic graphs (DAGs) that separate elements of a hypothesis, calibrate quantitative values at different points, and validate mechanisms across various scales, from cell to patient.

Patient Subgroup Identification Transforms Disease Understanding

Complex diseases like Parkinson’s and Alzheimer’s are not monolithic; patients exhibit highly variable progression and symptom constellations. By segmenting patient populations based on their distinct real-world experiences, Valo Health uncovers specific unmet needs and unique underlying mechanisms. This granular understanding allows for the identification of targeted patient subgroups where therapeutic interventions are most likely to have a meaningful impact, streamlining development efforts.

Target Prioritization Requires Advanceability Assessment Early On

Identifying a strong biological mechanism is only the first step in drug discovery. Valo Health critically evaluates the ‘advanceability’ and ‘druggability’ of potential targets in parallel with biological validation. This involves considering practical factors like tissue specificity, intervention type, and the complexity of therapeutic modulation early in the process, ensuring resources are focused on discoveries with a clear path to becoming medicines.

Related: CorrDyn provides data engineering expertise to structure complex datasets and performs thorough data assessments to ensure data quality and utility. Our technology strategy services help organizations build strong data foundations within the biotech and life sciences sector.

Full Transcript

Brant Peterson: I think there’s just a lot to be said for being able to integrate specific data sets around specific patient populations. The first question isn’t what’s the machine learning model? The first question is what are the data sets in which the most informative kinds of information are generated in the least dangerously biased way?

Jason: Welcome to Data in Biotech, a podcast from CorrDyn where we explore how companies leverage data to drive innovation in life sciences. Every two weeks we sit down with an expert from the world of biotechnology to understand how they’re using data science to solve technical challenges, streamline operations, and further innovation in their business. Here we go.

Ross Katz: Brant Peterson, welcome to the Data in Biotech podcast.

Brant Peterson: Thank you.

Ross Katz: Awesome. Well, just to kick us off, would you mind giving us an introduction to your background and what brought you to Valo Health?

Brant Peterson: Sure. I went through the list of all the things you can do basic science on. I worked on worms and flies and mice, Ph.D. and postdoc mostly, somewhere around information encoding in biological systems. So my training is basically in genetics and molecular biology.

Ross Katz: Yeah, so you’ve got this genetic background. Can you talk a little bit about how your background in genetics and your experience at Novartis framed your approach to drug discovery?

Brant Peterson: Yeah, absolutely. When you’re trained in genetics, even from the very beginning — I think I did a fruit fly lab in high school — it’s distinctive from the rest of the bio-education in that you start at the end. You’re thinking about how this complex causal chain that you have no hope of getting your arms around ends up in an organismal phenotype. And so then the other most important thing is the phenotype. And you pretty much immediately learn that’s the hard part. You’re sorting flies, you’re not sorting them on genomes. You’re sitting there looking at how many bristles there are on their shoulders. And I think that painfully directly translates into the really hard parts about biomedical research. And rightfully so. The things we care about are what a patient experiences in the course of health and disease. We run our clinical trials on how a patient feels, functions, or survives. We have to start at the end. And I think that’s actually really helpful when you’re thinking through some of the really hard steps in drug discovery, like do we go ahead here? Do we need more data here? Are we ready with this molecule? These are all questions that are fundamentally informed by whether or not we’re ready for where this is going to end. And then I think the other piece that ended up being maybe more relevant than I thought it would be at the outset is I worked a lot in grad school especially on the mechanisms that underlie molecular evolution in traits between closely related pairs of species. And that evolution is primarily non-coding. So most of what that mechanistically is based on is variation in the non-coding genome, in the regulatory genome. It just turns out, didn’t have to be this way, although I think population geneticists would argue that in fact selection means it has to be this way, that those are a lot of the same mechanistic processes that underlie variation in health and disease in people that we see when we do these very large-scale genetics studies. So the mindset, yes, but also maybe some of the praxis turns out to be really directly translatable.

Ross Katz: Awesome. Can you talk a little bit about your journey with Valo Health and how what you just said is embodied in the way that Valo approaches drug discovery?

Brant Peterson: Yeah, absolutely. The pitch was more or less start with the end. It was almost as simple as that. I started here just over six years ago, and when I interviewed it was a fairly small place. I think fewer than 40 employees. And very much still figuring out exactly how the business would work. But the notion at the core of it was that there was something about the combination of the advanced computing that we can do and the data that we had access to that would let us derisk much earlier the kinds of things that kill projects in drug discovery. So I had had this experience at Novartis, really rewarding experience actually, building machine learning models that predicted safety outcomes. And that was super exciting. And we learned it was going to be complicated to figure out how to get that into production, in a sense. For good reasons, for all the right reasons, smart people doing their jobs well in a structure that didn’t give us a path. And Valo immediately, just at the outset, was built to make sure that we were pulling that information as early as we could, that the integration of kinds of evidence with the points at which that evidence could have the greatest impact would be globally optimized rather than locally optimized, so that we would always be thinking about where we’re heading.

Ross Katz: I would love to hear from you, from the highest level view, what does the data ecosystem look like that feeds that on a tactical level?

Brant Peterson: I love this phrase data ecosystem. I think this is actually one of the really early things we figured out. But just to rewind, I think your question is the answer to the question. We are looking for the data that will at the earliest point allow us to derisk the core hypotheses that are going to impact the readout in the clinic. We’re thinking about what’s a Phase 1B, what’s a Phase 2 readout, who are these patients, and how do we find them before we ever make a molecule, so that we can understand what’s going to drive the disease state, what’s going to restore health in that patient population in ways that matter to patients in clinical development. So that’s the problem statement. The data flow directly from that. If we want to understand causality in patients, we need data from patients. So where is that? We can talk a lot more about the details, but fundamentally we’re looking for the kinds of data that give us points of leverage over understanding causality in the patient populations that we intend to treat, and understanding how to bring causal perturbations we may be excited about to the right patient population. I think there’s a little bit of a joint optimization there. Sometimes you don’t start with the patient, have zero ideas in the universe before you go look at the patient data. It is often the case that you’re working with smart, motivated biologists with deep knowledge of a particular biological process and lots of really solid ideas for how we could modulate that process. And then the task becomes how do we find where exactly in health and disease the ability to perturb that process is going to matter the most? How do we prove it? How do we get it to patients? So that’s this concept of we can notice patterns in patients, we can resolve causal structures in those patterns, and then we can use it a few different ways, different but similar ways. And then the other major fork of the work is a bit more closely related to that side project I described earlier. There’s a whole ecosystem here we’d call it closed-loop discovery or closed-loop chemistry, but doing essentially that same thing with the molecule. Starting where we mean to end, modeling the properties that are going to make a difference to our ability to get to the clinic — not just does it bind, but can we develop a particular hit into something that we can synthesize at scale, that is going to be soluble, that is going to be safe, that’s going to have good ADME properties? So these are the kinds of things that we can model on data before we have a hit, and then continue to refine those models as we move through the drug discovery process. Analogous to having a starting point before we get started and continuing to refine those models in the patient world, we have the same mindset and process in the compound world. And then you ask what’s the infrastructure, what’s the data architecture? I think this is again one of those things we came in the door with — we’re going to use the same machine learning models that allow us to predict whether or not a small molecule is going to bind a particular protein. We’re going to use those models for drug discovery, for asking does this series look like it’s got a good future for modulating this target? We’re also going to take those exact same machine learning models and we’re going to apply them not even in the clinic, in practice, in the real world to the likely off-target, undocumented off-target properties of already approved medicines that patients are taking out there in the world.

Ross Katz: So as I understand it, the clinical data that we’re talking about is coming from electronic health records or EHRs, is that correct? Or is there other clinical data going into it? And I’m interested in how that feeds the discovery process for you.

Brant Peterson: Yeah, I think the vast — 99 plus percent — portion of the data that we’re going to talk about by pure volume in any given context will be real world, will be EHR. I think there’s just a lot to be said for being able to integrate specific data sets around specific patient populations. A lot has been derived from research in UK Biobank. We obviously want to have that as part of the story. But in specific cases, I think about ALS. This is a relatively rare disease. Participation of patients and their families in these data collection efforts specifically targeted around what matters in this disease is huge. So even if it’s only a few thousand people compared to millions and millions of people and tens of millions of patient years — these numbers look big — but it’s really important to plan for where these relatively smaller data sets can have outsized impact on specific questions. That aside, yes, the primary data source we work with as we’re setting up our core research, as we’re building machine learning models, these are electronic health records, real world data drawn from larger healthcare systems for the most part.

Ross Katz: And my understanding is that the EHR data that you’re working with, one of the things that is unique about how you’re using it is that you’re not aggregating the data and attempting to make predictions in the aggregate, but that you’re developing models that are making predictions on an individual level based on features of that individual from the longitudinal course of their EHR records. Am I understanding that correctly? Or how does that work on your end?

Brant Peterson: Yeah, that’s right. And I think this is in many ways driven by the specific questions we’re asking. Sometimes we’re asking these very wide aperture questions. We’re trying to understand the disease course in very large patient populations or sometimes even the course of many — call separate diseases, they’re not really — multi-morbidity and the progression through different symptomatic presentation or chains of drivers. We think a lot about metabolic disease and the consequences of metabolic disease. If we think about heart failure and kidney failure and which causes which, you have to take a broad aperture on these questions. In these cases, yeah, we’re looking over decades. We’re looking over hundreds of thousands of patients. We’re looking over as much of the health information as we can pull in frankly. We’re looking over tens of thousands of features. In other cases we might have a pretty specific question. We might have a known off-target of a drug that is of interest and we’ll be looking at a target trial emulation approach, where we know exactly what we need to know. We know what the clinical study would look like and we’re really laser-focused on getting those data as good as we can.

Ross Katz: Right. So there’s this element of yes, you’re creating the general infrastructure to process large quantities of these longitudinal EHR records at the patient level for the purpose of training larger models, but also you’re creating the infrastructure that you need to ask those more targeted questions when you have a hypothesis that you want to test and understand based on the data you have available. Am I thinking about that right?

Brant Peterson: Yeah, that’s absolutely right. And that’s as much a technology problem as it is a culture and process problem. A lot of what we’ve built up is really an institutional knowledge base of how to start asking one of these questions and then how do we know whether the question will benefit from a decade of information prior to a diagnosis or a follow-up period longer than the baseline period? These are the sorts of things where knowing how to ask these questions and extract early answers quickly is something that we’ve built a good amount of technology to support, but the use of the technology requires that understanding of the question in order to translate. So a lot of it’s been building that hybrid engineer scientist. 100%. The thing we inherit from epidemiology most strongly is the drive to understand the generative process for the data before we start asking questions. One of the pieces of culture or institutional knowledge is understanding who do we need to talk to in order to take that step from the data as we’ve structured and curated it in a general way to the data set that’s going to best inform the next project we’re going to do at whatever scale. Is it going to be a language model trained over tens of millions of visits or is it going to be some very specific analysis in 5,000 very carefully selected people? The first question isn’t what’s the machine learning model? The first question is what are the data sets in which the most informative kinds of information are generated in the least dangerously biased way? And this means we need to understand bias inherent in the generation of the data from the healthcare process, from the coding process, from how we see it. That upfront, almost “don’t touch the keyboard yet” activity, we’ve learned a lot about how to do well and we’ve learned some things that don’t work as well. One of the things that I’ll just say, you have to talk to experts. If you’re talking about data generation, you have to talk to clinicians. If you’re talking about the mechanisms you want to extract from the data, you have to talk to your biologists. It’s fundamentally an extraction and elicitation of that knowledge followed by encoding that knowledge into our analyses that keeps us from riding off the rails immediately.

Ross Katz: Yeah, that’s really interesting and also flies in the face of a lot of what people want to say about — can you just give an idea of the scale of the patient population data that you’re working with and then maybe your thoughts on how you think about the process you just described relative to the “throw a transformer at it, let it extract all the features” approach?

Brant Peterson: Yeah, the first thing I did when I walked through the door here was got into one of these data environments and set up the biggest GCN model I could, and it ran for weeks, and it sure did make clusters. We got to the end and I started showing them to people and there were two really unfortunate things about that experiment. One was that the UMAP projection looked really unfortunately like the chocolate ice cream emoji that we use for other things, which was just the chance happenstance of UMAP. But the other was that it was shockingly uninformative. It told us that old people see the doctor more often than young people. And it taught us that there are a whole category of diagnostic codes and medications and even procedures that only occur in women between an average of 20 and 35 years of age. And then you go unpack that and you’re like oh they had babies. So this was a really fantastic early wake-up call, and I’m glad it was early. We went back actually with that exact same architecture in the context of this big broad-sweeping question about what is true across cardiovascular and metabolic and renal disease, what’s shared, what’s not shared, what defines trajectories? And it took us several months prior to even starting to structure the input data for PyTorch just to figure out which people not to use, what stratifications to impose initially, which features were and weren’t informative, even in how we included them in the model. What is a patient-level feature versus what is a visit-level feature required really sitting down and thinking about what is actually time-varying, what matters, what do we take out, what’s confounding? And we learned that set of lessons and now it’s almost baked into how you do this stuff. We know we’re going to see correlates of socioeconomic deprivation. We know they’re going to show up in different ways. We have an idea for what the hallmarks are and we have a toolkit for removing those axes of variance from the procedure, informed both by what we see in the data and also by what we know will be there from talking to folks who see a given patient population.

Ross Katz: That’s really interesting and just to clarify, you all are of a relatively large size and you’re talking about millions of patients over decades of —

Brant Peterson: Yeah, depending on the data set. In some cases we’re looking at a couple of million people over an average of 20 years, in other cases we might be looking at more like eight to 12 million people but where we have a shorter period of good solid data from those populations. On those scales.

Ross Katz: But even at that scale what I’m hearing is that embedding your prior understanding of what are the confounders and what are the mechanisms into not just the model you’re using but the data set that you’re presenting to that model, and setting up the model for success is a key learning from early on in this phase. Is that right?

Brant Peterson: Yeah, I think the other thing is really thinking hard about what we’re after here. Again, start with the end. It is easy to learn about patterns of care from data collected largely to track patterns of care. It’s comparatively more difficult to learn about fundamental biology from these data. It’s not something that most folks that walk in the door here have done before, it’s not the most common thing that a healthcare outcomes person would have worked on. If you worked on mechanisms of disease as a biologist you probably didn’t spend a lot of time with healthcare data. This activity does require us to tune in some specific ways how we ask the questions. How do we remove from the data, like we talked about socioeconomic deprivation, how do we remove from the data patterns that are reflective of care when we’re not looking for care? Or how do we make sure that when we’re asking about specific patterns of care, we’re asking in ways that aren’t trivial antecedents of the disease we’re asking about in the first place? These kinds of threats to inference around triviality matter a lot when what we’re after is mechanisms.

Ross Katz: Yeah, that makes a lot of sense. I want to talk about two aspects of your data ecosystem that we haven’t spent as much time on yet. There’s the wet lab component that you mentioned a little bit earlier, and then Valo also has really deep roots and you have deep roots in genetics. I’m interested in how the wet lab feedback loop works and how the genetics component also enters the picture.

Brant Peterson: Yeah, the answer is in different places. I’ll start with how we think about the integration with the wet lab. I think this was again extremely straightforward when the basis for discovery was the models that we had in the lab. When the basis for discovery is patient data, the activity of bridging becomes meaningfully more difficult, less obvious I’ll say. And we found a couple of places, some of them I think anticipated, some of them a bit surprising, where the ability to weave these two together at specific points in drug discovery has really made a difference. When we’re thinking about out of all the genes in the genome or out of hundreds of potential genes in a pathway, how do we prioritize biological hypotheses? It’s relatively hard actually to integrate the wet lab at that scale. In these differentiated human-centric organoid models the match isn’t ideal. But as soon as we step into a specific hypothesis, as soon as we have an idea about a pathway, as soon as we have an idea about a target or a process, we’ve found surprisingly often somewhere in the data in the real world there are proxies for that process. And often those proxies are either drugs specifically modulating that process or known off-targets. Well, when I talk about genetics I’ll talk about pleiotropy, but I’ll surface it here. Drugs are pleiotropic as well. Drugs impact multiple mechanisms through known and unknown off-targets. We have the ability to discover some of those through that chemistry platform I mentioned, but a lot of it’s just in the literature. And we’ve been able to use that as a pivot point for specific molecular investigation into non-obvious pathways. As soon as we’re there, the ability to calibrate what those drugs are doing in specific cells, in specific disease-relevant circumstances becomes massively enabling for those real world evidence studies. As we’re building up the DAG of the hypothesis – the graph that says if you twiddle this gene in this way, if it goes up, if it goes down, we expect these markers to change in a cell, we expect these pathophysiological processes to change in a patient, we expect these outcomes to be impacted on an 18-month timeframe. That’s a hypothesis. Some of those arrows early on we’re not going to get in a person, maybe ever, certainly not until we go to the clinic and collect specific data. But they’re cheap in a cell. That ability to formulate the hypothesis and then inform different elements of the hypothesis with different pieces of evidence has been really valuable in thinking about which of these potential real world evidence studies are going to get us closest to the mechanism we’re after. That may not be obvious in what is known in the world, that may not be obvious in the patient data, but we may be able to really incisively separate those when we go to look in the right cell type in vitro.

Ross Katz: Yeah, that’s interesting and it also teases us up for the discussion of genetic causal models and pleiotropy and I would love to open that up. But just first what I’m understanding is that the causal DAG that you’re creating, the DAG for directed acyclic graph that says these genes impact these mechanisms, which impact these mechanisms, which impact these mechanisms, which ultimately impact the therapeutic outcome that we do or do not want to create, having that written down is not just useful for testing the entire DAG itself, but it’s also useful for identifying opportunities to gather wet lab data that validate or invalidate any leg of that DAG, like the relationship between one point and another. Am I understanding that correctly?

Brant Peterson: Yeah, that’s right. And it’s just often the case that these endophenotypic measures – what’s going on below the surface, in the tissue, in the cell – are really hard to get in a patient and really easy to get in the lab. And then vice versa. We cannot ask what will the patient outcome be in cell culture? That just is translation. It’s almost trivial to say. But the structure lets us separate out elements of the hypothesis that we can test and more than test that we can calibrate, that we can benchmark and put quantitative values to that allow us to estimate. If we can get 50% inhibition we see a twofold change in this or a fourfold change in this biomarker, we can monitor that biomarker in plasma, say. And we can say this is how much of that change we need to see in order to see a 13% reduction in risk of this event on a two-year timeframe, and that’s our registration endpoint for the CVOT.

Ross Katz: Yeah, that’s interesting and it strikes me that there’s a lot of “what would we have to believe is true in order to believe this entire mechanism” and then targeted experiments that can help along the way. What I’m hearing is that some of those experiments are actually wet lab experiments, but sometimes you determine you can ask that question of the EHR data with previous drugs that have been brought to market that may modulate a similar mechanism, and so there’s ways to collate this information in order to weigh the evidence across all of the methods you have of getting at what the truth is.

Brant Peterson: That’s right. And we’ve alluded to it a couple times – genetics gives us another tool in this toolkit. If we think about the canonical intersection between genetics and drug discovery, it’s this idea that we can track variation in populations in the genome to associate that variation with variation in clinical outcomes. We can say get all the people that have ever had a heart attack and get all the people that have never had a heart attack despite living long enough to be about the same age as the people that have had a heart attack, and sequence them at 10 million odd variants in the genome that we know how to measure, and then just ask one by one hey, are you associated with heart attack? How about you? How about you? Do it 10 million times and we’ll come back with a list of votes. And then we have an infrastructure built over decades and many hundreds of people’s work for essentially attributing the variation in the genome that is observed to be associated with variation in health outcomes to probable mechanisms that we could imagine intervening on for drug discovery. There are lots of ways to do this. I think a reasonable approach that is fairly systematically applicable is to use the genetics of variation in those perturbable mechanisms and match it to the genetics of variation in outcomes. And this turns into the instrumental variables, Mendelian randomization story that I think a lot of folks are familiar with where we say we have the genetics of outcomes, health and disease outcomes, we have the genetics of something intermediate – the level of a protein, the level of a transcript, the level of some biomarker like CRP or something like that, something that gives us a hint about what’s going on inside the box – and we can combine those two data sets to say if a particular variant in the genome impacts that endophenotype, gene expression or the level of a biomarker or something, and also disease outcomes, well we’re reasonably certain CRP doesn’t change the genome, we’re reasonably certain heart failure doesn’t change the genome, and we’re reasonably certain that we can at least direct the arrow between the two, between the gene and the disease say. Then we can start to ask this causal question, canonically backdoor style, Judea Pearl causal DAG type question. And when you start to do this, you’re already back to building DAGs. This has a really natural extension once you move beyond the simple pairwise case to similarly populate our confidence and help us parameterize the magnitude of effect on those edges on that hypothesis that we’re expressing.

Ross Katz: So it makes sense that you’re constructing this DAG, that you’re using the correlations between the nodes in the DAG with each other and with the outcomes to construct this causal story, and then you’re bringing information to bear to either support or refute that causal story. But there’s assumptions that are built into that kind of model. It would just be interesting in understanding for you what are some of the key assumptions that you’re testing along the way.

Brant Peterson: Yeah, absolutely. We were just talking about genetics. And we’re talking about an approach that has seen enough popularity that I no longer feel like I need to try to say the word Mendelian randomization or even spell it. I just say MR in meetings and expect people to know what that means. I think when something reaches that kind of penetration in the field, it is because we know how to use it safely, and this is one where the assumptions are very clearly stated. The approach itself, instrumental variables analysis, is almost 100 years old. I think the first person that used the term in print is Philip Wright, who is Sewall Wright’s dad. Sewall Wright, founder of population genetics. This is OG math. And it does a nice job of stating that one of the things that can’t happen is there can’t be a way to get from the instrument, the genetic variant, to the outcome any other way than through the cause. And that’s not something the method proves, that’s something the method assumes. And when you’re dealing with genomes and biology in general, the idea that there’s only one path to anything is fundamentally bankrupt. As soon as you’re looking at what does this variant do, the answer is likely a lot. And that a lot is this concept of pleiotropy we mentioned. The idea from an organismal biology evolution perspective is that a single gene can influence many outcomes. Canonical case: alleles predisposing to sickle cell anemia are also protective against malaria. This produces a canonical balancing selection situation because those two things are inseparable – evolution cannot help us out with sickle cell without making us at higher risk for malaria. In the case of this causal reasoning, the idea that a single gene could act on multiple traits gets translated down to a single variant could act on multiple genes. You could have an apparent action on multiple genes through the invention of linkage disequilibrium – and again I think we’re going to end up with a pretty big term sheet on this – but linkage disequilibrium being the idea that variants that are near each other in the genome cosegregate across populations. It might not be actually that the causal variant for this gene is also the causal variant for this gene, they might just be close to each other. They might be the same. There are some really fantastic, delightful examples in the literature. I think maybe my favorite two are the endothelin-1 factor-1 locus in blood pressure, which is this case of a single region in a bit of chromatin that is accessible in endothelial cells that clearly has an association with blood pressure, and you can read a lot of papers, I think very probably has that association through both endothelin-1 and factor-1 action, possibly in different tissues. Factor-1 maybe is a muscle effect. There’s this really nice story that absolutely flies in the face of all of these core assumptions of these methods. And it’s not that you don’t see a signal when these things happen, it’s just that you cannot interpret the signal on its own. And this gets back to the word picture you’re painting of you have to enumerate the other paths in the graph. Even if we have tricks – we like the instrumental variables approach especially when there are multiple instruments because it lets us do a trick in the regression that essentially helps us feel better about the risk of pleiotropy. We do this Egger test and it’s real, it’s telling you something, but it’s not a substitute for understanding the causal structure you’re imposing. The other example I really like is this obesity association with the gene FTO. FTO is in the brain, it’s probably doing something – there’s a whole literature on how that could work – but there’s also this really beautiful paper from 2015 showing that that enhancer is in physical contact with a locus containing two transcription factors that are responsible for browning in adipose tissue that’s directly relevant for energy expenditure and metabolism. You can debate but you have to understand that the mechanism could easily be yes. We may not be looking for the gene, we’re looking for how this locus is assembled to causally impact by multiple paths, pleiotropic paths, this one outcome.

Ross Katz: From a drug discovery perspective, what does that imply about what you’re looking for in the data, in the DAG, in order to know that you’re onto something?

Brant Peterson: I think again this is that “a gene” versus “the gene.” Maybe one of the genes in the locus that could be associated is a structural protein, is a transcription factor expressed in embryonic development – all these things that aren’t drug targets. We haven’t talked about this so much yet, but I’ve sketched around a few ways of generating causal evidence in humans. I’ve said I like genetics, I’ve admitted to being a geneticist whether that’s a good thing or a bad thing. I’ve talked a lot about start at the end and the phenotype and all that. I even said a good thing about a genome for the purposes of causal inference is that it doesn’t change. That actually makes it a terrible proxy for drug discovery. In drug discovery what we’re going to do is make an acute intervention. In the Judea Pearl causal reasoning sense, we’re going to have a do intervention that wasn’t present at some point and then a patient’s going to start taking a drug and that will induce some change. The genome cannot model this. It’s a lifelong exposure. When we start to talk about this, we need to be thinking that there are obviously true causal signals that are not vulnerable to therapeutic modulation. And maybe it’s not that they’re not theoretically modifiable, but we also have to think practically about what drug discovery’s going to look like. We could have an example of something where we see that when we look at the regulation of a particular gene in a particular tissue, the direction of effect with respect to disease through these causal models goes in one way. We look in another tissue, it goes the opposite way. We could envision a therapeutic that would be tissue-specific, maybe there’s some antibody-drug conjugate approach, but these are hard. They’re just hard. When we’re thinking about prioritization of therapeutic hypotheses, I think we have to be reasoning about the advance-ability of the program that we would envision launching based on this resulting DAG. And it may be that when we look at the evidence, we see a really robust association with something that happens in embryonic development or with a structural protein, and a quantitatively weaker observed variation association with outcome in something that we think we could drug. And then for instance, we might look at the data and see that we observe more constraint in one or the other of these and that may tell us something – we see more tissue specificity versus broad tissue expression. A lot of other factors go into thinking about what’s going to make medicine. And ultimately I think the decision to prioritize targets is something we spend a lot of time on frankly – building systems for extracting this integration of evidence downstream of all of this causal reasoning and causal DAG. Building the biology gets us to the first checkbox on drug discovery.

Ross Katz: Right. You build the biology, you go through all of this work, you create the mechanism, you validate the mechanism, but you’re also – you have limited resources and time and so you’re trying to devote your time and attention to the places where you actually think the mechanism, if valid, can be drugged, has a path forward. There might be many biological discoveries that you could make that would be interesting in and of themselves but could not lead to the end goal of what you’re investing to do. That’s interesting in and of itself, it makes a lot of sense.

Brant Peterson: A lot of the architecture of how we implement these things in projects has to do with what we can do systematically versus what has a higher individual unit cost around a target. There are in some cases trivially-ish scalable activities. When we talk about running Mendelian randomization we know how to do that. When we talk about differential expression in tissues or even sequencing patient genomes – we haven’t talked too much about this, but one of the really exciting things that we’re up to most of the time is identifying the most interesting patients in the real world cohorts that we have access to and pairing those with biobank sample collection to get genome sequences, to get multiomics or at least proteomics from those patients usually from peripheral tissues or from peripheral fluids. Those all paint fairly straightforward pictures for how to do the analysis and how to interpret any given result. Some of the other stuff we’ve been talking about, a lot of this real world evidence, causal reasoning, target trial emulation, these are very difficult to scale. I think there’s an art to when specifically in a drug discovery project each of these tools can enter the toolkit, specifically to avoid that “everything about a gene and it doesn’t help you make a medicine.”

Ross Katz: Yeah, that’s interesting. And we’ve talked in general terms about how you’re doing drug discovery with the phenotype in mind. I would love to focus on the example of diseases like Parkinson’s and Alzheimer’s, and if you’re able, I would love to hear the story of how the methods that you’ve shared and the data that you’ve gathered are applied to study diseases like those.

Brant Peterson: My personal interest, I’ll admit it, the reason I wanted to start at Valo was because I thought there were a couple of these diseases, and Parkinson’s front of line, where the variation between people was the dominant characteristic of the disease. And there’s this throwaway quote – if you talk to folks at Michael J. Fox Foundation, you’ll find this quote: if you’ve met one Parkinson’s patient, you’ve met one Parkinson’s patient. I think there’s a lot of diseases for which that’s true but Parkinson’s especially. The way the disease progresses is so variable. We have these incredibly poignant patient days where folks would come in and three or four people would come in and talk about their experience of the disease and they were just completely – what this person’s biggest challenge is, this second person looks genuinely surprised, like “I hadn’t thought of that.” And then you talk about what are you worried about? What scares you? And across the board the answer is I don’t want to lose my edge. I’m worried about cognition. But the disease progression, UPDRS part 3, what we measure in the clinic when we do a clinical trial – motor symptoms. It’s almost a different disease that folks are talking about worrying about versus what we’re measuring when we do clinical trials. It just felt like this has to be something we approach from the level of individual patients. This has to be something we first break apart in the experience of patients in the real world before we ever start thinking about medicine. One of the first approaches we took was that plus – why not include Alzheimer’s disease? Why not include MS? Why not understand this full neurodegenerative spectrum, neuro-inflammatory spectrum? We have so much evidence from the genetics of these diseases from the 2010s to really understand that there are overlapping drivers. Some of the same genes are coming up in all the canonical hit lists. Alzheimer’s has just been I think for me fundamentally recast as a neuro-inflammatory disease no matter how hard we try to knock down Abeta or tau. Putting these patients all together and asking what differs about these patients and what’s similar, we just immediately saw a really obvious constellation of dementia-related symptoms that co-clusters Alzheimer’s patients and not all but some Parkinson’s patients, as you’d expect. We saw subsets of Parkinson’s disease patients that just really clearly lit up an inflammatory profile. We saw subsets of Parkinson’s disease patients that looked very obviously mitochondrial metabolic just in what we saw. You could say oh Parkinson’s, the lysosome, the mitochondria, but it’s not everything in everyone. And you see it immediately in how it breaks out across patients. And then we’ve had some preliminary success even in trying to track those subgroups to progression patterns. We’ve been able to validate some of that in the sense that we saw some of the earliest onset patients were some of the highest risk for rapid progression, where that pattern’s totally flipped in multiple sclerosis. It’s the late-onset patients that progress the fastest there. That’s a really good orientation into the disease. I think where we are now is we’ve really focused in on a few of those patient subgroups – I can’t go into too deep of a detail there – but to understand that these are the unmet need. These are the patients that are experiencing progression despite medication, our best efforts do not meet the needs of these patients. And those are where we’re starting to map out how do those patients progress? We’re seeing that some of those subgroups really do have that canonical motor progression and that leads us to one path to the clinic for any mechanistic intervention we’d find there. We’re starting to look at the genomes of patients, the proteomes of patients from those particular subgroups and these become hooks for discovery there. On the other hand I’ll say we’ve also found some really unusual kinds of disease and I don’t know that we’ll be following up on them from a molecular perspective, from a target ID perspective, but one of the things I did not expect was that we also see a really clearly discrete subgroup of Parkinson’s patients that have non-temporally overlapping lifetime co-morbidity with neuro-psychiatric disease. When you really go digging in the literature there’s a couple of papers that are suggestive of this, but that wasn’t something I went in with. I think it immediately implies maybe not so much a drug discovery paradigm, although there could be shared mechanisms that you’d be led to if you thought about how those diseases interacted, but maybe more that this helps us appreciate a particular population of Parkinson’s disease patients that have in their deep past a fairly stigmatized medical history that defines a prodrome. We have this canonical constipation prodrome, loss of smell prodrome. I think there is for a subset of patients a neuro-psychiatric prodrome that’s just not as well understood that could really meaningfully enter into how we think about care for those patients.

Ross Katz: Thanks for the explanations and for the stories. What it really highlights for me based on what you’re saying is you’re going in and you’re exploring, first you expand the lens of what it means to study a disease and you understand the landscape of this constellation of neuro-degenerative diseases and then you break it apart and you understand what are the different manifestations of each of these diseases and who are the subpopulations of patients that are experiencing those different constellations of symptoms, and then choosing, as I understand it, subpopulations that have particular mechanisms that are potentially both understandable and potentially druggable for further exploration. At the very beginning you’re understanding not just what the mechanism is but who you might target and what the biomarkers are and the entire ecosystem. Am I thinking about that right?

Brant Peterson: And that takes us back to one of the things we said earlier. Sometimes it’ll be that this is a group of patients that we have to find a path forward for. These are the folks that nothing’s working. Sometimes it’s more like if I scan across this disease with a mechanism that I have a strong reason to believe is associated somewhere, where does it best sit? If I’m thinking about oxidative stress, mitochondrial mechanism, who am I looking for? And the answer is it’s not everybody. There are biases with respect to age of onset, there are biases with respect to male-female presentation differences, the constellation of specific progression to date in the disease will help you see. And then there are biomarkers that you can measure below that. We think of it as a map. It’s a way into the right neighborhood. And then turning that constellation of observations into a causal pattern through a lot of what we were talking about earlier starts to help you not just bring the hypotheses you have but generate new ones. I think that’s right.

Ross Katz: Yeah, that’s really interesting and I love how the framework that you’ve developed allows you to ask the entire landscape of questions that you would want to ask and attempt to ask them earlier in the process. Are there any applications of LLMs or AI that are really exciting from your perspective, ways that it’s being used internally that you feel like are helpful, or maybe ways where you feel like it’s not quite doing everything that it’s billed as doing? Just interested in your takes.

Brant Peterson: Super question. I think one version of the answer is that we’ve been training language models for over half a decade here. This is trivially true in the sense that the deep learning architectures that we’ve had some of the most success with on the really large scale data – when we’re taking a quarter of a million people for two decades, 60 million visits, 40,000 features per patient, when we’re taking these enormous data sets and trying to learn structure from this fundamentally longitudinally sparse, fundamentally conceptually sparse data – language is also fundamentally conceptually sparse, fundamentally longitudinally sparse, fundamentally non-real, non-matrix-like. It is an infuriating data type just like EHR. When we think about the modeling architectures that make sense, the NLP deep learning of the 2010s, early 2020s is a really appealing match. I think we haven’t so much seen a ton of success with direct token-to-token transformers. I think maybe that might be because we’ve got the wrong definition of tokens. It’s intuitive and natural to call the medical code the token, but it isn’t the order of medical codes in the record, not even temporally, that encodes meaning. When we think about meaning in patient journey, that meaning is much more like the episodes of care themselves. I think we’re tokenizing the wrong thing in the trivial transformer architecture. And in these older skip-gram architectures we’ve had a little bit more freedom to find the right level of meaning on which to encode the language-like concept. This is just a direction for us as a species to work harder on. We have some research directions here, we’re thinking about how to get that closer to the mark. But the goal remains the same: this extraction of a dense, real-valued numerical representation of synthetic concepts built up from and faithful to the input data but tuned by whatever additional kinds of tasks you want to put on the model architecture, in whatever end-to-end training you want to build. It is a remarkably flexible, remarkably powerful way to take a ton of patient data and turn it into something that’s amenable to more standard machine learning pipelines. From a language model-inspired approach, I think this is a real backbone workhorse of what we’ve done over the last half decade. From a language model as chatbot perspective, I will say I think if we’re appropriately skeptical and keep our heads on straight, there is a bunch of exciting stuff that actually does work. A trivial one is I just don’t write a lot of plotting code myself anymore. I just say I’d like these points to be green, please. I think it’s even more true when we’re thinking about visualizing graphs. I find NetworkX is not a trivial syntax to write and it is really easy to write in GitHub Copilot. These are simple things. I think there are some more interesting ones. I mentioned things that don’t scale well in our causal reasoning approaches especially in the real world include these really painstaking trench-by-trench real world evidence target trial emulation studies. We have some early results that I think are really promising around early study design details from general descriptions of hypotheses. This is not I think if you take a macro lens on the industry at all surprising given where a lot of pharma companies have seen success with AI and development is in protocol development. Writing these protocol documents – it’s not even a cottage industry, it’s an industry in the AI world. And I think we’re seeing a little slice of that as valuable here. It’s again early days. Really important to keep our metrics in front of us when we’re doing this, really important to be interrogating reasoning traces. But I think where we have these fledgling agentic systems for performing specific tasks, there really is an opportunity for those to change the scaling logic, change the scaling economics of some of these kinds of analyses.

Ross Katz: Yeah, that makes sense and also to leverage the expertise you have in-house more effectively because you have brilliant scientists spending less time writing visualization code and more time reviewing protocols that are being spit out and editing them appropriately. That makes a lot of sense. Brant, you’ve been a great guest. Thank you so much for joining today. Before I let you go, can you let everyone know where they can go to learn more about you and your work at Valo Health?

Brant Peterson: Yeah, we’ve got a website, valohealth.com. I think that’s a really good place to get started.

Ross Katz: Awesome. Well, Brant, thanks again for joining. I really appreciate the time and look forward to connecting down the line.

Brant Peterson: Thanks.

Jason: And that’s it for this episode of Data in Biotech. If you enjoyed the episode, please subscribe, rate, or leave a review in your podcast platform of choice. See you next time.

Frequently Asked
Questions

How does Valo Health ensure insights from patient data are genuinely meaningful, rather than just correlations?
Valo begins by deeply understanding the data's generative process, biases, and confounding factors before applying any models. They combine real-world evidence, genomic causal models like Mendelian randomization, and wet lab experiments to build a complete causal map, validating each link in the chain. This approach prioritizes understanding the 'why' behind the data over merely identifying 'what' patterns exist.
What's the biggest challenge in integrating real-world patient data with experimental lab data?
The primary challenge lies in bridging the conceptual gap: real-world patient outcomes are hard to replicate in a lab, while cellular mechanisms are difficult to observe in patients. Valo addresses this by using a structured causal DAG, where different elements of a hypothesis can be tested and calibrated using the most appropriate data source. This allows them to formulate specific questions that can be answered at either the lab or patient data level, then weave those answers together.
How does Valo decide which potential drug targets to pursue after identifying causal mechanisms?
Beyond biological validation, Valo assesses the 'advanceability' of a target. This involves considering practical factors such as whether the mechanism is amenable to therapeutic modulation, tissue specificity, and the overall feasibility of developing a drug. They prioritize targets where the evidence points to both a strong causal link to disease and a clear, pragmatic path to intervention and clinical development.

Need a data partner for life sciences?

CorrDyn helps biotech and pharma companies build the data infrastructure that accelerates research and operations.

Book an intro call