Listen on
Overview
The prevailing approach to applying foundation models in biotech, particularly LLM-centric methods, often treats a patient as a document—a collection of text reports. This fundamentally limits the depth of insight available for clinical decision-making and drug development. Reducing rich, multimodal data like imaging, genomics, and real-time physiological readings into text-based summaries loses critical information, leading to less accurate predictions and hindering the ability to understand a patient’s dynamic health trajectory.
This gap directly impacts R&D efficiency, the precision of clinical trial design, and the ability to predict treatment efficacy with confidence. It means biotech and healthcare organizations are missing opportunities to use data to its fullest extent. Kevin Brown, Founder and CEO at Standard Model Biomedicine, with his background in brain-computer interfaces, medical imaging AI at Siemens Healthineers, and oncology data at Bristol Myers Squibb, recognized this challenge after a “GPT-3 moment” that illuminated the path to a more sophisticated approach.
In this episode, host Ross Katz, Kevin discusses how Standard Model Bio addresses these limitations by building multimodal foundation models. He explains how these models integrate diverse data—from CT scans and genomic assays to electronic health records—into a unified, time-aware patient representation. The conversation covers why a shared latent space for patient data, coupled with predictive architecture, offers a superior pathway to understanding patient trajectories and the impact of interventions, contrasting this with the shortcomings of text-only models.
Key Takeaways
Beyond Text-Centric AI: Multimodal Models Capture Lost Patient Information
Treating a patient’s medical history solely as text documents or reports (e.g., radiology reports, clinical notes) discards valuable, raw signal data. Converting diverse modalities like imaging, genomics, and physiological readings into text is a lossy process, reducing the information content available for predictive models. Building a shared latent space that integrates these modalities directly, rather than through text, preserves data richness and leads to more accurate and complete patient representations.
Patient Embeddings Map Dynamic Health Trajectories for Predictive Insights
A patient is not a static snapshot, but a continuous journey. By constructing a single vector embedding for each patient, updated over time and across various modalities, models can predict future health states. This temporal understanding enables counterfactual reasoning, allowing clinicians and researchers to model the expected impact of interventions and assess if a patient’s trajectory aligns with desired outcomes, such as treatment efficacy or risk of adverse events.
Open-Sourcing Accelerates Validation and Broadens Foundation Model Utility
Despite commercial potential, open-sourcing foundational biomedical models is crucial for achieving widespread validation and diverse benchmarking. This strategy allows numerous academic and clinical institutions to apply and evaluate models on their specific patient populations and use cases. This collaborative, rapid feedback loop builds confidence in model reliability and accelerates the overall pace of discovery and application across the biotech community.
Foundation Models Require Domain Alignment and Local Evaluation for Specific Applications
While foundation models provide a powerful base, their predictive accuracy for specific conditions still relies on alignment with relevant, in-distribution data. A model trained primarily on oncology data will require fine-tuning to perform optimally for neurodegenerative or cardiovascular applications. Effective evaluation necessitates local benchmarks tailored to specific patient populations, rather than relying solely on broad, general benchmarks.
Related: CorrDyn partners with biotech and life sciences companies to architect data engineering systems and advise on AI strategy that truly reflects patient complexity. We focus on building data capabilities that support accurate, time-aware models.
Full Transcript
Jason: Welcome to Data and Biotech, a podcast from Courty, where we explore how companies leverage data to drive innovation in life sciences. Every two weeks, we sit down with an expert from the world of biotechnology to understand how they’re using data science to solve technical challenges, streamline operations, and further innovation in their business. Here we go.
Ross Katz: Kevin Brown, welcome to the Data and Biotech podcast.
Kevin Brown: Hey, thanks Ross. Great to be here.
Ross Katz: To kick us off, would you mind giving us an introduction to your background and what led you to Standard Model Bio?
Kevin Brown: My background is originally in pure math. Then I realized I did not want to be a mathematician, for a lot of reasons, although it was quite beautiful. I joined a brain-computer interface lab in college. That was absolutely awesome and very cool. I went to grad school for that and got to work on early versions of what we would call Neuralink today: invasive neural implants to control robotic arms or virtual versions of those. While I was there, we had to segment gray and white matter in some brains to prepare for surgeries. Deep learning was just taking off, and man, it worked! It was one of those “the world is different now” moments when I saw the first few results, and I thought this is what I should be spending the rest of my life doing. Some people have described that in neuroscience the first time they heard neural recordings. For me, it was the first time seeing segmentations that just worked. I joined Siemens Healthineers to work on computer-aided diagnosis, which was absolutely awesome. That led me all the way to foundation models, and ways we can unpack that.
Ross Katz: Was there a specific GPT-3 moment where you realized that foundation models were going to be inevitable in biology? Can you walk us through your thought process there?
Kevin Brown: It was really important to record not just from neural recordings, but neural recordings from all over the brain. Not just cells, like neurons firing, but also diffusion-weighted imaging, structural MRI, optogenetic stimulation of the brain. All those things kept adding more and more information. The more modalities and other sensing ways about a system that you could engineer and put together at the same time. While I was doing this brain segmentation, I realized deep learning would render some of the interpretable approaches we’d been taking a little bit less performant. When I had that “aha” moment, I left grad school to join Siemens Healthineers because I knew this was the future. When I was at Siemens, we had scale and we started to notice that it was less about thinking: what particular layers do we need? Should we use this activation function or that one? That period was rapidly ending. We realized we were building systems that could be used millions of times a year. We’re sitting on a lot of data. What can we do with that? Perhaps that’s the simplest thing to do. So we did that. We built a federated learning system that to my knowledge is still running, which is awesome. That was for computer-aided diagnosis for lung cancer. But I remember looking at a lung nodule in a CT scan for lung cancer diagnosis. I thought, it would be really cool if we could combine genomic assays, pathology information, or longitudinal electronic health records. All these are relevant and may even bear down on whether you expect that to actually be a lung nodule or a false positive, or indicate specific treatment regimes that would work for a patient. The response was, “Look, we make tools for scanners and for radiologists primarily.” That group, which is an incredibly noble thing. It drives down the cost of health care, it increased quality, it lets us spread globally in ways that are otherwise difficult to do given the labor shortage of radiologists and the technical expertise required. That’s a very awesome thing. But I really wanted to do this multimodal thing. I joined Bristol Myers Squibb, where they had one of the world’s richest oncology datasets, rigorously recorded at the cost of billions of dollars through many pivotal, really amazing clinical trials. That was an absolutely awesome experience, partly because, in Pharma, the domain expertise is just higher than almost anywhere else in the world. If you want to talk to a world expert in cell therapy or gene therapy, you put time on their calendar, you call them up and then you just ask questions and they’re happy to talk. You want to talk to someone who really understands how imaging works in clinical trials or should work or could work in the future, you talk to those folks. Likewise for biostatisticians. I was able to just have some of the most productive knowledge-immersed times of my entire life, including undergrad and grad school, just walking the halls of BMS. There was a GPT-3 moment, which was literally a GPT-3 moment, where I read the GPT-3 paper. We had been building data science models in drug development—human data predicting outcomes of clinical trials, or response to therapy for a patient. Those were great, but I saw this one figure in the GPT-3 paper in particular, where it shows the performance on several downstream tasks, like a hundred. Some of them go from zero percent accuracy all the way to a hundred at different scales; sometimes it takes six billion parameters, sometimes 13, sometimes more. I thought that’s going to happen in biology. We’re going to make these multimodal. It’s all going to go into one giant foundation model, and we’re going to figure out a way to make it work. I don’t know exactly how it’s going to look, but that’s what’s going to happen. That was the GPT-3 moment, and the rest of the next couple years were spent figuring out where and how to do that.
Ross Katz: Okay. With that background, can you give us an introduction to Standard Model Bio and how you think differently about foundation models for biology?
Kevin Brown: Certainly. I realized we needed multiple modalities, and that goes back to my time in grad school. Then I also realized we needed data scale. There was a third element that I started to understand when I was at BMS: that domain expertise, that last 10%, was critically important. It’s not something that every foundation model can do all of. You can’t own predicting outcomes in metastatic melanoma, non-small cell lung cancer, early breast cancer, and cardiovascular disease—especially if a patient has a previous history of a GLP-1 or a diabetic history. All of those details are critically important, and no one can build a model that does everything, but maybe you can build one that facilitates any downstream task. I thought, “Where would be the best place to do this? And where would I get the data to do it?” I thought, “Well, maybe I do this at BMS.” BMS is a fantastic place with incredible data, but that data, like every big Pharma’s data, is highly biased—and it should be, right? It’s mostly your drugs, mostly early trials, mostly failed trials, and older standard of care, not a lot of your competitors’ drugs. That limits the generalizability of any foundation model you want to be very general. I thought it needs to be more than just these vertically specific things.
Ross Katz: Okay. In terms of how you think about foundation models that facilitate downstream tasks and enable a variety of different use cases. One of the things I’ve heard you argue is that the patient is not a document. Can you unpack that and give us an understanding of how you think about the patient mathematically, and why thinking about a patient as a bundle of text is fundamentally limiting?
Kevin Brown: When we first started this, we thought we would use this with text, right? We thought we’d use a large language model backbone. You’d be silly not to use an LLM at all for biomedical applications, because they’ve already been trained on all of PubMed, right? And most of the scientifically published literature, for better or worse, makes them tremendously powerful tools. So we thought, language is a very general space, right? It’s almost by construction designed to describe anything. So maybe we should just map everything into it, whether it’s radiology images through radiology reports, whether it’s pathology images through digital pathology reports, whether it’s any other molecular report. That was fairly fruitful. It’s not that difficult to do. To give you an example, we used an existing pathology foundation model, published by Bio Optimus. We took that, we combined it with a Llama model, and in two hours, you can get state-of-the-art for medical image understanding at the time, on a single H100, including debugging, right? We were thinking, “This is the future!” Then we realized that when we wanted to incorporate other modalities that weren’t as readily mapped to text, we were running into some walls. My favorite example of this is whole genome sequencing, right? You have 3.6 billion base pairs, and you’re trying to compress that into text, but most of the text you have that’s linked for any one of those things is “KRAS positive,” “EGFR positive”—great, right? ROS1. That doesn’t really give you a rich enough way. If you take some of these contrastive approaches or the traditional vision-language or other modality-language models, it wasn’t really going to work to then feed it into a language-aligned embedding to add to the context of the LLM. We switched a little bit to a JEPA-style model. That was one of the first things that really felt right. We thought, “We are truly respecting the complexity of what it means to be a human patient, right?” Every physician, every clinician, knows that a patient is not just what’s written down in the notes or even just communicated at rounds. There are other subtle things. For example, “we assumed X, Y, and Z because of their previous history,” or things that don’t get written down, or maybe they just get said out loud separately. There are visual things, right? that an epileptologist can see when they look at EEG traces, right? And they don’t always make it into the report. So we thought, “This works really well for radiology and pathology, but for other things it doesn’t.” Let’s map it into a shared latent space. We still use language models. Those language models can interpret physician notes or electronic health records, map them into language space, and then that language gets projected into this shared embedding space. Likewise for images, likewise for digital pathology. This actually turns out not to be that difficult, because every single AI model, more or less, maps input data to some compressed representation that’s a vector. You just need to find a way to take that vector and map it onto the shared internal vector. That really clicked for us. We thought, “What’s the best way to pre-train this, right? How do we make sure that we’re getting a fairly good model?” Should we use data corruption techniques, like masked language modeling, for instance, or masked image modeling? We were really inspired by some of Yann LeCun’s work with the JEPA models, Joint Embedding Predictive Architecture. One of the things we loved about that is it naturally fit into this “patient is not a document” perspective, where we take many modalities—imaging, pathology, EKGs—map them into a shared space, and then predict what that space should look like at time T+1, T+2, and T+3. That almost by definition shows a trajectory of a patient in this abstract “patient space,” defined by the different measurements you could make about that patient. It’s never limited to the ones you may have for that patient. If you have an EKG, great, it can shed a little more light on where they are in that space. If you don’t, that’s fine; you marginalize over it. We felt like this was the first thing that really respected patients and also let us think about patients as moving temporally. They’re not just a snapshot of the document you have at a particular time. It’s not just about predicting the next word you’d see in a note or answering a question that may go into a radiology report. It was really about, “Here’s the patient today; where do we expect them to be?” That’s what we were doing.
Ross Katz: What I’d love to do is say back to you my understanding of how the model works and perhaps work with you to define some terms, just to make sure everyone is coming along with us. First, ultimately, what we’re trying to do is create an embedding for each patient based on their medical history up to that point. This embedding captures all the information from all the different modalities—all the different ways we measure that patient’s health throughout their journey—from CT scans to blood work, to genomic profile, clinical notes. All of that ends up in a single vector, a single list of numbers that says: this is who this patient is at this given point in time. The JEPA approach that you’re talking about, my understanding is that rather than trying to make the patient noisier and have the model predict who that patient is, what you’re doing is trying to predict what that embedding will be at some point in the future based on what’s happened previously. You’re stepping through time and saying, “In 2020, this is how we thought the patient would look in 2021.” In 2021, this is how we thought the patient would look in 2022, and so on, up to the current moment. You can then take that embedding and predict what that patient might look like in 2027 based on their medical trajectory to date. Am I thinking about that correctly?
Kevin Brown: Yes, exactly, you’re thinking about that correctly. One of the things it affords once you get into this regime is that you can build anything off of these embeddings that you want. Just like every AI model takes the data it has seen, produces a compressed representation, and then does something with it—whether it’s a classifier head (basically putting logistic regression or a Softmax head on top of a classifier), regression to a single number, or a diffusion head for producing an image, like in a stable diffusion-style model. We know that’s almost like Play-Doh that anyone can build things off of, which is really cool. But the temporal nature of it—what gets us excited—is how in physics, sometimes working in the space of differential equations, makes things easier, right? When you think about the equations of motion, it’s easier to specify them in terms of change over time and space rather than mapping out the explicit trajectory, right? Integration is hard, but differentiation is easy. It’s easy to say, “If I perturb this a little bit this way, this is where I expect the ball to roll or the planets to move.” I really like this very loose physics analogy, right? But when you think about the patient, I feel a little bit sometimes like Kepler looking at the stars and trying to figure out the equations of motion, right? Writing down in a notebook, “Here the stars were in this position and that position and this position,” and then trying to figure it out—and then once it clicks that you can describe these things in these very elegant mathematical ways in terms of how you expect them to vary over time, everything makes sense, right? Going through classical Newtonian mechanics. We thought, “I want to get to that place for a patient where we say, ‘This is where you are, this is where you’re going,’ and then you can start doing these counterfactual games, right?” With the appropriate statistical rigor—that was so drilled into me at BMS—I came out with massive respect for the biostatisticians there. You think about the difficulties of really designing systems that ascertain what’s true. But I love this idea of counterfactual reasoning: “Here’s the trajectory you’re on.” If we do this intervention—whether it’s exercise, a GLP-1, or a cancer treatment—we expect you to go over here. Then you can say, “Not only did I do counterfactual reasoning, or I jittered some things to see how likely it was that I was going to go each time with Monte Carlo sampling.” You can say, “Am I on the right track? I did this intervention, which means I expect to be over here. Am I actually over there? Is this therapy working?” I love this idea of thinking about the good spaces of patient space, where longevity is expected. And these bad attractors—like serious adverse events, cardiovascular events, a heart attack, or metastatic disease—and trying to push a system away from that. You can start thinking, “Maybe these therapies are basically bumps, and they work for a period of time—like a targeted therapy that works great in lung cancer for a few years before you develop resistance.” This is a system, a way of thinking about it, that is elegant and also respects the fact that patients change over time. They have many measurements that matter. They’re not just what’s written down in one specific document. They’re not an ICD-10 code.
Ross Katz: I want to spend some time on the modalities aspect of this, then return to the temporal nature of it and how you think about counterfactuals and how a patient moves through time. My understanding is that you’re taking these very different data types, from radiology images to clinical notes, and treating them differently within the model architecture. You’re utilizing different assumptions in the way you process each modality to get the most information from that particular modality as it’s being brought in. Also, there’s a way of handling this that means you don’t have to have every modality for every patient to effectively model that patient’s trajectory. Can you help me understand how those different modalities are treated?
Kevin Brown: Each modality gets its own specific encoder to begin with, right? EHRs can be ingested by language models readily. Then likewise, let’s say a radiology image. We use vision transformers in 3D; it takes about three lines of code to build them today, right? It’s a fairly straightforward thing. Of course, when you want to do a little more fancy thing, you have to think about various data augmentation techniques. We’ve, for instance, trained on a million CTs—maybe 1.5 million now, but a million oncology-related ones. That allows us to encode various kinds of CTs and then put them into that same latent space. We’ll be coming public with that soon. The idea being that we treat computer vision in a way that we know has good standards for how to treat it. We don’t have to reinvent the wheel there. If there’s a better radiology model, we can also use that. Likewise digital pathology, we use other digital pathology models right now. The field is awash in them. It’s a really intense and crazy space where there’s all this churn; one week you see a $50 million model, and then the next week you see someone training a very similar one for a few thousand. We thought, “You know what? We’re just going to use what other people have produced at this point.” And that’s totally fine, right? You can ingest digital pathology data using best practices. With all the careful thought that has to go into that, and then map that into that shared latent space. Likewise with an EKG. It could be EEGs, brain MRIs, chest X-rays. The important thing is we don’t have to develop every single encoder ourselves, because that’s too big a bite to chew. But we do that in imaging. We do that in genomics. We are doing that soon in multiomics and other things. And there’s a number of reasons for that. The other thing I love about it is that if you’re missing a modality, that’s okay, right? This naturally pairs with the way clinicians have to operate. Sometimes you don’t have a test you’d want to do for a variety of reasons, or it’s missing, or it didn’t come at the right time, or maybe the endoscopy didn’t get what you wanted to see, and you have to think about how that information would get built in otherwise. That means you’re not totally bound to these very low-scale regimes where you only have highly paired data for every single sample. That’s unrealistic, and you’re not going to get the scale you need to build the size models you want to really reason about these things. Then I want to spend a little time on why we care so much about so many different modalities. What if you just did a vertical model: “Here’s my radiology model. Here’s my pathology model. Here’s my EKG model. Here’s my genomics model,” and then just integrate them later? I think that doesn’t work for a number of reasons. One, because even at a low level, you get some interesting cross-attention and cross-feedback across these different things. Two, clinicians know they don’t even operate in these intense verticals. There are some silos, but even a radiologist doesn’t just produce a report, and that’s the end of it. Often at a serious hospital, there will be feedback. You can call them up. You can ask for clarification. You can double down. You can say, “The patient was reporting this; take a look.” There is that kind of crosstalk. So it makes sense to build models that also have that communication built in. I think that is the future. One of my favorite examples of this is cardiotoxicity, to show just how broad these things can be. Cardiotoxicity in oncology. In radiation therapy, you can have adverse events from radiation therapy, especially in the cardiac sense. You can triage patients according to how likely you think they’re going to have that kind of event. You can look at pericardial effusion on a CT scan. You can combine that with routine electrocardiograms. You can also combine that with the longitudinal EHR, which would obviously show different kinds of risk factors—whether it’s as simple as body mass index, previous obesity history, or smoking status. All of those are relevant. The clinician is going to compare all of those things, and why not do it in a raw signal space if you’re building a model? It also shows how interdisciplinary we know this model has to be. So the verticals can’t even be oncology, right? Because even in oncology, now you’re in cardio, right? And those cardio models are going to get better, even if they’re just trained on general cardio patients. Likewise, you’re also going to be in immunology. Not only because of immunotherapies and the side effects you have there, but there are gastroenterologists that only see oncology patients. There are cardiologists that only see oncology patients, and it’s only going to get more complex. So we didn’t really see there being a limiting factor there. Then, once you go that wide in the horizontal, there’s no way you’re going to own all these application layers. You cannot credibly claim you’re an expert in cardiotoxicity for early lung cancer, progressive metastatic disease in melanoma, and general cardiology. That’s insane. But you can do that first 90% in a very similar way.
Ross Katz: That makes a lot of sense. The goal is a foundation model that is the best, most complete representation of the patient possible. And there’s information in all the different modalities that can inform that representation of the patient through time. To ignore that is to fundamentally restrict yourself from getting that full representation of the patient. Another question that I had about the modalities is, how do you prevent a high-signal modality like imaging from dominating noisier modalities like genomics?
Kevin Brown: That’s a great point. This actually plays into the temporal aspect. This is a problem we’re actively working on. If you’re an oncology patient, you’re going to get several CT scans most likely. They’re going to be at regular enough intervals, so CT is a great place to go. That’s why it’s the other modality we went into, besides longitudinal EHRs, because they’re the most regularly paired. Then, at baseline, you may have a sequencing run. But it’s not going to be all the time. Likewise, a resection with pathology is very important at baseline. It’s not going to happen at the same frequency as a CT scan. There are ways you can get better clarity of where a patient is going on the trajectory, and that can be updated with seeing what you expect to see in the MRI. The other thing we do, which is important, is that you have to ground the model a little bit. You need to either be able to reconstruct the signal you would expect to see. Am I seeing a plausible MRI at this point in time for where I am in embedding space? Am I seeing what would be a plausible doctor’s note or ICD-10 code at that point in time? Am I seeing an EKG that looks like I expect? That’s a pretty important thing to make sure the embedding space stays relevant and doesn’t collapse into repeating trivial embeddings. And that was a major unlock for us.
Ross Katz: Right, if I’m understanding correctly, you’re using a joint objective when training the model. You’re using the supervised version of the model, where you’re trying to predict the next “token”—the next thing that would be in that patient’s medical record. And you’re also using JEPA as you described earlier.
Kevin Brown: Yes, that’s right. What I also love about that, which I think makes it possible, is asking: if I take this additional measurement, will I meaningfully change where that embedding would be or what something else would look like in the future? Is it going to inform what I would see with other modalities? That lets you optimize the kinds of tests you would run. Whether that’s in clinical trial development, picking the biomarkers you would want. You can figure out which ones are redundant because they’re not changing the expected trajectory of a patient. So don’t do a measurement. Or whether it’s in a system that’s eventually clinically facing. That makes me pretty excited. Because everyone likes to say, “the right drug for the right patient at the right time,” which is cool. But it’s also “the right measurement for the right patient at the right time,” which I think informs all of those decisions.
Ross Katz: Another question I had about the modalities is: is one of the downstream tasks figuring out what the hypothetical information gain would be from taking a particular test, from a particular modality, for a particular patient, in order to improve the embeddings?
Kevin Brown: Yes, we would love to do that. There are all these nuances for doing it. We can play around with them, do these basic assessments, and we do. We are actively looking for really awesome biostats collaborators that would love to take it and use it. I’d also mention that in terms of these downstream tasks, all the models are open source with very permissive commercial licenses, right? You can download them. You can use them and do various things with them. So we really want everyone to take it and do beautiful validations with it, and beautiful downstream tasks. We’ll even help people do it. We worked with one group; their paper is not public yet, but it took them a couple of hours to get a positive result. Then they submitted the paper a month or two later, which was fantastic. Likewise, we’ve repeated that with some other groups. The other day, one of the highlights for me was I just told Claude, “Go download the Standard Model and apply it to this dataset from this paper.” I didn’t provide any links, right? It searched the docs, pulled it down, and it didn’t require any intervention. I thought, “Okay, that’s cool,” right? I feel we’re now really speeding up biomedically in a meaningful way. It made me happy.
Ross Katz: The tools available to fine-tune foundation models across tasks have improved dramatically. What that ends up doing is expanding the potential list of end-users who could think of a problem that a time-aware patient embedding could help inform. I’d love to get to some of the tasks you think would be most appropriate later, but one of the things I’d love to understand beforehand is, when you think about the temporal nature of the model, how do you separate something like disease progression? The model is picking up on: “this patient is likely to have some progression in some disease they’re experiencing,” versus “there’s some treatment effect that happened earlier that is altering the patient’s trajectory.” I was having trouble thinking of: if the tumor is shrinking after chemo, is the model learning chemo works, or is the model learning this type of tumor was regressing? How do you think about counterfactuals in that kind of situation?
Kevin Brown: We think deeply enough about that to know that we should have other people think even more deeply. Part of that is, causal data analysis is really hard. Understanding the heterogeneous treatment effects—like treatment effects for a specific patient—is really difficult because you can’t go back and give the same patient a different treatment and see which one worked better. You can’t go back in time. That makes it really difficult. There are very good statistical techniques to do that. They do require a lot of nuance. I think it does permit that kind of analysis, but you have to be really careful—going back to the BMS days—or you will lead yourself down a road that is not statistically sound.
Ross Katz: Another thing about the temporal aspect of it: we talk about time in terms of T-minus one, T, and T plus one. But sometimes you’re trying to predict what the patient is going to look like in two years, and sometimes you’re trying to predict what the patient is going to look like in six months. For downstream users of a model like this that are trying to predict the next thing happening, how do you advise people to think about time in the context of the model?
Kevin Brown: That’s a great point. We would love to make it very time-agnostic, from very small scales to very high scales. Where I would love to see it go is: “In the next hour, what should I do with this patient?” all the way to “In the next year, what should I do with this patient?” Or for people building systems that do that. I do think that will be difficult. We’re not in the real-time sense yet of, “Should I do this test?” But there are scenarios where you would want to be able to make those decisions, right? With a patient that’s likely to go into status, and that’s going to change different treatment things you could do, or has a sepsis likelihood. But I think one place it could be cool is applied in a cool way. There was some LLM-based work at NYU Langone that did a lot of this, which I thought was really awesome. The senior PIs on that work are Keng Anjo and Eric Orman there. Absolutely fantastic researchers. They did implement, in real time, being able to predict things like mortality in prospective shadow mode, which I thought was great. I would love to see it used to reduce alarm fatigue, right? What is the most important thing to alert a clinical team to for this patient at this point in time? Or what is the test I need to do to avoid those things, and how can I queue up the requisitions for the things that will need to be done? We’re not there yet, but I would love to see that because you’re right, patient-relevant time scales can go from seconds—in a stroke—to years in terms of an exercise intervention.
Ross Katz: Even though you can’t know what the time scale you’re operating on is, it allows you to predict T+1, see what the embedding looks like, then predict T+2, T+3, T+4, and so on. In the absence of intervention, this is what the patient’s trajectory looks like through amorphous time. You can also inject a study or an intervention in the middle of it and see how that influences the patient’s trajectory. Am I thinking about that correctly?
Kevin Brown: Yes, that’s right.
Ross Katz: Let’s go back to foundation models versus vertical models for a minute. I know you talked a little bit earlier about how medicine is not naturally narrow to these disciplines; there’s information from each discipline and each modality within each discipline that feeds into the overall picture of the patient. Can you talk about what evidence out there supports the claim that narrow models are more brittle, or that the approach you’re taking for these broader foundation models is definitely the right way to do it?
Kevin Brown: Formal evidence? I wouldn’t go so far as to say there’s a registered study for it. But a lot of it is just reasoning through it from first principles. We know that multiple modalities are relevant for patients, or we wouldn’t have tests from those modalities. The question is, what is the best point to integrate them? The way they’re integrated right now is largely text, and sometimes follow-up text, right? A radiology report that gets read, a few words that describe a genomic assay, right? For example, “EGFR positive.” Those are very valuable today, but integrating them at the text level is problematic; by the time you get something all the way down to text, you’ve reduced the content. Even in a normal LLM, right? Predicting the next token, one token has less information than the embedding that you use to predict that token. So it makes sense to communicate across these modalities in embeddings in one way, shape, or form, because there’s so much more richness there. You can only ever destroy information by converting it into text, even if it makes it more legible for a human. Then the question is, where should that communication happen? Should you do it all the way at the end, like with the vector you would use to predict the next token? Should you do it earlier? If you do it earlier, do you get benefits? That’s what hasn’t really been rigorously assessed yet. I know there are some papers and research that look at these levels of fusion, but I think some of it hasn’t really explored the scale at which you could or should be doing these things. So I think it’s going to make sense to have multimodal models. Now, medicine does get siloed, right? There are tracks, obviously, and there need to be for reasons of education. But I do think that in the future—and I don’t know when—some of the communication that happens right now by text, either by physicians or by agents, right? Or you could have a radiology-to-text report generator. Those reports then get fed either to another human or an agent that is interpreting all that text and combining it with text from the genomic assay. I think in the future, it’s not going to be an agent or a human at the top pulling the text. I think the communication is going to happen implicitly at a lower level from modality to modality, from a model that is naturally communicating because there are embeddings that are talking to each other—from radiology to pathology, to genomics, to general patient information. The human will definitely always be in the loop, but I think the communication can happen at a lower level because text is such a lossy medium. Right.
Ross Katz: Right, the argument here is that every small piece of information gathered throughout the patient journey should be treated as gold, as really valuable. And that this embedding space you’re constructing across modalities is designed to capture as much of that information as possible and not lose any of it through the ways humans communicate with each other inside hospitals, inside EHR systems. And that EHR systems could potentially benefit from having systems that reference this condensed representation of all that information.
Kevin Brown: Yes, that’s right. Because this isn’t new when you talk to a radiologist. They know there are things about that scan that they’re looking at that don’t make it into the radiology report for a number of reasons, right? It’s there. They can see it, right? They can look at a tumor margin and think, “Hmm, that doesn’t look necessarily as good as I would want,” but how do you quantify “looks icky,” right? You know what I mean? But it’s there. Likewise with digital pathology, it’s not just about counting a PD-L1 measurement. It’s a little more nuanced, and some of that nuance doesn’t get put in for a number of reasons, but it is there and can be picked up. I think the labels for it are difficult, but if you have a self-supervised task—like predicting where this patient is going—that’s highly patient-relevant. So it’s meaningful. It also, I think, naturally pulls out those things that matter for that patient, even if they don’t get written down explicitly.
Ross Katz: That makes a lot of sense, and it also highlights for me the way our healthcare system is trying to get everything into discrete space. Everything gets a code, everything gets a category. It’s a binary: you’re either diagnosed with the thing or you’re not. But the embedding representation can capture a more continuous journey of what those different radiology images look like, which say, “We’re moving in the direction of a diagnosis that looks like this,” even if it’s not diagnosable at time T-minus three. Yes, that’s right.
Kevin Brown: I completely agree.
Ross Katz: I want to talk a little about the training data that goes into this, and also how that plays into how a researcher—who’s sitting on a bunch of patient data—would interact with the model to get these patient embedding representations out. How does the way the data is prepared to go into this model relate to that?
Kevin Brown: That’s a great question. It takes way less data normalization or curation than it used to. In fact, when we’re faced with a table, we back it out into text anyway because it’s easier for the LLM to ingest. So don’t be scared away by thinking, “I’m going to need to spend hours making sure all my columns line up to whatever I think the Standard Model is using.” I don’t think that’s the best way to think about it, because it’s going to go back into text anyway. Anytime you have those raw modalities, it’s fine. There’s a relatively limited amount of preprocessing you need to do. And I think that makes it really good, as long as the data is accurate and high-fidelity. That part, of course, still needs to be done, because sometimes things get mislabeled. We’ve seen people accidentally write micrograms instead of milligrams, and those are obvious to correct when you see them, but the model may not know that. So there’s a need for review, but broadly speaking, it’s not that difficult. And we’re happy to help anyone do it, but it’s not crazy.
Ross Katz: Is there any concern either on your part, or do you think there should be concern on the part of potential users of these models? The data you’ve trained the model on is from a fundamentally different distribution than the data they might be feeding into it. Is there some data shift happening there?
Kevin Brown: That is a fantastic concern. Right now, we are oncology-forward, although we’re moving rapidly into immunology and cardiology as well. That being said, if you downloaded this and said, “I’m going to use this model, which is really good at oncology, to predict neurodegenerative outcomes in my Alzheimer’s disease trial…” You’re going to have some issues without fine-tuning that model on some Alzheimer’s data. I think that would be one thing. I think it goes to a broader point, right? Any foundation model is a compressed representation of the data it was trained on, right? That’s by definition. So we want to be a latent, pseudo-data layer so that you don’t have to go out and buy all that data necessarily. To train your model, you don’t need to buy thousands of CT scans, or millions. And I think that makes it exciting or easy for people to get up and use, but if you’re trying to use that model—no matter how good it is—if it wasn’t trained on cardio data, it’s not going to get a cardio.
Ross Katz: If it’s a task that is not in distribution for the data you originally brought in when you trained it, you should be thinking about how you can fine-tune this model and then evaluate it. Which brings me to the question: what are some of the ways you evaluate the quality of the patient embeddings you’re producing, and how do you think about interpreting how good the model is at different downstream tasks it might be asked to do?
Kevin Brown: That’s one of the core questions, which actually doesn’t get answered enough, right? We’re at this period where everyone says, “I want to build a foundation model.” They train a model, and sometimes it’s only evaluated on one or two things. And that’s not really what you would necessarily want or need. You really need to evaluate it on lots of things. But the problem—the reason why people don’t do that—isn’t that they’re lazy; it’s that they don’t have the data, right? So you need to go to other institutions and say, “Can you evaluate this on your use case, on your patient population?” Taking one step back, if I’m a patient and I want a model that’s powering some AI that’s making decisions for me, what makes it good, right? Okay, we say the embeddings that application layer is using are good, fine. But how do we know that? It doesn’t really matter to a patient how well it works on average. It matters how well it works for patients like them, right? This is a fundamental problem everyone has. It’s not new, right? When you go to a clinical trial, you have a particular population, but that may not be what it’s like to be applied for a lifelong smoker who grew up in, say, Greenville, Alabama or New York City, right? We know that socioeconomics and various demographic factors have real effects in patient populations and affect the way things are applied, whether at the drug level or the AI level. We need to evaluate these things very locally. I’m not the only person who said this. I think Stanford has been saying things like this. Local evaluations are good. We shouldn’t just have huge benchmarks that say, “This is the lung cancer benchmark. This is the interstitial lung disease benchmark.” We need ones that are really specific to the patient populations they serve, which is why we’re happy to give the model to different people to evaluate it. I actually think that needs to be done more, because you see these patients and it’s, “Well, we trained it on this data.” It was really good. We got good results, huge scientific insight, but we really only know if it works at this one institution, right? It needs to be validated on more places, I think.
Ross Katz: Practically speaking, based on my understanding, it looks like you use the data you have to get embeddings out of the model, either through inference or through fine-tuning and then inference. Then you use a prediction head on top of those embeddings. It could be as simple as a regression on the different features in the embeddings to predict a particular outcome, and you ask, “How predictive is this linear model, or more complicated model, of the outcome you care about?” And that’s how one would think about evaluating the quality of the underlying embedding representations. Am I thinking about that correctly?
Kevin Brown: Yes, you’re exactly right. So, can it do something useful? In oncology, let’s say I’m doing time to event. If I put a Cox proportional hazards head on top of this model that takes these embeddings and spits out the hazards… Can I use that to get a higher concordance index, or time to AOC, or whatever the appropriate metric for that particular question is? Then some other model, right? That would indicate whether it’s useful, right?
Ross Katz: That makes a lot of sense. From a use case perspective, can you highlight some of the ways your open-source models have been used in the wild so far?
Kevin Brown: Okay. One is early-onset pancreatic cancer. Another is thinking about cardiotoxicity. The other is toxicity for general treatment, whether it’s sarcomas, pancreatic cancer, or non-small cell lung cancer. Those are the first places where we’ve seen a lot of the applications. Some things relevant to clinical development: Can you predict when a patient is going to have a line of therapy transfer that would make them eligible for a clinical trial? Can you predict whether they’re likely to meet the inclusion/exclusion criteria for that trial at a specific time in the future, rather than just right now? Eventually, HCPs can take the model or applications built on that model and start those conversations early. Because you expect a patient to go on a certain trajectory, you can start figuring out what the best trial for them might be, for instance. Or eventually, one of my favorite places is thinking about metastatic disease for oncology: can you predict certain kinds of adverse events? I think it’s going to be a huge thing, right? And I think it’s going to be fun.
Ross Katz: The other thing that’s occurred to me as we’ve been talking about the embedding space is that the embeddings are so flexible in terms of the things you could do with it. Yes, you can feed it to a deep learning model to predict something. Yes, you can run a regression on them to predict something. Yes, you can do a clustering and segmentation analysis to see who’s close in embedding space and who might have the right patient profile for a clinical trial. There’s a lot of potential downstream use cases for this kind of thing.
Kevin Brown: That’s right, even if it’s just for communication. For communication, you could say, “Here’s a patient that was similar to you; here’s what happened with them.” I think that could be a powerful thing.
Ross Katz: Very interesting. You’ve done a lot of work on these models, but you’ve chosen to open-source them. Can you give us some insight into why you decided to go the open-source route?
Kevin Brown: For now, honestly, it comes down to validation and benchmarking. We need models like these to be broadly benchmarked by not just one academic medical institution, no matter how prestigious, not just 10. By the end of the year, we want this validated in one way, shape, or form on a hundred different academic medical centers. That’s absolutely doable, but it’s also really important. It makes it frustrating as a scientist if I put my scientist—not my CEO—hat on. If I’m thinking, “Here’s this other model. Is it better than ours? Is it not better than ours? Is there a place for some overlap? Can we make a comparison? Which one should I use?” I can’t get their data, and I can’t get their model, so I don’t know, right? The only thing we could do when faced with “you can’t control what other people do, you can only control what you do” is… What we can do is give them our model. We can say, “At least you’re aware of it.” You can take it, benchmark it, and we will help you. We’ll dedicate hours from some of the finest ML engineers in the world, and we’ll help you do that. That makes me feel pretty good at the end of the day.
Ross Katz: What I’m interested in from your perspective: you already mentioned you’re oncology-forward. How would you want people to think about the limitations of the Standard Model Bio models that are being released?
Kevin Brown: Right now, they’re oncology-forward. We chose that for one business reason: oncology is a huge market. But it’s also the thin end of the wedge when it comes to precision medicine. You will have people write grants that say, “as done in precision oncology,” for instance. So we knew that was a good place to start if we wanted to work with people who have been thinking about this very deeply for many years. That means if you’re using it, it should be oncology-adjacent. It doesn’t have to be oncology right now, right? In the sense that if you want to predict cardiotoxicity, perhaps that would make sense, even though it’s not necessarily a cardio thing. Or if you wanted to use your own cardio model to predict how things would happen in an oncology space, that could work. We are moving into general chest disease with a collaborator we’d love to announce soon, which we’re really excited about. And I think that’s going to be a pretty powerful thing because it’s going to get us into some of the immune-related things, which are also cancer-relevant, right? Interstitial lung disease can be a devastating side effect to certain immunotherapies. We are thinking about oncology today; in a couple of months, it’s expected to be a lot more.
Ross Katz: Should I think about this as: you have a model architecture that you think is universally applicable, that models the patient in a really great way, but your journey is one of incorporating more patients with different therapeutic issues and potentially weighting modalities differently based on that cohort of patients and their journey through their patient experience? Is that right?
Kevin: Yes, that’s right. If someone says, “I would love to use this, but your model isn’t working that well…” We want to know. Sometimes we use that to identify what data to go after next, right? Someone says, “I really wanted to use this on predicting myocarditis, and I can’t.” Tell us, and we can try to hunt that down.
Ross Katz: Kevin, you’ve been an excellent guest. I really appreciate your time. If listeners want to learn more about your methods, the work you’re doing, or where you’re heading from an organizational perspective, where should they start?
Kevin Brown: Honestly, they can go to our Substack. That’s where we publish a lot of stuff even before it makes it into papers. So, blog.standardmodel.bio is one of my favorite places, or just email us.
Ross Katz: I highly recommend that Substack. That was how I discovered your work and really enjoyed having you on to talk about it today. Thank you very much, and I look forward to connecting down the line.
Kevin Brown: All right, thank you, Ross.
Jason: And that’s it for this episode of Data in Biotech. If you enjoyed the episode, please subscribe, rate, or leave a review in your podcast platform of choice. See you next time.






