Skip to content
Naren Tallapragada — Organoids and Active Learning with Naren Tallapragada
Data in BiotechEpisode 39

Organoids and Active Learning with Naren Tallapragada

Naren Tallapragada of Tessel Bio discusses using organoids and active learning to reverse-engineer chronic disease treatments.

50:25Full transcript below
NT

Naren Tallapragada

CEO & Co-Founder at Tessel Bio

Overview

Biotech leaders face a critical dilemma: drug discovery is slow, expensive, and often fails because early-stage models don’t accurately reflect human disease. Traditional approaches force a choice between scalable, easy-to-run experiments that lack biological relevance, and complex human models too costly and time-consuming for broad screening. This inefficiency wastes billions and delays the development of life-saving therapies.

Host Ross Katz speaks with Naren Tallapragada, CEO and Co-founder of Tessel Bio, who explains how his team directly tackles this challenge by bridging the gap between biological accuracy and experimental efficiency. With a background blending electrical engineering, physics, and systems biology, Naren provides a quantitative lens on how to reverse-engineer chronic disease. His company employs mini human tissues—organoids—as more predictive models, then applies an active learning platform, Tessalogic, to accelerate target discovery.

This episode reveals how Tessalogic’s AI-driven approach selects the most informative experiments, achieving 5x to 30x efficiency gains over brute-force methods. Naren discusses the practicalities of working with complex biological systems, the explore-exploit trade-offs in experimental design, and where the industry’s focus on AI in biotech currently misses opportunities for deeper mechanistic understanding.

Key Takeaways

High-Fidelity Models Demand Smart Experimentation

Biotech’s challenge lies in balancing biological accuracy with experimental scalability. Organoids, as mini human tissues, offer superior disease relevance over traditional cell lines. However, their inherent complexity—from handling viscous Matrigel to measuring nuanced biophysical phenotypes like mucus flow—makes brute-force screening impractical. Success hinges on a method that intelligently navigates these constraints, selecting only the most informative experiments.

Active Learning Prioritizes Efficiency by Eliminating Redundancy

In drug target discovery, active learning platforms like Tessalogic don’t just identify the next best experiment; they systematically reduce the number of necessary experiments. By continuously updating predictions based on prior results and biological data, this approach helps scientists avoid unproductive experimental paths. Tessel Bio reports 5x to 30x efficiency gains, translating directly into conserved resources and faster progress in resource-constrained environments.

Benchmarking AI in Biology Requires Custom Metrics

Unlike fields with standardized datasets (e.g., protein folding), many complex biology problems lack universal benchmarks for AI performance. Tessel Bio addresses this by internally comparing its active learning against historical whole-genome brute-force screens. This allows them to quantify efficiency—for example, achieving the same top target identification with 3-20% of the experiments—providing concrete evidence of value for stakeholders.

Human Biological Knowledge Fuels AI-Driven Discovery

Active learning models become more effective when seeded with “inductive bias” from human biological understanding. Incorporating known signaling pathways, gene interaction graphs, and tissue-specific expression data as soft priors guides the model, even when starting a search. This integration helps both in generating initial hypotheses and in validating AI outputs, ensuring predictions are mechanistically plausible and reducing the experimental data required.

Related: CorrDyn helps companies in biotech and life sciences develop strong AI strategies to solve complex problems and reduce operational costs through data cost optimization. Read more about how biotech manufacturers gain more from their data.

Full Transcript

Jason: Welcome to Data in Biotech, the podcast from CorrDyn where we explore how companies leverage data to drive innovation in life sciences. Every two weeks, we sit down with an expert from the world of biotechnology to understand how they’re using data science to solve technical challenges, streamline operations, and further innovation in their business. Today we’re joined by Naren Tallapragada, CEO and co-founder of Tessel Bio, a company that is revolutionizing drug discovery through active learning and organoid models. He walks us through how Tessel Bio uses AI and experimental biology to reverse engineer disease, identifying new drug targets more efficiently than traditional methods. It starts with an overview on the power of organoids, mini human tissues grown in a dish, and how they provide a more accurate model of disease than other conventional approaches. He also breaks down Tesselogic, Tessel Bio’s active learning platform, and how it accelerates target discovery by optimizing which experiments are run next. If you’re curious about the intersection of AI, biotech, and drug development, this is an episode you won’t want to miss.

Ross Katz: Naren Tallapragada, welcome to the Data in Biotech podcast.

Naren Tallapragada: Thanks for having me, Ross. Really excited to be here.

Ross Katz: Well, just to kick us off, can you give us a brief introduction to your background and what brought you here?

Naren Tallapragada: Yeah, absolutely. My name is Naren Tallapragada, I’m the CEO and co-founder of Tessel Bio. We’re a drug discovery company based in Cambridge, Massachusetts, so it is sunny and warm and 25 degrees outside right now. Really my journey into biotech and into the role of CEO of a biotech startup has been a long and winding one, but really motivated by a lot of personal factors that made me make the professional decisions that ultimately led me here. I’m actually by training an electrical engineering and physics guy. That’s what I studied as an undergrad at MIT. I really thought I was either going to spend the rest of my life working on signal processing or solar panels, and then a couple of things happened. Most formatively in my sophomore year of college, my mom died of small bowel cancer, after a lifetime of having Crohn’s disease. She was also an intensive care physician who was a neonatologist—she took care of premature babies at the hospital. I spent most of my life, including the early years of college, really avoiding anything to do with biomedicine and gravitating towards the physics and engineering world that my dad lived in. But then I realized as a result of watching my mom’s experience as a doctor, as a patient with chronic disease and finally as a patient with cancer, that there’s a lot of things in biology and medicine that needed fixing, and perhaps hubristically, I believed that there was some way of taking all this math that I had learned and having an impact on the development of new therapies or at least deciphering and understanding the mechanisms of diseases like the one that my mom had. So that prompted me to start dabbling a bit in some bio research as an undergrad, ultimately applying to graduate school. I went to Harvard and got a PhD in systems biology. My particular flavor of that was basically a quantitative angle on how tissues regenerate and repair themselves. If you watch stem cells—this is a fun fact—we’ve all got stem cells throughout our body, even as adults, that are constantly regenerating and repairing all of or many of our organs. We know the ones that don’t get repaired quite well, for example, the neurons in your brain. But you have a brand new intestinal lining roughly every seven days. So there’s just this remarkable control problem, thinking about it through the lens of an engineer, of how you end up with the right number and mixture and placement and geometry and arrangement of cells and cell types with high fidelity in this system that’s undergoing constant turnover. That was a problem which was deeply relevant to the diseases that motivated me to get into biology in the first place, and that was also a problem that really appealed to my quantitative mind. Fast forward through a whole bunch of other experiences, and I end up here now as the CEO of Tessel Bio, where we are using many of the same techniques and experimental approaches that I first encountered and developed in grad school—so-called organoid culture—to model disease. I’m a big believer in seeing is believing, really trying to observe and understand how cells in a dish are growing and behaving in the context of health and disease and trying to use that as a very evocative and information-rich readout of how drugs work and how to make them better. The path that led me here has paid off in all sorts of very interesting and unexpected ways. There are things I did in grad school that have suddenly reappeared now—supplemental figures from my PhD which are suddenly relevant to our work—and it’s just a nice reminder that career arcs are long, but in the end, many things seem to happen for a reason. That’s perhaps a longer answer than you wanted, but hopefully one that gives us a lot to unpack.

Ross Katz: Not at all. Can you introduce us to Tessel Bio and how your background plays into the way you do drug discovery and build a platform there?

Naren Tallapragada: Absolutely. We like to say that we reverse engineer chronic disease—we harness the power of human physiology to reverse engineer chronic disease. What that really means is we are trying to discover new drug targets, new knobs and intervention points in human disease that we can actually develop therapies for—a small molecule or an antibody that touches on one of those knobs and has an impact on disease. That’s a universal quest in biotech. In our hands, we really believe that if you start with a disease-relevant phenotype—this is a word I’m going to be using a lot, phenotype—if you start with a readout in a dish which is as mechanistically similar to or as evocative of the actual clinical endpoint in a person as you are able to model in a dish, if you start with that and use that as the axis along which you are optimizing a therapy or deciding that a drug is working or not working or one drug is working better than another, you are likely to have a more impactful therapy that’s more likely to succeed in a person than if you go the other way around—scour the literature, put your head down, talk to as many academics as possible within a 20-mile radius of Boston because, obviously, there’s no world outside of Boston, Massachusetts. I’m kidding, by the way. But you talk to every key opinion leader and expert here and then you decide by some process of data integration that gene ABC1 is the most important driver or the most targetable therapeutic driver of a disease like COPD, and then maybe you spend a lot of time and money developing a drug against that target, only to find in a phase two clinical trial that perhaps you were quite wrong. Instead of doing things that way, we would rather take something that we know is broken in a patient with a disease like COPD—for example, how much mucus does your lung produce? It turns out you can grow these human lung cells in a dish and they produce a lot of mucus and you can measure things like the viscosity of the mucus or how quickly the mucus is flowing. These are biophysical readouts that are the same as what happens in a person with thick mucus they’re trying to cough up or thick mucus that might be blocking some of their airways in their lung. From that, you essentially do a systematic process of perturbing the system, taking some readouts, integrating that data to say, well, of all the knobs that we could turn here, of all the protein-coding genes in the genome or all the different intervention points we think exist in the system, what is the subset that is the most likely to be playing a role here? From 20,000 genes, can we then keep iteratively shortlisting down to a thousand and a hundred and ten until we find a top target that we believe is both disease-relevant, has the effect that we want, and is druggable. From there, you’re off to the races, and you develop one or more drugs and take them to the clinic—God willing and funding permitting. That’s our approach: reverse engineering the disease from the outcome to the targets, rather than from the targets in anticipation of the outcome.

Ross Katz: Can you introduce the audience to organoids and explain how you work with them?

Naren Tallapragada: Organoids—a nice and somewhat futuristic-sounding term for mini human tissues in a dish. There are many ways of taking what is ultimately derived from, or modified to be like, human cells in a patient. Let’s take an example. Let’s imagine, Ross, that you had COPD and that I did not. If we wanted to understand differences between your lungs and my lungs in this organoid culture context, you might take some cells from my airway—the biopsy process is rather painful—and there are many ways in which you can source these tissues. You could take some lung cells or airway cells from me, some from you, and then grow them in an environment outside of the body where the cells grow and can be expanded and passaged and frozen and all these other handling techniques that enable your typical cell culture, except that these are primary cells that have come from you as a person with disease or me as a healthy person. In coming from the original human, you are theoretically preserving many aspects of the disease context and the way that these cells would respond to drugs or the way that these cells are dysfunctional, and you can make a fair comparison of what is the same or different between a healthy lung and a diseased lung without having to go and directly experiment on a person. Fancy term for mini human tissue in the dish. Organoids generally refer to taking these cells and embedding them in—it’s basically tumor ooze from a mouse, but it’s representative of the connective tissue, the extracellular matrix in your body. If you embed cells there, they tend to grow into these beautiful three-dimensional structures that have all sorts of interesting complexity and behaviors and sometimes even geometrically look like little portions of an organ. If you take cells from a mammary gland, for example—representative of the human breast—you end up with these structures called acini in 3D, which look like the milk-producing substructures of the breast. If you take intestinal cells, which I’m quite familiar with, and you grow them in a dish, you don’t get a nice planar epithelium, you don’t get a tube unless you pattern it into a tube. Your intestine has this wavy pattern where you have these long villus-like fingers which give you a lot of surface area to absorb nutrients, and then you have these smaller horseshoes called crypts, which have the stem cells that sustain the entire tissue. If you take these cells from a human intestine and grow them in a dish, you end up with these little sacks which contain many of these little crypts and you can try to coax them to make those villus-like projections as well—there’s all sorts of fun biophysics around when that does or does not happen. You basically end up with organotypic cultures that look like the original organ in some ways, have the same cells as the original organ in some ways, express many of the same genes—that’s exciting—but they also carry out many of the same behaviors. In the lung setting, these cells produce and transport mucus, which is the same thing that happens in you or me when we’re sneezing and coughing. In the intestinal context, these cells actually transport things across the barrier, so you can model things like nutrient transport or drug delivery. In a fibrotic context—relevant to many diseases across many tissues—these organoids actually get stiffer and more tied up and twisted and harder to deform in the same way that the parent tissue would. You’re basically trying to mimic as much as possible about the parent tissue, but in a way and in a setting that you can actually perturb these things—poke them, see what they do, try something again based off of what did or didn’t work the first time.

Ross Katz: Could you talk a little bit about how you’re gathering the data and why your approach to drawing conclusions from that data is different than other companies that do what you do?

Naren Tallapragada: I think what you’re alluding to here is something that we like to say at Tessel quite a bit, which is that the best models of biology are really the hardest ones to deploy and scale. If you take that statement to its logical conclusion, a human being is the ultimate representative of that. You can’t really try a hundred or a thousand or a million different drugs in one human being with a disease and see which one they respond to the best. There’s both practical and ethical reasons why you can’t do that. On the flip side, what’s generally the workhorse of the drug discovery process is usually some sort of a cancer cell line which is a living human cell and contains many of the same proteins and components and various aspects of biological gadgetry responsible for the function and dysfunction of a living being. But it may be very alien to the context of the actual disease you’re studying. If you’re looking at a chronic disease like Crohn’s disease, sure, you could use a colon cancer cell line to study what’s broken about the intestinal lining in inflammatory Crohn’s disease, and there are certainly linkages in the biology of inflammation and cancer in that tissue. But it’s just not really representative of what’s happening in a 30-year-old patient with Crohn’s who doesn’t have any malignant transformation, doesn’t have any cancerous tumors growing in the tissue, but has a totally different set of processes at work that make the tissue leaky and inflamed and subject to constant onslaught from the patient’s own immune system. You really want to find a system that strikes a balance between the scalability of one of these cell lines—don’t literally sneeze on the plate, please, if anybody goes and does these experiments—but even if you sneezed on the plate, the cells would be okay and they’d keep growing and you could poke them ad infinitum. On the one hand, you want as much of that flexibility as possible, but on the other hand, you want something that’s as representative of the human biology as possible so that you’re not predicting things about the wrong mechanism entirely. Organoids really do give you that nice happy medium. But they’re still not going to be as scalable as the cell lines. There’s all sorts of considerations here. If you’re growing an organoid in the tumor ooze I mentioned—people particularly use this substance called Matrigel—that’s one of these really weird substances that’s actually a liquid when it’s cooler and solidifies when it’s warmer. You actually have to keep these cells in these reactions on ice at 4 degrees C and it’s a very viscous fluid that’s hard to pipette, and if it’s hard for a human to do that, it’s really hard for liquid handling robots to do that. Not to mention all the phenotypes I mentioned—all the things you might want to measure about what the cells are doing. Those are pretty complicated phenotypes. If you want to measure a cancer-related phenotype where you’ve got a bunch of cells growing in a dish and you add some drug and now they divide less or they grow less or more of them died in response to the treatment, that’s actually fairly easy to quantify and there are ways of doing that in a robotic and somewhat routine way these days. But if you want to measure the speed at which mucus is flowing on top of a complex three-dimensional layer of lung cells, that is a far more experimental readout, far more complicated, subject to all sorts of technical considerations around optics and keeping the camera in focus, and subject to all sorts of biological and practical considerations as well. There’s not really a machine-readable SOP that you can just stick into a robot, set it, forget it, walk away, come back the next day and have tons and tons of data that you can play with either conventionally or in a machine-learning setting. What we think about is what can you do to try to get the best of both worlds? What can you do to work within the limitations of sample constraints, handling constraints, the fundamental complexity of these complex in vitro models, but make them a little bit more amenable to the scale of perturbations and the repeated follow-up that’s required to actually discover an interesting target or have enough confidence in a disease mechanism in order to go and develop a drug against it? That leads us into our active learning platform, which we call Tesselogic, which is really all about making predictions of what the best next experiment is that you should do based off of the history of everything you’ve tried before and the priors you may already come into the experiment with—I think these pathways are important, I think these genes are important, I just tried knocking out ten genes that I selected randomly or with a very strong bias towards a pathway or a mechanism I think is very disease-relevant. Whatever the case, wherever it came from—did it come from a machine’s prediction or from a human scientist’s brain?—you should be able to synthesize that information and intelligently make a decision about what to test next. This gets into all sorts of fun things around explore-exploit trade-offs and whether you want to do an experiment in order to clarify your model of the world and which genes you think play a role in this disease, or whether you’re trying to do a speed run to your druggable target as quickly as possible because you’re trying to spend as little time and money on discovery as possible and get to the stage of clinical development as quickly as you can. These are all ultimately considerations and trade-offs in anyone’s drug discovery program—it’s not unique to Tessel—but the way we approach it is using what is effectively a co-pilot that empowers our scientists to make the best use of these complicated systems and complicated readouts that we think are really predictive of human outcomes, but where you might only have a hundred or a thousand shots on goal or you might have to do things iteratively in batches. You can’t go and brute-force them by dumping every chemical that’s synthesizable by some vendor somewhere on these cells and then stick it into a robot and hope for a dataset that helps you in a brute-force way find exactly the well and exactly the concentration that is your magical hit for a drug campaign.

Ross Katz: My understanding of active learning is that somewhere in your model you need to have a set of the universe of possible perturbations that you might do to the system, and then some either prior or informed estimate of what the impact of each of those perturbations might be, and then error bars surrounding that impact across the entire universe of perturbations. Is my mental model of what information that model has available correct, or in what ways are the models that you’re using somewhat more complex?

Naren Tallapragada: Your intuition is spot on. You have a parameter space that you are optimizing over. In our case, we’re saying out of all of the genes that might be interesting drug targets, or out of all of the genes that might have some causal effect on the phenotype we’re looking at—the mucus transport, the fibrotic stiffness of the tissue—which of those genes is the most likely to have the largest effect size with the lowest uncertainty in that prediction? Ultimately, that’s what you’re after, that’s what an optimal output of this model would be. The thing about active learning is you generate that iteratively and you can do experiments that are focused on reducing the uncertainty and you can do experiments that are focused on having a more accurate estimation of the effect size. But you do need to have some prior or a predefined landscape of the knobs that you can turn. In our case, the simplest way of thinking about it is that there’s roughly 20,000 protein-coding genes. Viewers from the future who read a paper that said there are a hundred thousand human genes, or viewers from the past who said there were a hundred thousand human genes, who want to say I told you so—please don’t hold this against me—but right now in 2025, call it roughly 20,000 protein-coding genes. We know that maybe our universe of targets sits somewhere in that set. We know from lots and lots of work by lots and lots of groups—standing on the shoulders of giants—that these genes interact with one another. There’s prior information in the form of protein-protein interaction graphs. You know that A and B actually interact with each other and maybe that interaction was observed, if you’re very lucky, in the tissue and cell type of interest—maybe somebody went and did a study in human primary intestinal cells and observed that these two proteins actually interact, or maybe somebody saw that it was theoretically possible in a test tube. Prior information and public data around how different aspects of genetics and biochemistry work really spans that gamut. The pre-training data you have available to you is quite variable in type and quality and context. It tells you what’s maybe theoretically possible; it’s not necessarily complete. There may be interactions or complexes or reactions where A actually interacts with D or there’s this other protein E that plays a role, which are not represented in public data. It’s absolutely possible that you’re missing information, it’s also absolutely possible that you have too many edges and your priors are wrong. What we do is we basically define the space over which we’re trying to find an optimum—which of these genes is the best target, or how can we generate a list of these genes in some rank order by effect size and uncertainty. But then as far as what we can use to explore that space and update the model based off of the data we collect, we provide all of this public and proprietary information that we generate as we do experiments as a soft prior. So A is known to interact with B and C and D, and if I do an experiment and I happen to knock out gene A and there is a very strong phenotypic effect, that might make me upweight the genes in its neighborhood, B, C, and D. If it doesn’t have a phenotypic effect, it may make me downweight those neighbors or reconsider the structure of the graph—maybe A sits downstream of B, they do interact but there’s some causal ordering here and A is actually an upstream or downstream effector depending on what you see. You can make those sorts of updates to your prior knowledge just as much as you’re making updates to the likely effect size of intervening on any one of these nodes. It’s a very interesting and complex optimization problem. It’s not the smoothest optimization problem. As far as what the possible answers are—the ultimate outputs of the algorithm, what gene is the most important—you do have to define that space a priori. As they currently stand, our model is not going to come out and say there’s some splicing variant in some mRNA and that’s your drug target. In theory you could make a prediction about something like that if you incorporated that level of detail or complexity into our system. We choose not to do that for now for the sake of actually accomplishing our goals and getting a directional cue as to where there’s an interesting therapeutic target. But in principle you could make this as complicated as you want. There’s a whole bunch of publicly available information, as well as information you end up generating along the way as you do your own experiments, of how different genes interact with each other, which genes have correlated expression in cells, which genes may or may not actually be expressed in a tissue of interest. This is a common problem with public datasets—you’ll see that gene A and gene B interact, but then if you go and look at what genes are actually expressed in lung tissue, neither gene A nor gene B are expressed, so having that interaction in there was a total distraction. There’s all sorts of reasons why you might want to tweak your priors or update that landscape while you do the optimization. But the ultimate output of the optimization is constrained by what you put in as the potential answers—your model is going to predict what you allow it to predict, basically.

Ross Katz: I’m interested in how much of the mechanisms are embedded in the model and how much of that you rely on human intuition or human in the loop in terms of helping to take what’s coming out of the model and determine what the next best experiment might be.

Naren Tallapragada: It definitely helps to add—you might want to call it inductive bias—into your model. You know that there exist signaling pathways, you know that generally one gene is upstream of some other gene in that signaling pathway, so there’s at least some sense of a causal ordering that ultimately should lead you to predict that some gene at a certain point in that pathway would have a larger or smaller effect just because of the way that influence is propagated through the graph. That’s definitely something we incorporate into our modeling approach—the structure of signaling pathways, or at least the principle that there is such a thing as a signaling pathway. Generally biology tends to organize itself that way and solutions that pop out that look like that are preferred to solutions that are a lot messier and less structured. I’d say this works at two levels. In order to observe that a gene of interest is having an effect on a phenotype of interest, all sorts of things need to be true about the gene, the way that gene is doing stuff in the cells, and the phenotype in the experimental system that you’re using. The gene needs to be expressed, it needs to have all the partners that it interacts with, the way that information is propagating downstream of that component need to be basically similar. And then the thing that you’re measuring needs to be similar to what happens in a human being. The ultimate mucus secretion or transport by coordinated action of cells that secrete mucus that’s just the right thickness and cells called ciliated cells that beat and transport that mucus—that coordinated action also needs to happen at least in some locally coordinated coherent way. That’s where you really lean on the organoids and the cell model and the experiment to make sure that all of the variables, parameters, and components that you might be optimizing are actually present and active and tunable. That’s really important—that’s one. The second is at the modeling level. I’ll give you an example in these lung cells. There’s this signaling pathway called Notch signaling. It’s been known for a really long time to have a whole bunch of important roles in a whole bunch of tissues in the body. In the lung context, it basically helps determine how many of these mucus-producing goblet cells are present. It’s a really interesting mechanism for a quantitatively minded person because it’s a little bit like the checkerboard pattern on your shirt. Imagine there was some self-organizing way by which the white squares and the blue squares, before they decide who’s white and blue, talk to each other and say, you’re already more in the direction of white, then I’m going to be blue, and just by virtue of that local coordination, you end up with this global pattern. It’s really quite elegant. This particular pathway is at work in so many different human tissues, which is why it ended up being a poor therapeutic target. But of course, when people said this modulates the production of these mucus-secreting cells in the lung, why don’t we develop a target that goes after Notch signaling? So you’ve got two things there. One is we know a lot about the Notch signaling pathway and how that intersects with the ultimate phenotype of producing mucus. That’s really useful prior information for how to structure the models, how to come up with initial guesses of what perturbations to try, because even if they’re not the drug targets, you know that they are going to generate a phenotype and if you can measure other data types or have a sense of what else might be connected to Notch signaling, you’re off to the races—you can find alternatives that are safer, better, whatever. But a lot of the pathways you know about that play a role in driving the phenotype are not the ones that you can safely drug, because Notch signaling is just as central to the formation of blood vessels in your body and to the function of your nervous system as it is to this beautiful mucus-patterning process I just told you about in your lung. You don’t want to trade mucus obstruction and chronic cough for seizures or a vascular tumor or severe diarrhea or any of the other consequences that might happen when you inhibit that pathway. That is where the signaling pathway gives you a good prior for initial perturbations to try, it gives you a structure from which to start doing a local search for other targets, but you’re not constrained by that—it just gives you a sense of directionally where to go next. It integrates or overlaps with human-in-the-loop insights for two reasons. One, the human might say, I remember from grad school or a paper that Notch signaling is important, so I’m going to try a Notch perturbation. That’s one way a human might help seed the active learning process. But on the flip side, if you didn’t provide that information and you said, computer, tell me your best guess as to the genes that have the biggest effect size on mucus production and cell type patterning in the lung, you would hope to see some Notch signaling-related things in that list as a sanity check that your model is not completely overfit to some weird feature of your data or giving you predictions that are in no way interpretable or verifiable. The priors help inform how you evaluate that your model is doing a good job, the priors help you come up with good initial guesses even if those aren’t the ultimate answers, they’re good places to start that you know are strong hammers with which you can whack the system and make it do stuff, and the priors also really do help with data efficiency. If you know there is some order or structure or hierarchy to how different genes interact, then you actually have some sense of which perturbations are the most efficient to try first, or where in genetic space it may not be worth looking. Those are the different ways in which all of that informs what we do.

Ross Katz: I want to zoom out a little bit. But before I do, one of the things that your white paper highlights about the Tesselogic platform is the efficiency gains you get with this active learning approach. Can you give us a brief window into what are the efficiency gains you saw and how you determined what those are?

Naren Tallapragada: In our white paper—I’ll zoom out even more and give you a meta-comment. There are two things I think are absolutely true about AI and bio that I wish more people focused on. One, there’s more to life than protein folding, really. There are phenotypes like the ones I’ve been going on about, like this mucus secretion and transport by the coordinated action of different cell types that are emerging from the stem cells and how they behave. That’s one aspect of it. The other aspect is that because there’s so much diversity in what you could measure, there are also not necessarily agreed-upon benchmarks for anything under the sun. At this point, if you look at something like AlphaFold or an AlphaFold competitor, there are all of these ways of assessing that one model is better than the other, more data efficient, runs faster—all these really useful benchmarks. When it comes to most problems in AI for biology, especially the ones focused on the biology rather than the chemistry—focused on the targets and modeling what the cells are doing more so than, did I predict the strength with which this small molecule stuck to this protein—in that setting, you have to come up with the benchmarks yourselves. This is a comment on how we think about benchmarking our performance. The ultimate true honest benchmark of any biology project or target discovery project is: did you take that target, turn it into a drug, test it in a person, and see that it succeeded in a phase two clinical trial? That is the honest way of assessing that your target discovery campaign worked. The problem is that takes a long time, it’s very expensive, and if you hold yourself prisoner to that benchmark, you’re going to go nowhere fast or let the naysayers win and just continue to do drug discovery the way that it’s done right now, which I would argue has been variably successful. So how do we benchmark our performance on a much shorter time scale with less cost? One fair head-to-head comparison is: if I am a discovery scientist doing an experiment in a system like an organoid or a cancer cell line, how much time could I save myself? How many experiments could I have gone without? How much in the way of reagents did I not need to purchase? How many nucleic acid constructs did I not need to buy—guide RNAs I didn’t need to make and stick into cells in order to knock down or knock out some gene? On that time scale of doing a target discovery screen where I take a bunch of cells and hit them with some sort of a CRISPR knockout strategy or a whole bunch of compounds, how many fewer conditions could I have tested to get to the same top answers? The best way we could think of to publicly benchmark that was to compare ourselves against settings in which people did these large brute-force screens. There are public databases of these—there’s one called BioGRID ORCS, which is the one we really highlight in our white paper, where a lot of people did whole genome screens where they knocked out or knocked down genes in the genome one by one in a whole bunch of cell contexts. There’s a lot of cancer cells in there, there’s also people who were studying secretion of insulin by human pancreatic beta cells. There’s a cool diversity of tissue types and phenotypes represented. In that setting: could you have tried 3,000 knockouts and found the same top 100 targets? Could you have done 300 knockouts and found the same top 100? When we say efficiency gains, we were referring to this range basically from 5x to 30x efficiency gains with essentially V0 of our active learning algorithm. You could have done 20% of the experiments and found all the same top targets. You could have done 3% of the experiments and found all the same top targets. Our modal performance is somewhere around a 7x or 8x speedup, which is great if you’re in a resource-constrained environment. If I’m a small startup and I want to do a CRISPR screen, even if the cell culture system is not so complicated, even if I don’t care about combinatorics, I just want to get to some interesting answers with as little money and time as possible—this is really useful. But when you start thinking about those two big existential settings—the complex in vitro models, the organoids, the organotypic things, the ones where you can’t stick it on a robot, set it and forget it—or you think about combinatorics: nice, one gene, but now I want to test every pair or combination. Now you’re dealing with scenarios in which it’s just not practical for anyone, even if they had an infinite amount of money, to test every single condition. There are 20,000 genes, but 20,000 choose 2 is roughly 200 million. Or you have these organoid systems in which it maybe takes eight hours to do an experiment or to get the Matrigel to pipette in just the right way, and you can maybe set up a thousand wells but you can’t set up a million and walk away overnight. These are settings where the efficiency gains really translate into even being able to do the experiment at all. There’s really a sense of: if you used our algorithms and fed the data from the screen into our model iteratively gene by gene or batch by batch, as you would if you had been running the experiment yourself—you make a prediction and say okay now I’m going to try this next set, now this next set—how quickly do you end up with the same top N targets? You can vary different parameters and see how that varies across top N. You can do fun things where you start off by essentially saying I got my model of the world completely wrong—what happens if you got the ranking in exactly the opposite order? If the first gene you fed into the system was actually the one that ranked the lowest in the list? And yes, it’s obviously not as high performance as if you happened to be absolutely right with the first gene that you picked, but we’re talking about 6x versus 7x efficiency gains. There’s very clearly a sense in which biology is compressible, in which you don’t need to try everything in every system, and in which the history of what you’ve already tried is a useful predictor—perhaps even the most useful predictor—of what you should try next. That’s true across a very wide range of parameters and hyperparameters and philosophical choices of how to do science. We’ve benchmarked it publicly, but then we are using it. We dogfood our own product. We are using it to drive our own drug discovery campaigns. You might guess correctly that we’re working on lung disease, and there are other tissues as well where we’re thinking about this both internally and coming up soon with some pharma partners, knock on wood. We’re using it and benchmarking in real-world performance prospectively on targets we’re discovering, on experiments we’re running, and reassuringly, it’s already better than just brute-forcing it. Our own prospective measurements of efficiency based on the data we’re generating are lining up nicely with the retrospective estimates we made from public data—admittedly with a benchmark that we cooked up, but I think it’s a decent benchmark, and if anybody out there wants to talk about benchmarks for biology prediction problems, that’s one of many fun things I’d be happy to discuss over a beer or a coffee.

Ross Katz: We have to bring this conversation to a close, but I want to touch on one of the things that you just said—you were talking about the AlphaFold and the protein-folding approach, and then you talked about your approach and this cost, quality, and speed trade-off that everyone’s trying to navigate, where on your end it’s higher quality and slower but you can manage costs using that active learning component, and on the AlphaFold end it’s faster and relatively inexpensive to get lots of data in there. Can you give us your perspective on where the attention is going in this ecosystem and where you think the opportunities are?

Naren Tallapragada: That’s an astute observation about the simplex over which everybody is optimizing—that nice triangle. I think there’s room for everybody. I’m not knocking approaches like AlphaFold which have really been transformational in so many different ways. You no longer need, for an initial guess at how a protein looks, some poor postdoc to spend years using heavy metals to try to crystallize a protein. That’s still really important and useful in the process of understanding the biology of that thing. But the fact that a computer can give you a decent guess pretty fast and might help accelerate your own hypothesis generation and research—especially for this large unannotated part of the genome—is very important and useful. There’s been a lot of attention in that direction, especially from investors and ML types, for a variety of reasons. One, there’s this very large high-quality public dataset in the form of the Protein Data Bank, which went into training these models. There are all of these benchmarks, a lot of things can be done essentially purely computationally, and the end result of those models is understandable to multiple types of investors in terms of how you’d actually make money on it. If I come up with a very interesting protein or drug that binds to the protein at the end of the prediction problem, I have hard IP that I can go and try to commercialize. And if I don’t want to make IP, I’ve got a very interesting software tool with all this whiz-bang ML in it, so maybe I can commercialize access to the model or the software. I understand why people’s mental models align well with commercializing protein language models and AlphaFold and all of that. But if we really care about the ultimate reason why we’re all in biotech in the first place—which is to make a difference in the lives of patients with disease—you have to ask yourself when is it really important that we have new targets, that we really understand the mechanism of biology, that we don’t wait for 30 or 40 years’ worth of academic research to come out before we try something new. Where can we maybe systematically take a few risks or go a little bit faster in probing new biology? Not just the stuff that happens to be served to you on a silver platter in a format that is amenable to computer modeling and that makes sense to the computer scientist’s brain. There’s no easy answer to that, there’s no template for how to take complex messy squishy biology problems and turn them into something that is machine learning-friendly or accessible to both the engineering and biology communities. But I believe that is where the key challenge lies for using AI to transform drug discovery and development and make a difference in the lives of people. That is where Tessel is doing our small part, trying to be a little contrarian and push the envelope where we feel like there’s not necessarily been enough attention.

Ross Katz: For listeners who are interested in learning more about Tessel Bio, where should they go?

Naren Tallapragada: You can go to our website, www.tessel.bio. The white paper that Ross has mentioned is linked there. You can always email me at Naren, N-A-R-E-N at tessel.bio. I’m also occasionally active on Twitter at @ntallapragada. Three different ways you can reach out, and I do encourage you to reach out, whether you’re interested in discussing some of the ideas I’ve brought up here or think you might have something useful to contribute to our team or the effort I’ve put forth in today’s conversation.

Ross Katz: Naren, it’s been a pleasure having you on the podcast. Really appreciate the time and look forward to connecting down the line.

Naren Tallapragada: Thanks so much, Ross. Really appreciated that and loved the opportunity to speak with you today.

Jason: And that’s it for this episode of Data in Biotech. If you enjoyed the episode, please subscribe, rate, or leave a review in your podcast platform of choice. See you next time.

Frequently Asked
Questions

How do organoids improve drug discovery over traditional models?
Organoids are mini human tissues grown in a dish, mimicking human physiology more closely than immortalized cell lines or animal models. This biological relevance allows for more accurate disease modeling, predicting drug efficacy and toxicity with higher fidelity, and ultimately reducing the risk of failure in clinical trials.
What kind of efficiency gains can our drug discovery programs realistically expect from active learning?
Tessel Bio's Tessalogic platform has demonstrated 5x to 30x efficiency gains compared to brute-force screening methods in target discovery. This means achieving the same high-quality list of top drug targets with a significantly smaller number of experiments, saving substantial time, resources, and reagent costs.
How much human expertise is still required when using an active learning platform like Tessalogic?
Human expertise is crucial for active learning. Scientists provide "soft priors" like known signaling pathways and gene interactions to guide the model, which helps structure the search space and validate predictions. This collaboration ensures that the AI's recommendations are biologically plausible and relevant, while enabling scientists to make more informed decisions faster.

Need a data partner for life sciences?

CorrDyn helps biotech and pharma companies build the data infrastructure that accelerates research and operations.

Book an intro call