Listen on
Overview
Most drugs work by holding onto one protein. Stop taking the pill and the disease returns, because the protein was blocked rather than the thing producing the problem removed. Adam Freund’s position is that the target was wrong: cells are the functional units of tissues, so they are also the functional units of disease, and a cell’s behavior is governed by a signaling network evolution made redundant on purpose. Block one node and the network routes around it, which is a reasonable account of why chronic disease pharmacology delivers modest gains on a permanent dosing schedule. Remove the cell and every pathway it was driving goes with it.
Adam Freund is Founder and CEO of Arda Therapeutics, which uses single-cell sequencing to find and target the pathogenic cells behind chronic disease and aging. He holds a PhD in Molecular and Cell Biology from UC Berkeley and spent seven years as a Principal Investigator at Calico Life Sciences, where he founded a research lab on the biology of aging and helped grow the company from 15 people to more than 200. His work spans experimental and computational biology, including functional genomics, single-cell RNA sequencing, and machine learning.
In this episode, host Ross Katz and Adam work through what made the approach practical and where the analysis goes wrong. They cover why B-cell depletion is the proof case and why Rituximab’s failures in lupus taught the field where to measure, a statistical shortcut in single-cell analysis that lets one oversampled patient produce the significance of a hundred-donor study, the neighborhood analysis Arda uses instead of clustering, the six converging axes that move a population from enriched to credibly causal, and why depleting cells can hold a chronic disease in remission for months after the drug has cleared.
Key Takeaways
Redundancy is why single-target drugs underperform, and why removing the cell works
Cells resist perturbation by design. Buffered stability is the foundation of homeostasis, and the corollary is that a highly effective single-target drug is hard to find, because blocking one node leaves nine other routes open. Freund’s read is that this explains the shape of chronic disease pharmacology: incremental benefit, dosed forever, fighting the stability the network is built for. If the cell has taken on an aberrant set of characteristics, taking out the whole node beats arguing over which of its ten redundant pathways to inhibit.
B-cell depletion is the proof case, and Rituximab is the lesson about where to measure
Oncology established that removing cells works, but cancer is the easy case because the bad cells are the ones dividing uncontrollably. B cells were the first setting where the pathogenic population had to be identified and then proven causal by depleting it in the clinic, and George Schett’s CD19 CAR-T lupus case study only published in 2021. Rituximab had repeatedly failed in lupus, so it took years for the field to believe CAR-Ts and T-cell engagers could succeed where it had not. The list now looks obvious, covering lupus, systemic sclerosis, MS, RA, and myasthenia gravis, and Freund’s point is that the clarity is recent, which is why extending depletion past B cells is only being attempted now.
The number of cell states is a dial the analyst sets
Clustering groups cells by transcriptome similarity, which is a sensible way to reason about a million individual cells. The problem is that cell state is an abstraction rather than a biological ground truth, so the number of clusters is decided by a metaparameter. Freund has turned that knob from two clusters to 2,000 on the same data with his own hands. He also names the incentive: one reliable way to get a Nature paper is to find a new cell type, and one reliable way to find a new cell type is to keep raising the cluster count. His standard for where to stop is the specificity a therapy can achieve, since resolution beyond that adds nothing you can act on.
If significance does not scale with donor count, you are not measuring disease biology
The usual next step is a chi-squared test per cluster, comparing disease cells against control cells. That test draws power from cell count, not donor count, and one patient can contribute tens of thousands of cells. Hold cells constant and a dataset with one patient and one control returns the same p-values as a study spanning hundreds of donors. Microarray, bulk RNA-seq, and qPCR were all already powered by donor number; single-cell data introduced a second dimension that makes it easy to treat the cell as the unit of variability, which is the right unit for basic cell biology and the wrong one for telling disease from health.
Neighborhood analysis scores cells without drawing boundaries
Arda’s alternative starts from what the ideal output would look like: a per-cell measure of how disease-enriched that region of the expression space is, crossing clusters and cell types rather than being confined by them. Neighborhood analysis gives each cell a score from the donor identity of its nearest neighbors, then propagates outward through second, third, and fourth-degree neighborhoods. Guilt by association, with a correction for donor representation so that five cells drawn from five different disease donors score differently from five cells drawn from one. The result is a continuous enrichment map over the whole dataset, independent of any cluster annotation.
Enrichment is a filter; causality takes six converging axes and a molecule you can build fast
Causal cells are almost certainly inside the set of expanded populations, the same way the protein you want to block is almost certainly among the upregulated ones, so enrichment narrows the field without settling anything. Arda layers on which populations produce factors already known to drive the disease, cytokines in inflammatory disorders and extracellular matrix in fibrosis, plus GWAS loci weighted by how strong the heritability signal is for that condition. Roughly six axes converge to leave one to three suspect populations. Freund is direct that the hard part comes after: a dozen plausible surface markers usually fall out of the analysis, and narrowing twelve to the one worth betting millions on, then pairing it with the right modality, is where the work lives. A modular depletion library of benchmark depleters, CD3 binders, validated ADC payloads, and CAR-T-ready formats is what keeps the platform repeatable rather than a run of one-offs.
Circulating cell counts are a red herring, so test in human tissue
Rituximab depletes 100 percent of B cells in peripheral blood, down to undetectable. So do CD19 CAR-Ts. Rituximab does not fix lupus and CAR-Ts do, because the causal cells sit in lymphoid organs, marrow, or the affected tissue rather than in circulation. Freund expects the same to hold for the T cells, mast cells, myeloid cells, and fibroblasts Arda is going after, and notes that depleting them in circulation is both less likely to help and more likely to be toxic. Arda pushed precision-cut tissue slice models past their published limitations, maintaining viability and tracking populations over time to bridge rodent, human ex vivo, and the clinic, and reports deep reduction of fibrosis markers in human IPF lungs with its lead T-cell engager.
Partial depletion is often enough, which widens the therapeutic index
For a pathogenic population that is not rapidly dividing, a fibroblast rather than a cancer cell, the dose-response is roughly linear: remove half the disease-driving cells and you see about half the benefit. Cancer denies you that because the cells regrow quickly, and B cells deny it for a different reason, since only about 0.01 percent are pathogenic and clearing that fraction means clearing all of them. Correctly identifying the causal population therefore buys room to tune a molecule for efficacy without toxicity. The related surprise from B-cell depletion is that when cells return, the pathogenic ones often do not return at the same rate, which turns a chronic condition into episodic treatment: deplete, withdraw, monitor biomarkers, retreat on the cycle of reaccumulation.
One queryable corpus, and no bioinformatics queue
A standard biotech stack assumes small, bespoke, hypothesis-specific datasets, and Arda works the other way: thousands of donors and tens of millions of cells across internal and external studies, different technologies, different tissues, plus spatial, genetic, and clinical metadata. Off-the-shelf integration either leaves too much study-to-study variance behind or erases the biology by forcing everything to look alike, so Arda built custom ML integration, including a whole-organism specificity atlas holding roughly 15,000 donors’ worth of data that lets it compare a target’s expression against some 600 healthy cell types across 50 tissues. Every data object lives in one structured system with FAIR-compliant metadata and a controlled vocabulary, which is what lets AI tools query the corpus with no additional data engineering and lets any scientist pull a gene from a dataset without filing a request with the bioinformatics team.
Related: CorrDyn builds the data engineering and machine learning systems that biotech and life sciences organizations depend on, including the statistical discipline that separates a real disease signal from an artifact of how the analysis was parameterized. Also on single-cell and spatial data: Kenny Workman of LatchBio on spatial biology and Mo Jain of Sapient on genomic analysis.
Full Transcript
Jason: Welcome to Data in Biotech, brought to you by CorrDyn. This is the podcast where we explore the intersection of data science and drug discovery, through the eyes of the people who are living it every day.
Ross Katz: Welcome back to Data in Biotech. I’m Ross Katz. Most drugs work by grabbing one protein and holding on. You take the pill, the protein stays blocked, and the day you stop, the disease comes back. My guest thinks that whole approach is aimed at the wrong target. Adam Freund is the founder and CEO of Arda Therapeutics. His bet is that the real unit of disease isn’t a protein or a pathway. It’s a cell. Certain cells turn pathogenic and drive the disease, and the more logically sound approach is to find those cells and remove them, the way oncology has removed cancer cells for years. What makes that possible now is the availability of new types of data. Single-cell sequencing can turn a piece of tissue into millions of measurements, enough to point at the exact cells that are doing the damage. We get into three things on the episode. How Arda finds those cells, why the standard way of analyzing the data quietly misleads a lot of careful people, and what it means to treat a disease in a handful of doses instead of every day for life. Adam Freund, welcome to Data in Biotech.
Adam Freund: Thanks very much, Ross. Glad to be here.
Ross Katz: Awesome. Just to kick us off, so you’ve spent over 20 years in R&D, including helping scale Calico from 15 to over 220 people, and you’ve built Arda around the thesis that cells, not pathways, are the functional units of disease. Can you walk us through your background and the intellectual journey that led you to where you are today?
Adam Freund: Happy to. So my intellectual passion has always been the biology of aging and the diseases that come with it. And so when I was a grad student at Berkeley and then my first industry role at Calico, I spent both of those roles studying for years, studying the cells that go wrong as tissue deteriorates. So cells that accumulate and drive dysfunction in these various tissues. And so the obvious therapeutic move if the cell is the problem, is to remove it. And that’s exactly what we do in oncology. It’s the standard approach therapeutically there. But outside of cancer, almost nobody was doing it and that’s because we lacked two important things. We didn’t really have a map of which cells were pathogenic in a given disease, and we didn’t have the tools to eliminate specific cell subsets or cell states. And so what changed, and what I think allowed us to start Arda, is resolution and the ability to actually create those maps and then the tools to enable that. So in the last decade or so, single-cell technologies have reached the scale and the cost dropped enough so that we can turn a lump of tissue into millions of individual data points and thereby read the full cellular architecture of the disease. And that’s not just from a single sample from a single donor, which is of course where single-cell technology started, where all technology starts, is with Ns of one. But now we have data sets of hundreds or thousands of donors, both at Arda and externally. And so we can get that true sampling of human variability. And so that gives us two things at the same time. The first is the identity of the pathogenic populations and secondarily, the surface markers that actually help us distinguish them, which we can get into in more detail. But those, those surface markers are basically the coordinates by which we actually target and remove the cells. And so once that technical, technological revolution happened, once that data existed at scale, building a company around cell depletion went from this idea that had been in my head for a very long time, to something that could actually be done practically.
Ross Katz: Awesome. So if you look at the by comparison, the other companies in the in the biotech space the way that we talk about how to treat disease is by blocking individual individual factors. But there’s limitations to this approach where we’re blocking individual factors versus depleting an entire category of cells. Can you talk a little bit about how you think about that problem and why this cell depletion approach is the right is the right therapeutic approach for certain disease types?
Adam Freund: This was the whole, this was the whole foundation for the therapeutic hypothesis. A cell’s behavior, cells first of all are the functional units of life. They’re the functional units of tissues, and thereby they’re also the functional units of disease. Diseases happen because cells go wrong. And yet when we think about going in and treating disease, we’re often targeting the activity and function of single proteins and single pathways because that’s what we know how to do. But a cell’s behavior is governed by a signaling network, and we’ve evolved checks and balances that make that network deliberately redundant so that cells are resistant to perturbation. Buffered stability is an important part of life. It’s the foundation of homeostasis. The unfortunate corollary is that redundancy is exactly what makes highly effective single-target drugs hard to find. If you block one node, the network routes around it. And so consequently, much of chronic disease pharmacology delivers these modest incremental benefits, and not to mention has to be dosed continuously, so you’re fighting the network’s own robustness, and you’re fighting it forever. But if the cell itself is pathogenic, if the behavior of the cell has taken on this aberrant set of characteristics that you’d rather avoid, the higher efficacy move is to just take out the whole node. Rather than argue about which of its 10 redundant pathways to inhibit or try combinations of those, etc. Etc. If you remove the cell, you shut down all of its disease-driving pathways at once.
Ross Katz: Yeah, I think that makes sense. And I think it gets to some of the some of the challenges that we’ll talk about that we’ll talk about in a bit. But my understanding is that the that sort of the validation case for this cell depletion method was in was in B cells with Rituximab and its successors, which have generated a bunch of revenue and have been able to treat a several different autoimmune diseases. So we’ve got this we’ve got this success case that says cell depletion can work. But I guess what interests me is how are you building on top of top of that, and are there any are there any sort of general limitations to the cell depletion approach that sort of guides where you can take this in the therapeutic ecosystem?
Adam Freund: That’s, that’s a great question. Originally, I think, of course, the proof case was cancer. But cancer’s unique in that the cells are we understand the bad cells. They’re the ones that are growing uncontrollably. But B-cell depletion represents one of the first situations where we weren’t totally clear what the pathogenic population was and we had to identify it and basically prove to ourselves that it was causal by depleting it in the clinic and seeing the benefit. It’s worth reminding everybody that George Schett’s first CD19 CAR-T paper which was a lupus case study, was only published in 2021. So we’re only like five years into the second phase of the B-cell depletion revolution. Obviously, Rituximab came first, but it was not effective in a lot of the diseases where we now are seeing efficacy with these more potent depleters. And so it took time for people to believe these newer modalities like CAR-Ts and T-cell engagers could work because Rituximab had repeatedly failed in lupus and other conditions, though it did work in some autoimmune diseases. And so there was this genuine causality question. And I remember debating this in 2021 and 2022 when these case studies first started coming out, like which diseases are actually B-cell driven? And it now seems almost obvious that a whole set of them, lupus, systemic sclerosis, MS, RA, myasthenia gravis, can be treated with these drugs, but that clarity is recent. And so it’s not surprising at all to me that we’re only now thinking about how we can extend the concept of cell depletion beyond B cells. Once you start to think about targeted cell depletion as a therapeutic strategy, though, you start to see evidence that extends beyond B cells. Most of it was not intentional per se as a as a targeted cell depletion strategy. It just happened that’s what the drugs do. So mostly these are crude, broad-spectrum depleters. So think about like anti-thymocyte globulin or CD52 or CD3 antibodies that wipe out all T cells and prevent transplant rejection, or CSF1R antibodies deplete macrophages, almost all macrophages, to treat GVHD, or like c-Kit antibodies for mast cells and mastocytosis, or IL-5 receptor antibodies for eosinophils in eosinophilic asthma. Like, there are actually a lot of examples mostly centered around these different subsets of immune cells. But most of those came from discovering that a cell type depended on continuous signaling through some pathway for survival and then blocking that pathway, which resulted in death of the cell. That works, but it’s a very constrained way to build depleters because it relies on this deep pathway knowledge and it’s hard to get specificity of subtypes of those cells. You end up depleting entire cell populations. Which is true of B-cell depleters, too, by the way. That’s how CD19, CD20, BCMA, these don’t target the pathogenic subset. They target all the cells and we just hope we get the bad ones along with the good ones. What we want to do at Arda is to make this modular and repeatable. And so for that, we take our cue from oncology. In oncology, we kill the cells by finding the unique surface markers, even when they play no functional role in the cell’s behavior. That’s true largely of CD19, BCMA, which were originally cancer targets, of course CD22, CD33, DLL3, etc. Blocking these things doesn’t kill the cancer cell. They’re just addresses. And so we can take that same approach to pathogenic cell states in immunology and inflammation and we can co-opt the modalities that our oncology colleagues have built, like ADCs, CAR-Ts, and T-cell engagers. We just need to know the right cells to target and their unique surface markers. And that’s that’s what we’ve built a discovery engine to do. Your additional question was about the limitations of this approach. And I think we won’t know the true limitations until we test clinically all these different cell states. But I think the fundamental the difficulty outside of oncology is, as I alluded to earlier, finding out which cells are causally pathogenic. We don’t always know. We have good hypotheses, but we can and we can test that preclinically. But ultimately, those tests have to be done clinically. The other difficulty is understanding what modalities are the right modalities to deplete different cell types. If you’re targeting a solid tissue fibroblast, that is a fundamentally different problem than targeting a circulating B cell. And you have to think about designing your therapeutic modalities and selecting your targets in a way that is aware of that difference.
Ross Katz: Yeah, I think that’s a great jumping-off point to move into the discovery engine and the way that you and the way that you go in the direction of selecting which cells to target. So before before single-cell sequencing, you had to as I understand it pre-specify a hypothesis that you were looking for. So you only so you could reject the reject the null hypothesis, or you could confirm or you could have enough information to continue with your experimentation. And so there was this idea But I’m interested in what changes you make inside of your discovery engine that allows you to that allows you to work from a hypothesis-free discovery engine or hypothesis-light discovery engine.
Adam Freund: It’s analogous to going from a qPCR approach to understanding gene expression to a whole transcriptome sequencing approach to understanding gene expression. In the first, you have to pre-decide what probes you want to look at, and there’s some practical limit to how many you can look at. Whereas in the second case, when you’re actually sequencing the whole transcriptome, you don’t have to have any pre-existing hypotheses. You just get to get to know everything that’s there, and you get to then ask what’s changed. The same thing is true now when we think about the cellular architecture of tissues and of diseases. So before single-cell sequencing let’s think about what our options were. We really had histology where you pick the antibodies that you want to stain, immunohistochemistry and things like that, where you pick the antibodies you want to stain against and then see where they are in the tissue. And then a little bit later, we had flow cytometry where you do the same thing, but you can generally stack more antibodies and compensate for things and so you can get more populations of cells. But you’re still limited to a handful, really, and you have to have these validated antibodies, and anybody who’s worked with antibodies knows validating them is definitely challenging. So I think one actually one really good example of this in the specific case of fibrosis, which is an area that we focus on, is Alpha SMA. So Alpha SMA was treated as the marker for pathogenic fibroblasts, for myofibroblasts, for decades. But it’s only when you look at the single-cell data, and anybody can verify this in a in a single-cell data set, that you see that Alpha SMA or ACTA2, which is the gene, expression is actually anti-correlated with fibrosis gene expression in disease tissues. If you look within the fibroblast population, the Alpha SMA high and the collagen fibrosis high populations are different. Single-cell data shows us this, but we never really appreciated this just using Alpha SMA as a general marker of fibroblasts of activated fibroblasts histologically. And so, right, unbiased single-cell sequencing flips this whole thing. You can sequence every cell irrespective of its identity. Obviously, there are sampling biases in any technology, so it’s not purely unbiased, but you don’t have to go in with a pre-specified hypothesis. You can actually then ask once you have all the data, which cell populations shift in disease, which ones expand, which ones contract. You can read the actual language of the cells. And then taking it one step further, spatial transcriptomics adds the next layer. Now we get back that almost anatomical information about where those cells are in an intact tissue, but we also get, once that technology reaches true single-cell resolution and we’re basically there which cells are present. The full accounting of the cells and where each and every one of them are within the tissue, and that’s that’s an incredibly powerful type of data.
Ross Katz: And so if once you have this single-cell data, and I think you I think you outlined really nicely what the availability of this data has given you from the perspective of basically trying to understand the biases of the data but also use it as a data exploration to create hypotheses as you’re exploring through it. But when most groups are using single-cell transcriptomic data, they’re they’re starting by clustering, and one of the things you’ve talked about is the limits to the typical clustering approaches that people will use when they’re when they’re digging into this data. Can you talk about what the what the typical approaches are and what the limitations to those approaches are and how Arda thinks about things differently?
Adam Freund: This definitely gets a little bit technical but ends up being quite important. And so clustering, when you have single-cell data, what you have is the transcriptome of a bunch of individual cells. And one of the things that we try to do with that data is understand with all of these individual cell data points, rows in a gigantic spreadsheet where the columns are genes, which of these cells are similar to each other because it’s really hard to think about a million individual cells. And so we say, well, let’s think about what the cell states are. There’s not a million individual cell states, there’s some smaller number of that, probably in the dozens, and clustering is a way to group cells by their phenotypic similarity across their transcriptome. And the problem, this is a very sensible approach and it’s a smart way to think about biology, but the problem is that the number of clusters and how you divide all these cells into these clusters and these cell states isn’t handed to you by biology because cell state isn’t a biological ground truth. It’s an abstraction that helps us conceptualize what’s going on in the tissue. And so consequently, the number of cell states in a tissue doesn’t have a true answer. It ends up being driven by a metaparameter that the analysis that the analyst sets. And if you turn the knob one way, you can get two clusters. If you turn the knob the other way, you can get 2,000. Like, I’ve done this with my own hands. It’s very easy. And so that’s the first part of the issue because the biology that you are interested in doesn’t necessarily cut at the same level or in the same way that the phenotypic clustering does. The second, and I think it’s actually it gets a little bit worse. So the next step when people do this and they’re thinking about what cells are enriched in disease is they do a chi-squared test in each cluster where you’re basically asking how many cells in this cluster come from disease patients and how many cells in this cluster come from healthy patients. And if you have a lot more cells from disease patients than healthy patients, you say that population is enriched in disease. So the problem with that is that a chi-squared test derives its statistical power from the number of cells in the cluster, not from the number of donors that contribute to that cluster. And so a single patient can contribute tens of thousands of cells, and you could have a data set with one patient and one control. And you would get, if you hold the number of cells constant, the same p-values as a data set with hundreds of donors. And you don’t need me to tell you, nobody needs me to tell you that doesn’t make any sense. We all intuitively grasp that the real unit of biological variability is the donor, and all other transcriptomic analyses are already powered by donor number: microarray, bulk RNA-seq, qPCR. It’s only single-cell data set it’s only single-cell data that enables this new goalpost because the data now contain this other dimension, the cellular unit, and it’s so easy to pretend that the cell is the unit of variability that we care about. And if we’re trying to understand basic biology of cells, yes, that is what we care about, but if we’re trying to understand the difference between disease and health, that is not what we care about. And I think the tell here is always very simple: if your significance doesn’t scale with the number of donors, then you’re not measuring disease biology.
Ross Katz: Yeah, I think that makes a lot of sense, and you’re right about the intuitive nature of it. So you’ve got this you’ve got this single-cell data set where typical clustering approaches require you to require you to select the number of clusters in advance, but then also there’s there’s a data science hack that’s often out there where people will just iterate through the numbers of clusters until until the statistics show that what they want what they want the clusters to show, and but then you’re not only making inferences that are wrong based on the statistical power that utilizes cells instead of donors, you’re also making inferences that unless you control for it are not are not accounting for the fact that you’re that you’re testing a variety of hypotheses back to back to back to back. So you’ve got this go ahead please.
Adam Freund: I was going to say one of the one of the jokes that we make internally is like the one of the best ways to get a nature paper is to find a new cell type. And so if you find the way to find a new cell type is to just increase the number of clusters until you find a new cell type. And so yeah, it’s very there really is and I’ll just repeat it, like there is no ground truth here. And it’s really about what is the question we’re trying to ask, and from a therapeutic standpoint, what is the specificity we can even achieve therapeutically. And clustering cell states beyond that level of therapeutic resolution isn’t really helpful. So yeah, it’s just
Ross Katz: No, that makes a lot of sense, and I think that so what I’m interested in from the perspective of Arda’s method is how do you marry these two things together? You’ve got the you’ve got the resolution that you get from single-cell transcriptomic data and spatial transcriptomic data, and then you’ve also got this idea of the donor as really the as really the individual source of statistical power. So can you like walk us through what is what is your approach look like?
Adam Freund: Yeah. So if you if when we first developed these approaches as part of the platform, we spent actually quite a bit of time thinking about what the optimal output of some algorithm that assessed cell expansion in disease would look like. And what it would look like is on a UMAP or other low-dimensional representation of single-cell data, it would be per cell a measurement of how disease enriched that area of the UMAP is, all. And it would cross clusters and cell types. It wouldn’t be constrained by these boundaries. It would it could be tiny, it could be huge, and it really would be this sort of independent axis around which the identity of the cell is not driving is not so important. And so what we ended up using was this approach called neighborhood analysis. So instead of clustering, we look cell by cell and we give each cell a score based on its neighborhood. We look at its nearest neighbors in expression space and we ask which donors those neighbors come from. Are they disproportionately disease donors or controls? It’s guilt by association. A cell sitting in a heavily disease-enriched neighborhood gets a high disease-enrichment score, and that signal propagates through expanding neighborhoods. We don’t just look at the nearest neighbors, we look at the nearest nearest neighbors and the third degree and the fourth degree. There’s a very iterative and algorithmic way that you can do this. We correct for donor representation so that a single oversampled patient can’t dominate. So if you can think about a picture where you have one cell on the center and five cells around it and they’re all from disease. If you have five donors in your data set and each five disease donors in your data set and each donor has contributed one cell to the surrounding space here, then it’s good representation across all of your disease donors. But if you have five donors and all the surrounding cells are just one of those donors, that is less representative of the overall disease sampling that you’ve done, and so those should give you different statistical results, and indeed they do. The output of this is a continuous map of disease enrichment across the entire data set, completely independent of any cluster annotation. And so that gives us a view of how disease reshapes the cellular architecture of a tissue in a way that is as unbiased as we can be.
Ross Katz: And I and I mean it strikes me that as you think about that sort of disease-enriched cell population looks like is step one and understanding what how enrichment might be representative of causality is like the longer chain of inference that you’re trying to go on. So can you can you walk us through what that what the methods are between or how you think about moving from enrichment to causality?
Adam Freund: No, exactly. Correlation is not causation. So enrichment is just a starting point. It’s not proof. Causality ultimately requires experimental validation ultimately in the clinic. Of course, we build there through preclinical testing. But the job of Arda’s platform is not to prove causality computationally, but rather to shrink the candidate list to something that’s experimentally manageable, usually one to three potentially causal cell populations that we call our suspects, our suspect populations. Disease enrichment is one input and it’s a good first filter because the causal cells are almost certainly within the set of expanded cell populations. The same way that the protein you want to block is almost certainly within the set of upregulated proteins in the disease. But because correlation isn’t causation, we have to layer on more beyond that. And so we also look at what which cells produce factors that are already known to drive disease, for example, cytokines in inflammatory disorders or extracellular matrix factors in fibrosis. We look for cell populations that both express these and are enriched. We look at cells that express genetic loci that are associated with the disease in human populations through GWAS, so bringing in those those giant human population studies and layering them on top of the single-cell expression data to figure out what are the cell states that are driving that genetic association. How much weight we put on the human genetics, for example, depends on how strong the overall heritability signal is for that particular disease. For many inflammatory diseases, it’s quite strong, for some fibrotic diseases, it’s weaker, and we downweight it accordingly. But it’s the convergence across these axes, and I think there’s about six of them that we really focus on rather than any single one that moves a population from enriched to credibly causal and worth experimental resourcing to test.
Ross Katz: That’s interesting, and it strikes me that as you think about this sort of cell depletion approach as like the as like a platform from which to attack multiple therapeutic targets the understanding both the genetic loci and the factors that the cells are producing and like thin thinking through what is known about the about the disease in question and how to weight those different factors in which cell group which cell groups you want to pursue is part of the part of the art of the platform that you’re developing. Am I thinking about that right, or how would you correct that viewpoint?
Adam Freund: Absolutely. It would be nice to have a push-button platform where you just hit run and it spits out the causal cell populations, but really it’s a combination of computation and biology understanding of that given disease. And so how we weight those different axes absolutely depends on what we know about the disease, and we in no way should discount years of understanding just because we now have fancy computational methods. We just want to add on to them, stand on the shoulders of giants.
Ross Katz: Yeah. So I think I think now’s a good time to talk about a specific target. So you’ve got Arda 101 which is as I understand it a pathogenic fibroblast population with this reproducible signature and so and you’ve validated the target and you’ve built an optimized T-cell engager. Can you walk us through that entire arc like how the how the workload that you’ve developed has enabled you has enabled you to get to that point?
Adam Freund: Our lead program is a nice end-to-end example. So we start as we do for most of our programs with human single-cell data across many independent data sets and hundreds of donors. We identified we identified a fibroblast subpopulation that has reproducib a reproducible pathogenicity signature in pulmonary fibrosis. And we confirmed that there is a target that is on the surface of those cells that is enriched, it’s highly specific. We have whole organism specificity analyses that compare the expression of our target on our pathogenic population to all other healthy healthy cell types that we have data for in the human body, and there’s something like 600 across 50 different tissues. And so then we once we have a target that we like from from that computational analysis, we go on to confirm the target is actually on those cells in intact tissue, that it’s actually elevated in disease at the protein level. And then we build and screen depletion molecules against it, and in this case, we landed on a T-cell engager. And so this is where the other part, another half, I would say, of Arda’s platform really comes in. A target is only useful if you can act on it fast. It’s easy to derive lists of targets. It’s hard to build conviction in targets, and that requires experimental testing. And so we’ve built this modular depletion library so we’re not starting molecule engineering from scratch every time. We have dozens of benchmark depleters that span modalities including what we call internally, for example, our carpet bombers, which are these broad-spectrum depleters that we use to pressure test the causal role of an entire cell type before we commit to a precision molecule. So it’s if depleting all of the macrophages doesn’t do anything, depleting a subset of macrophages may not be the best approach. Obviously, there’s caveats to that, but it’s generally a good way to prioritize broad cell types. We also have a well-characterized panel of immune cell engaging binders, so for example, CD3 binders for T-cell engagers that we can pair with our different binding arms so that we get that right ratio of target binding to immune cell effector binding. We have validated ADC payloads and conjugation chemistry, we have in-vivo ready sorry, in-vivo CAR-T ready formats. And the whole point of all of this is to be modality agnostic because the biology needs to pick the format. An internalizing target might call for an ADC, a setting that needs potent tissue killing for if you’re in a disease where you need to get deep into a tissue and there’s a bunch of T cells present, that might enable a T-cell engager better than an ADC and so on. And so for our lead program, Arda 101, the biology pointed us to a T-cell engager, but that same target we assessed in other formats as well. And it’s what that format switching that makes the platform repeatable rather than a series of one-offs. I would say that what we’ve learned is that the hardest step actually isn’t the target nomination, it’s everything after. It’s narrowing from usually when we have a cell type that we are interested in, we can pretty readily nominate a dozen or so surface markers that look pretty good by our computational analysis. But narrowing from 10 or 12 to the one that you want to bet millions on, that’s the hard part. Pairing it with the right modality, deciding if you need combinatorial AND gate logic understanding what modality fits the cell type or the target tissue and then building a molecule that can deplete those cells in real human tissue rather than in just a dish, that all takes time and that takes work. And so we’ve designed the platform to make all of that as robust and straightforward and fast as we possibly can.
Ross Katz: That’s very interesting, and it and it also makes sense that so all the all the talk in the data space is on is on discovery, but that there’s there’s the specific disease and cell and binder and all of the all of the detailed aspects that you need to figure out after you’ve nominated you’ve nominated the candidate cell population to go after. The another differentiator that I think we talked about previously was this idea of solid tissue testing because a lot of cell depletion data data is measured in blood or in the circulatory system. So it’s why is measuring measuring from solid tissue more important, or why is measuring from the circulatory system misleading in a disease like IPF?
Adam Freund: So this is a lesson that we’ve largely learned from the B-cell space. And so zooming out just a touch, there’s two ways that people assess whether their depleters have worked. In oncology, it’s generally xenograft models. So you put a ball of cells under the skin of a immunodeficient mouse, and then you inject your depleting modality and you look for that tumor to shrink. That is obviously a highly artificial system, and we understand that. But when we moved as a field to B-cell depletion, one of the things that we started to do was look at B cells in circulation because the targets that we’ve discussed previously, CD19, CD20, they don’t just deplete the bad B cells, they deplete all B cells. And so where can you readily measure B cells? You can measure them in circulation. And these things look great. They completely eliminate B cells in circulation, undetectable levels down to effectively 0%. Problem is, so does Rituximab. So Rituximab depletes 100% of B cells in peripheral blood, so do CD19 CAR-Ts. But Rituximab doesn’t fix lupus and CAR-Ts do. Why? Right? Because the B cells in circulation are a red herring. They’re not the causal cell population. You have to deplete the B cells in the lymphoid organs, in the marrow, or in the affected tissue itself. And if that’s true for B cells, that’s almost certainly true for these other cell types that Arda is now going after: T cells, mast cells, myeloid cells, certainly fibroblasts. It’s much less likely that you’re going to derive a benefit in, for example, IBD or mastocytosis by depleting the mast cells or the T cells in circulation, it’s also much more likely to be toxic. If you could get those cells in the affected tissue and only the cells that are in the affected tissue, that is likely to be both more efficacious and better tolerated. And so we have developed a system to look in tissues. And ideally, this is human tissues, though we can do it with preclinical model rodent tissues as well. And so we’re not the first to do this. Precision-cut tissue slice models have been around for a number of years, but we’ve pushed these models well past their published limitations, and we’ve optimized them for the mechanisms and modalities that we care about. So we figured out how to maintain viability, spiking in appropriate controls, tracking the different populations in real time so we can build these time courses bridging the human versus rodent ex-vivo to in-vivo systems so that we can build that chain from a mouse model to an ex-vivo system to a human ex-vivo system and then hopefully to the clinic. And what we’ve seen is what I mentioned: deep reduction of fibrosis markers in human IPF lungs with our lead and we have the same approach that we’re now taking into other fibrotic programs. And it really gives us confidence that this will work in the clinic because we’re seeing it in real human tissue with real human target cells and real human effector cells.
Ross Katz: It’s really interesting and it and it also so there’s the hearing you talk about the computational approaches you’re using, I think I think another aspect of your computational approaches that you’ve that you’ve talked about previously is you’re able to computationally model what happens when you deplete target cell populations in your spatial transcriptomic data, so you can simulate you can sliding scale what percentage of the target sales you’re target cells you’re able to remove and ask how the tissue signaling changes over time. Can you outline what that modeling looks like and maybe a little bit about the assumptions or limitations of those kinds of models?
Adam Freund: It ties back to the previous question of a solid tissue testing platform. Those are really useful models, but they do take real experimental work, and often they take fresh human tissue, which is not always an easy supply. So at Arda, we think a lot about how we use the best parts of both computation and AI and the best parts of experimentation. And so you can run a simulated version of that human tissue depletion experiment if you have, for example, spatial transcriptomic data. So let’s say we have a fair amount of single-cell resolution spatial transcriptomic data, which we do internally. If we and so we can annotate every cell, we can determine what they express, and we can easily identify the cells that our therapeutic is likely to target. And then we can ask, if we remove those cells or some fraction of those cells, what happens to the tissue and how does that evolve over time? It involves modeling the response of the other cell populations, how the overall signaling networks rewire what happens to the effector populations that expand, and so you can simulate this sort of depletion at increasing depths the effector cell recruitment, the recompute the tissue composition, and then score it on a disease axis that’s anchored between a normal tissue and a disease tissue. So a lower score would mean the tissue is predicted to move towards a normal-like state. This is we run longer horizon homeostasis models as well, letting the cells regrow, letting the effector cells decay to understand whether the effect persists rather than snapping back. Look, the reality is all of this is based on a ton of assumptions about cell kinetics that we parameterize, and so it’s hypothesis-generating, it’s not ground truth, but it is useful for prioritization. It helps us reason about what targets and what depth of depletion are likely to actually move the needle before we commit to that wet lab development. And I think one of the important things that we’ve realized is that when you are if you are targeting a cell population that is truly pathogenic in a disease but is not rapidly dividing like a cancer say a fibroblast population, for example. While of course we would love to remove 100%, that is not actually required to see a beneficial effect on overall cellular architecture and tissue response. It’s a much more linear relationship where if you deplete half of the disease-driving cells, you see half of the potential benefit. In cancer that’s not the case because the cells regrow so fast, and in B cells it’s not the place it’s not the case because only 0.01% of the B cells are actually pathogenic, and so you have to get rid of 100% of them in order to hit that 0.01%. But if you’ve correctly identified the causal cell population, really you have a little bit more room in your therapeutic index to tune your molecules to achieve efficacy without toxicity, and that ends up being quite helpful.
Ross Katz: Right. And it makes sense that from the perspective of tuning or detuning, you’re you want to understand what the influences are of like which cell which cell groups you kill and then like what are the or you deplete and then like what which and what percentage you want to leave remaining. I want to zoom out a little bit. So from the perspective of cell depletion as a therapeutic modality one of the things that you’ve that you’ve mentioned is being a key advantage of cell depletion over the alternatives is that you have the opportunity to do this dose-and-wait treatment cycle, rather than having to continuously continuously dose the patient with the therapy. So I’m interested in what that means for patients and then what it means for Arda from a clinical development perspective.
Adam Freund: This is if you’re blocking a pathway or you’re blocking a protein with a molecule, generally that molecule has to be on board all the time. If the molecule goes away if you stop taking the drug, the activity of the protein comes back right away, and so then the disease symptoms come back as well. But one of the consequences, one of the obvious consequences of a depletion mechanism is that the disease only comes back when the cells come back, and the rate at which the cells come back very much depends on disease biology, cell biology, and the ability of the tissue to reset. And so this is something that we’ve seen for B-cell depletion. I don’t think I actually I don’t think anybody expected this, I certainly didn’t expect it, but it turns out that if you deplete all of the B cells and then six months later the B cells come back, often the pathogenic B cells don’t come back at the same rate. So you get healthy B cells, but the disease stays in remission because the bad B cells don’t come back. Now they may come back in the future, there was some initial trigger that caused them to form in the first place, that trigger may still live within that person, but this idea of an immunological reset, which is now the phrase du jour, which is it’s a great it’s great branding, like it’s really great for patients. They take a drug and it looks like, at least for a little while, a cure. And we are very careful about using that word, but it’s when you’re when you’re not on a drug and you don’t have any symptoms, that’s what it starts to look like. Now it may not be a permanent cure but it’s still quite powerful. It’s good for safety because the drug can’t be having side effects if it’s not on board. It’s good for patient convenience because they don’t have to be taking something every day or every week. And so we think there’s a real potential for this kind of intermittent dosing when you target other cell populations as well. And so you would have this episodic treatment. You knock the population down, you take the drug away, you monitor, and you retreat only if the cells reaccumulate, and you treat on the time cycle of that reaccumulation, which you can hopefully measure by biomarkers. And yeah, that’s the way that you assess this of course clinically is you look for persistence and durable pharmacodynamic effect after the drug is cleared, and that’s that’s something that we plan to do.
Ross Katz: It’s really interesting and it also so there’s the reemergence of the cell population but then there’s the working assumption that if the cell population were to reemerge that the that if you were able to successfully deplete that cell population previously that the same treatment should be able to then extend like the repeat the there it shouldn’t attenuate over time. Am I thinking about that right, that like the repeated the repeated administration over the course of years in between should be able to accomplish the same goal?
Adam Freund: Yeah. And this is some a question that we get a lot. It’s actually a very astute question. People ask, well, cancer evolves resistance to cell clearing mechanisms over time, so won’t the same thing happen in your case? And I think the thought is correct in that there’s the same selective pressure present. If you are depleting something, only the things that don’t get depleted will remain. And so over time, that’s what you would select for. The difference is that cancer cells are genomically unstable, and so they are much more able to evolve these clones or these sub categories that don’t express the protein or that have somehow evolved resistance to your thing. When you’re depleting at an immune population that has become hyperactivated or you’re depleting a fibroblast population that’s become hyperactivated, these are genomically stable. They are cells that are not doing what they should, but they are not malignant. Their DNA repair mechanisms are fully intact and they do not acquire mutations at a rate that is anywhere close. So yes, they will regrow, but the chance that they will somehow be genetically and unstably different than the cell population you previously eliminated is quite low.
Ross Katz: That makes sense and it’s also it’s I think we see this sort of across the biotech ecosystem that the weapons that have been developed to fight cancer are available and useful for the other the other diseases and therapeutic targets that we have across across the ecosystem. So if we take it let me just give one example go ahead please.
Adam Freund: I was just going to say it’s all great to talk in theory, but the practical example is once again B cells. When the B cells come back, they’re CD19 positive. They’re not CD19 negative. And that’s all the proof you need. That the concept is sound.
Ross Katz: Yeah. The approach that Arda’s taking is very unique and it’s relying on these emerging data gathering approaches in the form of in the form of thousands of donors, millions of cells, spatial transcriptomics genome-wide association studies integration, clinical metadata. I’m interested in what the data infrastructure ecosystem looks like at Arda to support the platform that you’re developing.
Adam Freund: This is especially in recent years. With AI, this has become all the more important, and a standard biotech stack assumes small, bespoke, hypothesis-specific data sets, but we’re doing as you said the opposite. We are taking thousands of donors, tens of millions of cells across studies, both internal and external, different technologies, different tissues spatial, genetic, clinical metadata on top. So the hard part of this is making the all of these different data types comparable, making it all speak the same language. And so we’ve had to build a lot of this internally because off-the-shelf integration methods leave too much study-to-study variance behind, or they end up eliminating the biology that you care about by virtue of trying to force everything to look the same. And so one example of this I mentioned previously is all our whole organism specificity atlas where we pull in data across platforms. I think we have like 15,000 donors’ worth of data in that atlas, and we have to get all those to talk to each other, so we built this custom ML integration that gives far better correlation across data sets while retaining the exact cell type different signals that we care about. And that’s what lets us compare a target’s expression on our pathogenic cells against every other human cell type with a fair degree of confidence. Underneath all of this is there the sort of maybe it’s boring, maybe it’s not to your listeners, but like the infrastructure, the data infrastructure. And so we have every data object within Arda lives inside this single structured system with this rich FAIR-compliant metadata and this controlled vocabulary so that the whole corpus is discoverable and queryable, rather than scattered across one-off files on people’s desktops. And so the payoff of that is that things like AI tools can interface with this corpus directly, no additional data engineering time is required. And so the other benefit to this is access. We’ve built tools that put most of our single-cell data objects at any scientist’s fingertips with easy intuitive interfaces, and these things again are internal that we’ve developed specifically for the purposes of what we’re trying to achieve. So we’re not bottlenecked with this whole query and response cycle with the bioinformatics team hey, will you look up this gene in this data set for me? Everybody at the company can do that themselves. And I think of this very much as a just asset that grows and compounds as we add more data. And so I would say the way we think about it is a lot of companies are built and are being built these days to generate huge amounts of data. We are very much generating, we’re we’re trying to build a system that allows us to generate drugs, not just lists of targets, not just data sets.
Ross Katz: Hmm. So just hearing you hearing you talk about it just made me coming from coming at this from the data side, I’m just interested in, should I think about this as a as a search system, as like a data warehouse type system, or as like or as like a data lake as a data lake system? Like how do you how do you describe the infrastructure that exists in order to enable what you just described of like FAIR FAIR data with metadata in between that allows both humans and AI to like query and get what they need without having to go through the bioinformatics team?
Adam Freund: It is lake, warehouse, I’m I’m not sure that I am able to fully distinguish the terminologies here. But what we do is there is a bit of a manual component to it when we add data sets to our warehouse, lake, whatever it is. We do annotate them with that shared vocabulary so that we understand where they come from, what they’re doing. And you can now automate this pretty well with AI because you pull from the original papers and you can used to take used to be a lot slower, I used to do it by hand. Now, I think generally we’ve automated it. But this then allows search or AI indexing to happen quickly. And so if you want to say, well, where are all my IPF-related data sets, even if somebody calls it progressive pulmonary fibrosis and somebody else calls it interstitial lung disease you can figure that out. And it’s a toy example, but that type of thing propagates across every single metadata parameter that we have. And that allows us to harmonize things much more rapidly.
Ross Katz: That makes sense, and it also like so a typical solution for that is like querying the you’re able to query the metadata to identify the file to identify the files that exist when you need to dig into like whatever the spatial transcriptomic data that was that was generated by a particular experiment or something along those lines. That makes sense. So as we head toward the end here single-cell and spatial are moving really fast with higher resolution, lower cost, richer multimodal integration. How would you how do you see the landscape evolving over the next few years and what do you think it means in terms of what Arda’s able to accomplish?
Adam Freund: The one of the biggest bets that we’re making is on spatial data. I mentioned this a little bit earlier, but it’s so much richer than dissociated single-cell data where because where a cell sits in the disease tissue is a really important component of its biology and it’s one of our key causality axes. And high-quality public spatial data is still very limited. So we’ve put a bunch of effort into internal workflows that generate this high-quality spatial data at throughput and at reasonable cost. We’ve actually gotten the cost down to almost equivalent on a per sample basis to regular single-cell sequencing. Now tissue structure information is much, much harder to merge across donors than is cell transcriptomic data. And if we can make strides in this direction, it will enable a much richer view of disease biology. I think this is a core area of future research and development that needs to happen, both within Arda and externally. The other frontier that I think a lot about is protein data. We can read nucleic acids far better than we can read amino acids. And so RNA dominates the single-cell landscape. But protein is what matters most for surface targeting. And we’re getting better at reading it. We have things like CITE-seq, for example, but it still requires that pre-selecting a panel of validated binders, so we’re back to the dark ages in some way of needing antibodies that bind the right things. And maybe AI-guided antibody design can change that by making it easy to produce libraries of binders against thousands of proteins and you integrate signal even though no one binder is perfectly validated. Or maybe mass spec approaches finally reach true single-cell resolution. I’m not super bullish on that, but you never know. The mass spec field has consistently surprised and impressed me, so we always hope. So five years out, I would expect higher resolution, lower cost, much richer multimodal integration to become routine. And that will enable the kind of cell-centric discovery we’re doing at Arda that is still hard today in many cases to become more systematic. And I hope will enable many more companies to take this cell-targeting approach because personally, I believe that this domain of medicine is just getting started and has the potential to be truly transformative across disease.
Ross Katz: Awesome. Before we let you go, any calls to action?
Adam Freund: Just those. Just those. I really I can’t stress enough how much the B-cell depletion B-cell depletion is like the most exciting thing that’s happened in the I&I space in years. And it’s just one of a dozen cell types that drive these diseases. And like, we’ve just barely touched the other ones. And Arda’s we’re going as fast and as hard as we can to identify causal populations in these other ones and build the molecules, but we’re only one small team and there’s so much room here for patient improvement and so I just hope more people get on board with this idea that blocking cells or depleting cells rather than pathways is a way to really reshape disease progression.
Ross Katz: Awesome. Adam, thank you so much for coming on. Really enjoyed the conversation.
Adam Freund: Thanks, Ross. Me too.
Ross Katz: So I have two takeaways for me from that conversation. The first is to go after the cell and not the pathway. A pathogenic cell usually has several redundant ways to cause harm. So blocking any single harmful pathway only yields a small effect. If you remove the cell, then all of the pathways go away with it. And because the cell is gone rather than blocked, the benefit can outlast the drug. If you deplete the bad cells, then when they grow back, the harmful ones often don’t return in the same way. So the disease can stay in remission with no one having to take a daily pill. Adam’s careful not to refer to that as a cure but it certainly starts to look like one if what Arda is doing is correct. My second takeaway is a warning for data people. When you analyze single-cell data, the software often gives you clusters, and the number of clusters is a dial that you set, not something that biology is necessarily telling you. So you can set the clusters really low and get only two, you can set the number of clusters really high, you get 2,000, and you’re analyzing the outputs to determine whether the number of clusters is the right number. Arda’s approach relies on having the cells implicate each other which is an interesting counter to the typical clustering approaches. But if your standard test for which clusters matter is drawing its confidence from the number of cells, then you might neglect that the thing that needs to vary from patient to patient in order to detect disease is the number of is the number of sample donors. So it’s great that cells are voting in a certain way, but you need to control for the number of sample donors as well. So it just illustrates for me the level of care that needs to be applied when understanding under the hood what algorithmic approach you’re using and then what assumptions it relies on and then how those assumptions might need to be supported or might be violated statistically in trying to draw the conclusions that you need to draw. So I really like Arda’s approach methodologically speaking to drug discovery, and I hope you enjoyed listening to it. I’m Ross Katz, this has been Data in Biotech, and I’ll see you next time.
Jason: And that’s it for this episode of Data in Biotech. If you enjoyed the episode, please subscribe, rate, or leave a review in your podcast platform of choice. See you next time.







