Skip to content
John Androsavich — Why Biotech Talks About AI But Won't Pay for the Data It Needs
Data in BiotechEpisode 75

Why Biotech Talks About AI But Won't Pay for the Data It Needs

Ginkgo Datapoints GM John Androsavich on why biotech underfunds the biological data its AI needs, and what $199 ADME testing changes.

62:56Full transcript below
JA

John Androsavich

General Manager at Ginkgo Datapoints

Overview

Ask a computational biologist what would most improve their models and the answer is almost always more data. Ask who is paying to generate it and the answers thin out fast. John Androsavich, General Manager of Ginkgo Datapoints, states the gap in numbers: three data-labeling vendors in tech pull three to four billion dollars a year labeling text that already exists on the internet, and Meta’s roughly fourteen-billion-dollar minority stake in Scale AI is a single check worth more than a full year of AI drug discovery venture funding and several times the size of the entire single-cell data market. Against that, biotech looks like it is carrying on as usual and flirting with AI.

John Androsavich runs Ginkgo Datapoints, the bio-AI data arm Ginkgo Bioworks launched in late 2024 to sell biological data generated for model training rather than as a byproduct of a drug program. He holds a PhD in chemical biology from the University of Michigan and spent fifteen years in RNA therapeutics, including a stint at Pfizer as Global Head, RNA Medicine Lead, and earlier roles at Translate Bio, RaNA Therapeutics, and Regulus Therapeutics. He has sat on both sides of the table, evaluating which technologies pharma should buy and now selling the raw material everyone says they want.

In this episode, host Ross Katz and John work through why the underinvestment persists and what closes it. They cover ADME-1 at $199 a compound, about a tenth of what the assay costs anywhere else, and what happens to a discovery program when teams stop triaging molecules before they have generated the negative data models need. They get into two cases where more data did not help, the Virtual Cell Pharmacology Initiative and Ginkgo’s case for DRUG-seq over single-cell transcriptomics, a federated antibody consortium built on data generated for the purpose instead of decades of inherited records, and an autonomous lab where GPT-5 wrote its own experimental protocols and set a new industry low for cost per titer.

Key Takeaways

Biotech’s data spend is a rounding error next to tech’s

Three vendors labeling text that already exists on the internet pull three to four billion dollars in annual revenue, and that figure is growing. Meta’s roughly fourteen-billion-dollar minority stake in Scale AI, taken on its own, exceeds a full year of AI drug discovery venture funding and runs several times the size of the whole single-cell data market. Androsavich is careful to name the exceptions, Chai Discovery’s $400 million raise and Isomorphic’s $2.1 billion, and equally careful to say they are exceptions. None of this counts the hundreds of billions in data center and compute capital expenditure sitting behind tech’s number.

A tenth of the price changes when teams test, which changes what data exists

ADME-1 runs at about $199 a compound, roughly a tenth of the going rate anywhere in the world, built on Ginkgo’s automation plus partnerships with Inductive Bio and Tangible Scientific. The interesting consequence is not the saving. Tier 1 ADME turns out to be defined by budget rather than science, one or two assays for some teams and three or four for others, so a lower price lets a team run the full panel earlier and across more molecules. Triaging early is what destroys the negative data that downstream models need, so buying more sooner produces a training set that would not otherwise have been generated.

Volume stopped paying off around one to ten percent of the data

A 2026 Microsoft Research paper found that single-cell foundation models trained on 20 million cells saturate their learning somewhere between 200,000 and 2 million cells, which means the remaining 90 to 99 percent contributes nothing. Ginkgo hit a version of the same wall internally: a 2,500-antibody training set outperformed a 250-antibody set, but by a margin nowhere near ten times. The explanation was diversity. The larger set was less varied than the smaller one, which reframes the design question from how many measurements to which measurements.

Partnering decisions run on narrative, and that suppresses data spend

Androsavich has done search and evaluation on the pharma buy-side and now sells from the biotech side, and his read is that biopharma partnering is a marketplace, with all the irrationality Richard Thaler documented in marketplaces generally. Strategy fit, the internal champion’s motives, C-suite endorsement, existing relationships, and geopolitics all weigh on a deal alongside model performance. Pedigree and investor names give a company a real edge over a rival with comparable data. Without an empirical way to compare models head to head, a developer talking to its board has no clear argument for spending more on data, so the field settles into partner first and improve the model later.

DRUG-seq trades whole-transcriptome coverage for signal that survives modeling

DRUG-seq runs one perturbation per well in 384-well plates, chemical or CRISPR, measuring roughly 10,000 genes per well at about $10 a well and under $5,000 for a full plate. Novartis invented it in 2018 and it stayed obscure while funders and institutions concentrated on single-cell methods. Androsavich’s objection to single-cell data for training is the sparse matrix: dropout correlates with expression level, so the moderately expressed genes most sensitive to perturbation are the ones most likely to read as zero. He calls single-cell high in calories and low in nutrition, while conceding DRUG-seq gives up the full-transcriptome coverage that bulk RNA-seq provides at $300 to $400 a sample. Early adopters combining Virtual Cell Pharmacology Initiative data with single-cell models report better performance from both, and the initiative screens outside compounds free and releases the results under an MIT license.

A consortium that generates data beats one that pools history

Federating decades of accumulated pharma records means training on a distribution you cannot inspect, which Androsavich likens to doing a puzzle in the dark. Assay parameters are opaque, and a single partner’s methods usually drifted over the years the records span. The Ginkgo and Apheris Antibody Developability Consortium instead builds a purpose-made dataset at the center of the network: about five founding members each contribute 2,000 antibody sequences, get the raw results back for their own 2,000, and train on all 10,000, every measurement run under one protocol. Ginkgo trains a central baseline model so a member whose own approach underperforms still has something usable, and a member with a better model keeps it.

GPT-5 designed the experiments, and the autonomous lab ran them

Ginkgo’s Nebula lab is built from reconfigurable automation carts that began development at Zymergen seven or eight years ago, running roughly a hundred instruments on one magnetic rail with software scheduling 80 to 100 protocols concurrently. In a collaboration with OpenAI, GPT-5 selected 36,000 to 40,000 reaction conditions to optimize a cell-free protein expression system, wrote its own electronic lab notebook entries explaining each choice, and proposed reagents the team had not considered and in some cases could not source. The run set the lowest reported cost per titer in the field. Humans still handle procurement and loading, and Androsavich’s own framing of what comes next is a bake-off between the model and a human expert like Stanford’s Michael Jewett given the same automation.

Related: CorrDyn helps biotech and life sciences organizations decide where data generation earns its cost, and builds the data engineering and machine learning systems that turn those measurements into models worth trusting. Also on federated approaches to biotech data: Robin Roehm of Apheris on federated learning and Aliza Apple on Eli Lilly’s TuneLab.

Full Transcript

Ross Katz: Welcome back to Data in Biotech. I’m Ross Katz, principal and data science lead at CorrDyn. Here’s a paradox sitting at the center of AI and biology. Ask almost any computational biologist what they need to build better models, and you get the same answer: more data. Yet, as an industry, biotech barely spends anything on generating data. My guest today puts it about as bluntly as it can be put: Everyone thinks AI is going to transform biotech, but no one is willing to pay for the biological data that AI needs. John Androsavich runs Ginkgo Datapoints, a company built on the bet that data, not models, is what’s holding the field back. He came up as an RNA scientist, spent years on the pharma side deciding which technologies to buy, and now sells the raw material everyone says they want and few are willing to find. We get into why that gap exists, what it’s costing the field, and what Ginkgo is doing about it. From ADME testing at a tenth of the going rate to an autonomous lab where GPT-5 designed its own experiments. Let’s get into it. John Androsavich, welcome to Data in Biotech.

John Androsavich: Hi Ross, great to be here.

Ross Katz: Well just to get us started, would you mind, so we’re talking about Ginkgo Datapoints today and so you launched in late 2024 and it’s been described as defining the category of being a bio-AI data provider that’s generating biological data specifically for AI model training rather than as a byproduct of a drug design program. I’m interested in from your perspective and from the perspective of Ginkgo Bioworks as an organization what led to the launch of this program? What were you seeing both internally and externally that made you think this was the right thing to be doing?

John Androsavich: It’s a great question. And I think Datapoints really was able to emerge at an important time in the AI story as it’s been applied to biology. And so this is 2024, just to set the time frame. AlphaFold went to win the Nobel Prize in October. This would be associated with Demis and John Jumper, David Baker. So a really big boost for the AI bio space. We started thinking about Datapoints probably the a few months before that the summer of 2024, and really, I joined Ginkgo in January that year and looking around and seeing the resources that were available here. We have 250,000 square feet of space in Seaport, Boston that’s chock-full of lab automation and advanced analytical equipment. It’s a really superb facility. Most of this has historically been applied to what we now call our solutions business. That has a slightly higher barrier to entry. You’re doing strain production, maybe you’re making antibody. There’s like a physical outcome that comes out of that. But we’re really thinking about what are ways that folks could take a smaller bite out of this kind of lab facility. And data number one seemed to be obvious. But also understanding that if AI were to become more performant in biology, it would need domain-specific data and that data would not only need to be scaled, but it would need to also maintain a very high quality control. This is only achievable through automation and not through pipetting. Now there are other ways to generate data, pooled experiments, these surrogate outcomes, etc. But what we wanted to do with Datapoints was enable the industry to buy the same datasets that they buy now at a lower scale, but now at a much higher volume. And that’s only enabled through the automation, all the great analytical instrumentation that’s available at Ginkgo today. Now what’s interesting, Ross, is that the market reception to this was a little soft in that there were early adopters that we were happy to serve. However, at that time there was, and this is only two years ago, it’s amazing how far the field has advanced, but at that time there were a number of companies that were really interested in working with us, they could see that they could benefit from us, and that was mutual. However, they were a little scared off by the AI. They didn’t see themselves as an AI company and so maybe that wasn’t for us and they just wanted to maybe be a little bit reluctant to jump in there. So we actually had to adapt our marketing a bit to soften that position. And I think that we’ve actually seen the, that now the tide turn and the AI space is much more ubiquitous. It’s becoming even more so. People aren’t turned off by it. And we serve a number of customers, some that want the large AI training data, as well as others that are using maybe AI models on their side, but they need validation, they need lab-in-the-loop and it’s lower scale. But this now offering has really turned on and grown with the adoption of AI within biology.

Ross Katz: The I’m really interested in looking under the hood of the work you’re, of the work you’re doing from an automation perspective. But before we do what are the types of organizations or the types of, or the types of programs that tend to benefit most from the data that you’re able to provide?

John Androsavich: Great question. There’s we work with everybody across the spectrum. That’s one of the things I love about this job. So we work with small biotech, large pharma companies, small tech-bio that are really more technically inclined, less biologically, and they’re usually the AI developers in the space. But then also you have a really strong entrant especially recently of large tech firms getting into the biology space too. So we help them across the board. We typically work in a few different areas. We work on those pillars that are really key right now for biotech AI training. That is in the protein space where we’re working on antibodies and antibody developability. This is really characterizing an antibody and does it have drug-like properties? Does it go from just being a biological molecule to potential therapeutic? We also do the same thing in the small molecule space. This is chemicals. Is it just a chemical or could it be a drug? And this is really measurement of those characteristics. And then we’re also looking at the cell level and understanding how are cells responding to different perturbations, whether that be chemical perturbations or genetic perturbations, and really trying to help train models for what are now called virtual cells.

Ross Katz: Interesting. And if when we were talking earlier, you talked about, you talked about this paradox in the ecosystem where everyone believes that AI is going to transform drug discovery but and everyone says that the major bottleneck to better AI is more data, yet the investment that people are willing to make in data generation in biology is still relatively limited. So would you mind just fleshing that out a little bit? What why does this paradox exist?

John Androsavich: It’s fascinating. Every time you see a survey and you ask computational biologists, AI researchers in the bio space, what do they need to really improve their models and get better? Almost uniformly you see more data. But let’s put it bluntly. And I think this is, I’m going to put it stated a certain way for the sake of illustration. Everyone thinks AI is going to transform biotech, but no one is willing to pay for the biological data that AI needs. That’s the paradox. And I’m going to repeat that last bit. No one is willing to pay for the biological data that AI needs. So that’s a bit of an exaggeration, obviously. In recent news Chai Discovery, a biotech modeling company, they just raised four hundred million dollars. I suspect a good portion of that will actually go to new data. Even bigger, Isomorphic a couple months ago raised two point one billion a month two point one billion, and that’s maybe some of that would go to data. Other actually would go to maybe clinical trials. They do want to develop drugs themselves, unlike Chai. So let’s say that there are important investments being made in this space and I’m extremely excited about those investments. But I think those are exceptional examples and it’s worth putting the entire industry into perspective. So let’s compare biotech to tech. Just to label training data that already exists on the internet, but then to label it for LLMs, three vendors alone pull in something like three to four billion dollars in annual revenue. And that’s growing fast. Meta paid about fourteen billion dollars for a minority stake in one of those, Scale AI. And so that one check actually represents more than a full year of all AI drug discovery venture-capital funding and several times the size of the entire single-cell data market. So what you get is even before you account for hundreds of billions of dollars of capital expenditures that tech is pouring into data centers and AI compute, Ross, by comparison biotech looks like it’s business as usual, happy to just flirt with AI.

Ross Katz: Yeah. No, that’s it’s it makes a lot of sense and it’s also counterintuitive because a lot of the conversation in biotech seems to be dominated by the idea that these computational tools, these computational tools are going to lead to less expensive drugs coming to market. Or that it’s not going to be ten billion dollars to bring a, bring a drug to market, it’s going to be something closer to, closer to one, and it’s not going to be a decade to bring a drug to market, it’s going to be, it’s going to be a much shorter period of time. So like the, can we talk a little bit about the reasons for this underinvestment? Like one of the reasons that you mentioned previously was this idea that price, that of price, that biological data is more expensive to generate than scraping the internet. But at Datapoints you have this offering of ADME-1 your product that offers one hundred and ninety-nine dollars per compound. So I’m interested in how you think the automation that you’re able to bring to bear is able to change the cost structure and what that means in terms of the level of investment that companies can make and what they’re able to see from their investments in data in biotech.

John Androsavich: Thanks for mentioning ADME-1. We’re really proud of that product. It’s about, just to put it in context, about one tenth the price of where you can get those services anywhere in the world. Anywhere in the world. And there’s some really extra features that are special that go along with that and that’s not even just counting the price alone. That was enabled through automation but also through smart partnerships. So I’ll give a call out to Inductive Bio and Tangible Scientific that are our partners in ADME-1. And if there’s interest in that certainly we’ve seen a really strong demand for that because it is such a stellar price point. And overall though, this is classic Jevons Paradox territory, okay? So the more efficient a resource becomes, the more money is spent on it because consumption goes up. And it’s not just that biological data is expensive, it’s that creating new data in general is expensive and also takes time. And so what we wanted to do with Datapoints is that we wanted to scale data production and we wanted to capture those efficiencies through that scale and through that automation and pass those savings down to the customer with the idea that the customers would effectively buy more. They would use their dollars that would be much more powerful they’d be able to train more performant models and this would start that flywheel. But the other part of it is that actually turnaround times are important too. So we’ve really stressed turnaround times. We make our business as seamless as possible so that you can get the data that you need, whether that be for a large training run or it needs to be for a validation lab-in-the-loop type setting. Automation, of course, is the linchpin behind all of this. We like to think about automation not as a human replacement, but as a multiplier on every FTE that we have. And so humans are still a really important factor into this and I can’t understate that or overstate it. I’d argue that they’re actually the most important factor. But each one of our scientists can actually accomplish much more than that they would have been able to do alone with a single pipette now that they have automation. And that has a huge impact on productivity, has a huge impact on direct costs, and of course, we hope to have a huge impact on the types of data we can generate at a price point and with a certain time turnaround that would allow our customers to really feel like they are enabled and empowered to train better models.

Ross Katz: It’s interesting to hear you we’re in AI data land so I guess you can’t you can’t record a podcast in AI data land without, without mentioning Jevons Paradox. But the but I’m interested in how you evaluate the validity of that paradox inside of, it inside of biotech. Like, what would you need to see in the marketplace or what would you want to see in the marketplace that would tell you that as the price goes, goes down that the demand will increase even beyond what you would expect?

John Androsavich: We’re already seeing it. We’re already seeing it in the way that folks are using this. And as I mentioned, there’s a range of customers that we work with and a lot of folks are using ADME-based prediction models and then they’re fine-tuning those ADME models using ADME-1 for their particular chemical series. What we’re allowing with this price point is allowing them to do that earlier and sooner so that they don’t need to triage their molecules so quickly. A lot of times we conduct a drug discovery campaign and the idea is that we have this funnel and very quickly what we want to narrow down to just a handful of molecules. In order to get there in an efficient manner, the conventional way is that you just make triaging decisions and the idea of ADME, it’s funny because there’s different tiers. You always talk about Tier 1 ADME. Now I’m mostly from the RNA space, I’ve done a lot of oligonucleotide discovery, a lot of RNA LNP-type discovery. Small molecules I’d have to go back to grad school when I was doing chemical biology work to really familiarize myself with the space. But now that I’m back in it with Datapoints, I’ve done so and it’s given me an opportunity to look at ADME with a fresh set of eyes and a different perspective. And when I first came in here I had the naive perspective that Tier 1 ADME was something that everybody agreed were the most important assays. What I learned is actually it’s just what your budget allows. And so Tier 1 ADME is one or two assays or maybe three or four depending on how much of a budget you have. Either way the idea is that you eventually find a number of compounds at the end and they get a full workup. What we’re seeing now and what we’re enabling is everyone to be able to make their measurements of the full Tier 1 panel that fits in everyone’s budget early on and in more molecules. And I think this is really important because it allows you to make more effective decisions on your drug discovery program. And so that’s like a one-time consequence or a one-time payoff. But even broader than that, it allows you to generate data that effectively could be used in model training. And so one of the things that’s missing if you triage everything first is you’re missing a lot of negative data. And so by allowing folks to actually buy more early on, it does give you that ability to have the negative data that can be used for downstream models. So I think it’s paying off and we’re certainly seeing this in the way that we’re engaging customers and that they’re coming to us earlier and asking for larger orders.

Ross Katz: That’s interesting. What I’m hearing is that basically the availability of the lower price ADME ADME assays gives you the ability to de-risk earlier in the in the development of the drug. And then also because you’re able to get a higher volume of data and these models that people are using for predicting, for predicting things like ADME are so data hungry by its, by its nature the drug discovery program writ large across organizations becomes, becomes one that relies on that data sooner and sooner, which logically leads to the conclusion that you need to acquire this data, you need to acquire this data as early as possible, which the price point allows. Am I thinking about that right?

John Androsavich: Yep, that’s exactly right.

Ross Katz: Yeah. So another another thing that I’ve heard you talk, talk about related to underinvestment in biological data is this idea that we don’t yet know the scaling laws for biological AI. How much, how much more data leads to better performance across tasks or in the various tasks that we have in the biotech data space. And at what point returns start to diminish. And so this uncertainty causes hesitation on the part of potential buyers for biological data. So I’m interested in how you advise biotech organizations to think differently about data acquisition and start to, and try to characterize what scale of data they need to generate in order to accomplish their goals.

John Androsavich: I think the debate is still out on this. You can look at recent papers and there’s some really good ones. But it depends on the domain that we’re in and certainly protein language models, they seem to scale better maybe than some other ones, similar to large language models. Two examples I want to, I think illustrate the way that I talk to customers and the way that we think about data coming in. So the first one is actually from Microsoft Research. And so and their academic partners, they just released a paper this month and what it shows is that single-cell foundation models, they don’t follow scaling laws. So I think we’re learning a lot more about what doesn’t work than what does work. So it’s still progress, maybe just slower progress. But this one’s really interesting because we always think about single-cell, or the field thinks about single-cell as being maybe the epitome of good biological data that you can train on because it’s relatively easy to get large amounts of data. We’re talking about 20 to 100 million cells, lots of cells, each one has multiple data points in it if you’re measuring the transcriptome or the RNA expression in each of these cells. And it also allows you to have heterogeneity. And so that’s a really powerful thing in biology because if we think about a tissue, that tissue depending on the tissue could have six or 12 or even more different cell types. And so having single-cell methods has been revolutionary for biology and I don’t want to dismiss that. It’s absolutely true. However, for modeling, it doesn’t seem to be as performant as what you’d expect. And so what this Microsoft Research paper showed is that you can take 20 million cells for training, but your learning saturation actually plateaus around 200,000 to 2 million cells. That’s a 1 to 10 percent, and then you don’t get any benefit beyond that. I think this is a really surprising outcome for many in the field, maybe a little disappointing but we’ll talk about other methods that actually Datapoints that we’ve been using that we really think are more exciting and probably more information-rich that could help get over this saturation gap. But I think that’s one example where maybe the biological scaling didn’t allow for that the AI performance that we would expect. Another one is in our own hands here at Datapoints and that is I mentioned we work on antibody developability. Now antibody developability through our best effort is still relatively expensive. You have to make the DNA, you have to make the protein and then you start to measure these characteristics. We are actually making gains in improving the efficiencies of that and I think there’s some tricks and we hope to announce some stuff very soon. But for now we’re talking about hundreds of dollars per antibody. So we test these and we have one example where we have 250 antibodies and we have some basic predictive models on that. You can learn on 250 antibodies. But we also have one that is 10 times higher of 2500 antibodies. That used allowed us to use some more advanced models like protein language models, we were starting to get into that big data paradigm but we compared the performance of each of these models, one trained on 250, the other one on 2500. What we found is that the improvement with 2500 was not 10x better than 250. It was better, but not 10x better. And I think this was a little bit surprising. However, we went back and we looked and the catch is that the 2500 antibody dataset was actually less diverse than the 250. And this really points to an important decision-making step when you are designing data for training and that is that it’s not just the abundance of data, single-cell being an example, antibody developability being another example, it’s not just the quantity, it’s the diversity of that data and it’s the types of the data that really matter.

Ross Katz: Yeah. Another thing you’ve talked about related to like biotech companies’ orientation toward data generation is that every company that has a foundation model believes their model is already the best and so there’s not an in there’s there’s not an incentive from the economic perspective to invest in improving it further with additional, with additional data. And there’s, this is partially because you’ve got a strategic orientation where companies are competing on the narrative underlying their model and the pedigree of the researchers that have released the model rather than the model performance itself. So I’m interested in your perspective on this dynamic and what it looks like in the field.

John Androsavich: Yeah, I have an insider perspective on this being on both sides of the table, that is pharma buy-side doing search and evaluation and also now on the biotech sell-side. And what’s funny is I saw a post from Alex, the CEO of Insilico on LinkedIn a few weeks back now. If you follow Alex he often offers some really hot takes and so very interesting guy to follow. But in this case what he suggested actually made a lot of sense and that was how about if we could create an AI agentic method for doing biopharma BD it’s such a wonderful techno-utilitarian view but the problem is that biopharma partnering is a marketplace. And like the economist Richard Thaler pointed out is that marketplaces are not always rational. It’s classic behavioral economics. And so that is there are so many different aspects at play in any pharma partnering decision that goes well beyond what is the best model. And we’ll talk about how do you even evaluate what is the best model. But it means that you have to start thinking about does it fit my strategy? Who is the internal champion? What are their motives? Is there endorsement from the C-suite? Are there preexisting relationships at play? How much might geopolitical issues play into this? That also has a lot to do with it. And so often it’s more than just having good data or good models and it’s that story component that really is fundamental if you have a leg up from your competition because of your pedigree, because of your investors because you have these big names. This is this is part of the idea that when VCs are investing, they invest in the jockey and not the horse. And so I think that having a story is really important. Now the question is how do search and evaluation folks actually make a decision on who they’re going to partner with if they truly want to find the best technical solution? And we can talk about some concepts for that, what are the best frameworks. But in general I would say that you’re limited as of today and you’re often extrapolating on potential. So that’s what everybody’s doing. And that can really be a, a situation for these model developers where they might be talking to their board, thinking about the best way to deploy their capital, and they’re not sure if unless they’re in a head-to-head competition with somebody, how do they demonstrate the value of doing additional data generation and having better performance. Certainly a competition could help there, but also it could hurt and I think there’s some risks there. So often times we get trapped in this narrative, I think both sides get trapped in this narrative of thinking that let’s do a BD deal, let’s find our pharma partners first and then we’ll invest more into our models if we need to. And I think that in general is a little bit of a brake on what would be some additional investments into this field if we did have a more empirical way to evaluate the winners and losers.

Ross Katz: Yeah there’s a, just hearing you talk about it, I can think about the strategic considerations from both sides and from like a modeling company perspective you’re trying to only incorporate the data that you need in order to get that market validation that you can and get to the level of partnering with the organizations so that you can solve real drug discovery, discovery problems that lead to revenue for the model and then you can tell yourself a story about how you’re going to program after program like climb your way climb your way up and broaden the broaden the ability of the, of the model. But the idea of purchasing enough data to feed a foundation model that’s going to characterize the breadth and diversity of drug space regardless of the modality that you’re that you’re using it’s it seems out of reach and then from the, from the pharma, from the pharma perspective you’re willing to try anything that might get you to a high-quality validated target that you can that you can bring to market faster. But the idea of, the idea of investing, investing the money without the validation upfront is too much to get over so there’s like the, I can see why that deadlock arises. Do you, do you think I’m thinking about that right?

John Androsavich: Absolutely. And I’m glad you mentioned money too because that’s one dynamic I hadn’t considered in this. In that is if you really want the best predictions from these best modelers, it is not free. Inference is expensive and so just coming up with the predictions is really a task that you have to figure out and you don’t want to be in a situation where companies are not putting forth their best prediction because maybe they’re running multiples of these. If they have to do this for every single pharma company, it becomes prohibitively expensive and then maybe they’re not putting their best foot forward. So I think there’s, there’s a lot of details to work out in this type of framework. But ultimately I do think that the field will evolve and eventually get to something that is more level setting and we’re all scientists at the end of the day and it will be more data-oriented than what it currently is now just through extrapolation.

Ross Katz: Yeah. And this sort of puts into relief the kind of rational defection that you’ve described about how like the there are players in the space that are more than happy to just let to let other players invest in the generation of data and the development of, and the development of models and wait to invest until the inexpensive solution is like is available to them. Is there a case that you would make against that kind of logic sitting on the sidelines?

John Androsavich: You know, this isn’t new to our industry, the biopharma industry. It’s been happening for some time now. You can actually look at over half the drugs that are approved today, they don’t originate in the pharma companies that sell them, okay? They come from biotech. And that’s now the working model. So VCs fund early innovation in biotech, pharma buys it up later, and that’s just the way it works. Now that’s not to say that pharma doesn’t do good research themselves, they absolutely do but their risks are a little bit more less and they tend to be more conservative and that’s buffered by the fact that they can always come in and buy it later. Now I think AI is actually accelerating this trend and also shifting it to some extent. I don’t think it’s necessarily a, like a free-loader problem or anything like that. I think it’s more of a paralysis of choice problem where everybody across the industry including pharma as well as smaller modelers is that in a field that moves so quickly, where can you go and invest your dollars that would be able to compete your, compete with your beat your competitors? So it compounds though, I think, in that there’s statements made by leaders in this space that they don’t need any new data. It’s all about model architecture. Now that statement was made a couple years back and it’s likely been walked back today, in fact I heard that the individual that made that statement actually they’re building their own labs and so maybe they’re coming around to it. But I think that type of statement, that kind of thought, whether true or false I don’t know, and it’s probably somewhere in between in some cases it is actually right. But I think it scares people into looking foolish. They don’t want to look foolish. They are concerned about taking the expensive approach, buying the data, and then being proven by someone wrong that if with better AI researchers they could have figured it out with the existing data today. And I do think that creates this second or third or fourth take on it as you’re making these investments. And that just slows things down and the longer it slows down then you think you’re even further behind then you don’t make the investment again so it is really compounding in that way. So I do think that this is a, a concern in the space, a problem in the space. I’m interested to see, Ross, how this actually evolves now that the large frontier hyperscalers are getting into the space. Anthropic announced that they’re getting into drug discovery. Now there was a time when IBM was also in drug discovery and clearly they didn’t make any drugs. And so who knows how real this is. But is it that you now get less VC funding into these smaller biotechs that are actually more domain-specific because the VCs are concerned that the frontier model companies are just going to overtake the whole industry. And so that’s a real thing and I think it’s a similar problem to what we’re discussing here.

Ross Katz: Yeah there’s a, it just hearing you talk about it, I can think about the strategic considerations from both sides. It is impossible to say how successful these non-biotech companies can expect to be in the biotech space. The biotech industry also has a long history of people who understand data science and machine learning and statistics and software engineering coming in and thinking that because they’ve solved complex problems in other domains that they can necessarily solve the most complex models in the biotech domain. So I think naming IBM in its previous drug discovery efforts is an appropriate grain of salt for all of us to take along with this evolution. I want to talk about some of the initiatives that Datapoints has in place and how they relate to the way you’re positioning yourself in the ecosystem. So you’ve launched this can you just introduce us to the Virtual Cell Pharmacology Initiative and what it’s doing and why you decided to launch it?

John Androsavich: So the Virtual Cell Pharmacology Initiative is really an attempt to get the industry to recognize and use a new method for transcriptomic measurements. This goes back to the single-cell data. There are approaches to measuring transcriptomics that are more high-fidelity, more high signal-to-noise than single-cell alone. One of those approaches that we use is called DRUG-seq. Now DRUG is actually a pretty contrived acronym, I won’t even explain what it is. But actually I think it’s on the nose because you can use this method to actually assess the effects of a chemical perturbation, that is you add a compound, a small molecule compound or chemical or drug onto a cell and you are able to measure the transcriptomic response, the change in RNA gene expression that happens as a result of that drug. So Drug is useful but you can use this in many different applications. So your perturbation could be a CRISPR perturbation or genetic perturbation where you knock out a gene and you see what the response is. So it is effectively a very broadly applicable technique. And we do this in 384-well microwell plates. And each microwell plate has its own unique perturbation, either drug or CRISPR gene knockout, or you can even do it in combination. You can add biologics to this, any perturbation you can think of you could do it. And that’s the advantage of doing this on an arrayed format. So that is instead of doing a single-cell experiment that is all pooled, it’s all in one tube and then you use the single-cell sorting to really see what the outcome is on each individual cell, these have a literal physical plastic between each experiment on that 384-well plate. So each well is an experiment. Now in each of these wells, we’re able to measure 10,000 genes on average. Those 10,000 genes we measure with a lot of reproducibility and they have very high signal-to-noise and we do that on the order of ten dollars per well. So it is a very efficient technique. And for a full 384-well plate, you’re able to test a lot of different conditions and generate a lot of data for less than five thousand dollars. And I think that’s a really important technique. It was not discovered by Ginkgo. We have, I would say, perfected it. I don’t know if that if we’ve reached that end stage, it gets better all the time actually. But certainly we’ve been able to scale it and it was originally invented in 2018 by Novartis. And it’s been under the radar. And I think it’s mostly been under the radar is because the field due to influence from a lot of influential institutions such as the NIH, has really focused solely and singularly on single-cell techniques. And so I think that single-cell again, really revolutionary way to measure biology, the heterogeneity of biology. It is alluring because it has such a high amount of data output. But I would actually say that it is high in calories, low in nutrition. It’s like the junk food, of transcriptomics. While with DRUG-seq, we’re actually able to have much higher fidelity measurements and sometimes we can get more than 10x the number of genes. The problem with single-cell is you get a lot of dropouts. And so you create this what’s called a sparse matrix and most of what you’re measuring are zeros, okay? So huge amounts of data, but you have a lot of these dropouts where you’re measuring the genes expressed from each one of these cells and even if it’s a highly expressed gene, sometimes it drops out. And of course, the correlation between expression and dropout is very significant. So you often have more dropout and more medium-lowly expressed genes which are often the most sensitive to these types of perturbations. Now with DRUG-seq we do make some compromises in that we don’t measure every single gene that you could with bulk RNA-seq. That is more on the order of three hundred to four hundred dollars a sample, remember we’re at ten dollars a sample. So there are some limitations to that. But we still get 10,000 gene measurements on each one of these these samples. So it’s a lot of data. We think it’s sufficient data to actually train models on. And the idea is with Virtual Cell Pharmacology Initiative is that if we’re able to not only tell people that this is the best method, but actually demonstrate it and get the data into their hands, they will start to adopt it more and see the advantages of it. And I don’t think it needs to be an or. I actually think an and is really useful. We’ve had some early adopters of the VCPI data that have combined it with single-cell models and they’re seeing actually improved performance across the board. So I think it could be complementary as well, which makes a lot of sense. And so I’m really excited about this initiative. There’s also a special aspect of it where investigators, scientists, hobbyists, whoever they m- be might be, now the hobbyist may not have access to these types of things, but the invitation is still open that you can send us your own chemical compounds and we will screen those for free. And so you can send us we’re doing 100,000 of these compounds. And so you can send us 10 or 100 or 1000, whatever it might be, and we will screen those for free and we will put them on the platform and then everybody can benefit from this data being out there. So I think it’s really unique in that way. It’s taking a different approach to transcriptomics which are really the bedrock of virtual cells. It’s offering an orthogonal measurement there. But also we’re making it fun in that it is open source, it is all MIT licensed you can use it to train models, you can use it for commercial purposes, and you can actively participate by actually submitting your own compounds.

Ross Katz: Yeah, I mean there’s so much there’s so much here I want to like the, there are so many questions that I want to ask. But the like as you’re talking what I’m hearing is the DRUG-seq at ten dollars per well is what you all view as a, as a Pareto optimal data point that is both highly informative and, well it’s highly informative, it’s highly reproducible, and it’s very cost, it’s very cost-efficient. So you’re able so the argument to the ecosystem at large is here look at this data, it’s very informative, train some models on it and see how informative it is and then that builds some momentum behind acquiring more of this data and starting that flywheel. Am I thinking about that right from the Datapoints perspective?

John Androsavich: That’s exactly right. And one of the challenges with DRUG-seq, I mentioned that there’s a focus on single-cell that maybe is causing folks to overlook DRUG-seq. I think the other thing too is that it was started in pharma companies and it really is most effectively implemented within companies that have large-scale automation. Now pharma companies are not going to open their doors to anybody who wants to do DRUG-seq and let them in. But Ginkgo Datapoints has and so I think for the first time we’re unlocking this technique we are really democratizing it and allowing people to take advantage of it. And part of it is just awareness and we’re really hopeful that once people not only get to understand the technology but they see it and they use it is that it gets adopted a lot more thoroughly across the industry for various applications. I think this is useful in tox screening, I think it’s useful in maybe cosmetic screening across the board. You could think about environmental toxins this could be useful for. It is very broadly applicable. And the great thing about it, Ross, is that you don’t need to retune it every time you have a new application. And so you can very quickly start to test different types of perturbations and you can use it in almost any different cell model, advanced organoids, primary cells, we do this in primary blood mononuclear cells, so PBMCs which are can be derived from patients and you can do it in simple immortalized cell lines too. So it’s applicable across the board.

Ross Katz: That makes a lot of sense. I’m interested in where you think VCPI fits into the sort of broader Virtual Cell ecosystem. So you’ve got the Arc Institute with the state the state datasets models and the Virtual Cell Challenge, you’ve got the Chan Zuckerberg Initiative, you’ve got Tahoe 100 million what do you think that this work that you’re doing with VCPI adds to that ecosystem or where does it, where do you would you differentiate it from the other initiatives that are underway?

John Androsavich: Diversity. Every single one of those initiatives that you just discussed are all based on single-cell techniques. There’s nothing wrong with that, although we do have more and more data showing that there are limits to that technology. But it’s really having that orthogonal measurement that’s really important. And I think that the complementarity of DRUG-seq is going to be important to virtual cells. What’s key and what’s missing, and this is something that in the future we need to do, is we need to bridge that gap as I alluded to before, what’s the power actually modeling on both these types of datasets? There are some aspects of single-cell that I just can’t replicate with DRUG-seq. I’m not saying it is great for every purpose. But I think for virtual cells it can really be a, an accelerant and it could be really a powerful tool for folks to improve models that maybe have fallen short before and just single-cell data. But if you’re looking for just pure volumes of data, lots of it’s zeros but also actual actual integers above that and maybe for heterogeneity too single-cell’s the right way to go and cell classification etc. So I think both methods are actually quite important and I love to see that come together and for these these initiatives also to bridge and come together too.

Ross Katz: And that and that makes sense that if there’s a lot of if there’s a lot of focus industry-wide on single-cell data and there are a lot of methodological problems that are yet to be resolved with how you work with single-cell data and also you’ve got the cost constraint of trying to trying to generate single-cell data that if there’s an opportunity to build on that single-cell data or integrate with that single-cell data in a way that is more cost-effective that is reproducible and you can target the dataset at the set of the set of compounds that you want to focus on then you’ve got you’ve got an opportunity to do something special at a at a very high value.

John Androsavich: That’s exactly right.

Ross Katz: Yeah. Awesome. So another initiative that you all have underway is your work on the Antibody Developability Consortium, this federated learning initiative with Apheris and people who have listened to multiples of the Data in Biotech podcast might remember Apheris CEO Robin Roehm came came on and talked about federated co-folding so I don’t want to talk about like what federated learning is or what Apheris does but you’ve got this you’ve got this theoretical fix for data sharing in biotech where everybody gets the benefit of the model, your data stays private but you’re able to you’re able everybody gets the benefit of the of training on each other’s data so that they can so that collaboratively they can get further than they would otherwise without giving away the proprietary data that they have internally. So I’m just interested in what your role is in the Antibody Developability Consortium and what you think the benefits and the limits of that model are.

John Androsavich: Yeah, Robin first off has been amazing thought leader in the field of federated learning and Apheris has been an incredible partner to us. We’re excited about the Antibody Developability Consortium because it is very unique in that we are not training on historical data. And so this is where Ginkgo’s role is really special. Most consortia today, in fact all of them I know of with a caveat we’ll go back to that are based off of historic data. So that is you have some also in developability that are a number of pharma partners coming together, they have historic data that’s been collected over years if not decades. They put this into a federated architecture and then they learn from those data. I actually compare that to doing a puzzle in the dark, okay? Meaning that you can’t see the data in that federated environment, but you can train on it. So you’re passing the weights back and forth, but you’re not actually seeing it yourself. You’re making a lot of assumptions into how that data was actually generated and the assay parameters on all the other things. You can see the distribution of those data, but you could don’t really have necessarily the details. And let’s face it, they’re probably are some limitations to the way you collected that data. In fact one company alone probably has a change in the methods that they used from the start of that historical data time point, right, to the end of it. And so there’s probably diversity even within one partner and then you’re you’re training on this across multiple partners. We’ll see how that turns out. I think that there will be limitations to it but certainly could be helpful and I think it’s an important it’s important for me to acknowledge the amount of effort and work that has already gone into establishing this type of relationship between multiple pharma partners. That is an incredible amount of progress to get pharma to partner in this way and to actually be pre-competitive and say let’s come together and learn together. That is not an inherent trait of pharma. Pharma often times is very hyper-competitive, they’re very secretive, and they don’t work well together. There are exceptions to that and I think in the AI field, we are certainly seeing consortia be an exception to the rule and I think that some of the work done with Amgen and others through either the FATE consortium or the MELLODDY consortium, while there may be limited returns in terms of model performance certainly just being able to do that activity with multiple pharma’s involved is impressive. So we’ve gone through those business development those partnering challenges as well. It’s worth the effort in our opinion. We’re right on the level or right at the stage of taking this to the next level where we’re going to be in the data generation phase. And this gets be gets back to like the special nature of our consortium and that is instead of using historic data, we are actually building a made-for-purpose training dataset that is going to sit in the center of this federated network. It is going to be designed through contributions of our member partners. Currently we’re looking at three or four founding members, we can go up to five, maybe six or seven, I think there’s there’s really no limit, but we’re planning this around the idea of five. We’ll have up to 10,000 antibody sequences total. Each partner will actually submit 2,000 sequences and they will be able to get the data back on those 2,000, the actual data, the actual kind of CSV file came back, the data return. But all the other ones the remaining 8,000 is that they will be able to train on that through this federated environment. And what’s neat about it is that this is meant to be very collaborative and there’s a lot of sharing of ideas and best practices, but at the same time it is reserving each member’s autonomy to be able to do it on their own too, meaning that you can bring your own AI/ML team to this and through the federated environment you can train on that central training dataset. At the same time Ginkgo with the help of our collaborators and our academic partners, we are actually building a central model ourselves. And so that will actually act as a baseline for everybody else and if their modeling approaches maybe aren’t as good as ours well at least they can fall back on ours. Or if their modeling approaches are great well then they keep their model and maybe their model’s better than anybody else in the consortium they don’t share that. So I think that it balances out those different incentive structures. But what I’m really encouraged by is that we’re going to develop all these data with the same exact protocol that we’ve developed here at Ginkgo and that we’ve been delivering data to our customers for over the last two years and we’re going to do this at a scale that I think has never been done before and we’re going to do it on behalf of our partners who are early adopters here and willing to work in this collaborative environment and I think that’s a really nice outcome that we’re right on the precipice of seeing.

Ross Katz: And if I and if I understand correctly if you’re able to demonstrate the value of the model slash models that you’re that you’re producing here then what it does is it is it creates a clear path that embeds your like the protocols that you’ve established, the data generation procedures that you’ve established in the heart of these really valuable models. Should I think about that as like a strategic consideration from Ginkgo’s perspective in DataPoint Datapoints involvement in this or how would you change that characterization?

John Androsavich: Well I think it’s always been a bit scary for me to think that there are certain standards of what we call a drug-like molecule when it comes to an antibody, yet we have different ways to measure those standards. Right? As a field, we can all agree that it should have a certain amount of hydrophobicity and polyspecificity and these types of measurements but if we’re actually taking different approaches to measuring those and having different outcomes, that’s probably a bit of a, of a exclamation park on the exclamation mark on the field in general. So I think that the need to standardize these assays and these outcomes are actually emergent and necessary for the industry as a whole even without AI. But it becomes even more important when you’re starting to model and train etc. Now the caveat that I showed before in talking about made-for-purpose training data and this is I think adjacent to it but it’s not on the nose here and that is I’m really excited to see the work that Lilly TuneLab is doing and the way that they’ve built out their federated learning. Now that one is not so central but rather they have their models that are available to I think now a hundred different smaller biotechs that are part of their network and the biotechs it is a kind of an exchange of services here where the biotechs can access the models that Lilly develops, but in exchange for actually contributing their data to it. Now the way that they do it is that they have certain multipliers or credits and you get a higher credit if you use a Lilly-like protocol. And I think that this is a really interesting thing. So I think the field is moving towards having these more established protocols, still leaving room you don’t have to use the Lilly protocol but you’re incentivized to do it, but still leaving room for some divergence from this. And I think what’s interesting for about the Ginkgo Apheris consortium is not so much that we’re trying to enforce our standard on everybody else, but rather we’re trying to actually take everybody else’s standard and come up with a central solution that is actually good for the industry and matches a lot of the data that they already have and that is one of the existing preliminary conditions I would say to participating in the consortium is that we’ve done these pilot experiments with all of our members, they already know what the coherence is with their existing data and so I think the good news is that we’re already pretty well aligned and it’s all about refinement from there.

Ross Katz: That makes a lot of sense and anybody who’s interested in TuneLab we did have Aliza Apple on the podcast to talk about to talk about TuneLab on a previous episode as well. I think the to if I were to update what I said earlier it would be that the establishment of mutually beneficial standards encourages everyone to invest in data generation and train more m- and train more models that can build on the can build on each other and where each data point that gets that gets generated is more valuable to each individual organization and more valuable to the ecosystem at large and so as the value of each data point in the market goes up that’s beneficial for data generation generally and so that’s the that makes a lot of sense. You’ve built this autonomous lab designed for machine-to-machine or you’re building this autonomous lab designed for machine-to-machine operation and you’ve got this collaboration with OpenAI. Can you talk a little bit about about the work you’re doing? And I’m really interested in what a lab-in-the-loop architecture looks like.

John Androsavich: This is really neat. You can come by our Seaport facility. This did not necessarily exist this level when Datapoints started. We have a lot of different automation throughout Ginkgo. One piece of automation that we’re we’re investing heavily in though is a unit called a RAC, a reconfigurable automation cart. We’ve actually been developing this internally for many years now. It was actually started at Zymergen and we acquired Zymergen. So it has about seven or eight years of product development to it. And it really finally reached the point now and I think that the field is actually ready for it in that we’re establishing these large system-like autonomous labs that are able to generate scientific data not just at high throughput. We’ve seen that before with work cells. That’s some of the technology that Datapoints uses from other automation providers. That’s not so new. What’s new about the autonomous lab is that we have a hundred different instruments on one system and it’s really meant for diverse data generation. And so it can do more than one protocol at a time, 80, 100 protocols at a time and it really is replicating what is happening in the lab today from a physical architecture perspective. That is if you go into a lab, you see this linear layout where you have these black bench tops with instruments on each on each section of it. And then you have a human that actually goes about and transfers samples from instrument to instrument in the order that they need it and at the time points that they need to do it. This replicates that. It is a linear setup. However, it is a magnetic rail that transfers the plates from one instrument to the next and instead of a human dictating when that should happen, it’s of course a human writing the protocol but actually you have then the Catalyst software that we have driving where the plates need to go at a specific time and it’s orchestrating all these in real time so you can have multiple protocols at once. Now, I mentioned and I said of course it’s a human designing the protocol, but that’s not always true. And so we did a really neat collaboration with OpenAI where they used GPT-5 to actually design experiments to optimize the output of a cell-free protein expression system. And so this is really important. Cell-free protein expression systems simply speaking are you take away the cell membranes, you have the rudimentary mix of what you need in order to translate proteins and make proteins and you’re able to do this with a number of components but those components need to be optimized. The ratios of those components, which ones you add, are there any that you can remove. This is a big multi-parameter problem. And of course there are consequences to what your recipe is. That is you can not only have higher or lower titers but there’s also a cost perspective. And so the goal for the industry and cell-free protein expression is used in some drug development and actually commercially approved drugs, it’s also used in other industrial applications. But the idea is that you get a higher titer out a higher titer of protein expression from each volume of lysate that you use and then you have a lower cost point. That’s the goal. So less more protein less cost. We did about 36,000, 40,000 different reaction conditions using GPT-5 to actually select those reaction conditions. What was amazing about this is that GPT-5 was writing its own ELN and so we always often talk about what is the interpretability of AI models. In this case you actually had the LLM telling us exactly why it chose to do what it did and it also was looking for new reagents. Like what are some reagents that we haven’t thought about before? And they actually suggested some pretty unique ones, some of them we couldn’t actually source. There is a human component to actually doing procurement for the reagents and loading the reagents onto the system. But for large for a large part of this GPT-5 was actually designing the experiments and the protocols and then the autonomous lab that we have, what we call Nebula based on these RAC setups, these reconfigurable automation carts, they were really doing the experimentation. And I think this is a really neat example of what can be done with autonomous labs. Now it’s not just a GPT-5 that can do this. The humans at Ginkgo do this. We also have a Cloud Lab interface where other scientists can actually access this and submit samples as well. And there’s even some sense that maybe you could have a bake-off one day almost like the go-type of competition where our cell-free protein expression did reach a new goal. We successfully had the lowest price point per titer in the industry that has ever been reported before and GPT-5 did that. There are some academics here in fact one that probably was influential in the way that GPT-5 learned what some of the best experiments could be Michael Jewett from Stanford who is an expert in the space. Maybe one day we could actually see him and you actually have a human that designs experiments and if he had the horsepower of Nebula behind him perhaps he could even come up with a better outcome. So I think it’s a really interesting way to think about where the field of autonomous lab is going and the role of humans in that loop.

Ross Katz: Yeah, I mean that’s an entire podcast episode to itself. But as we head toward the end here I just have a couple of a couple of final questions for you. Like the, if you were advising a mid-stage biotech on how to think about their data strategy in the context of the bio-AI landscape, what would you tell them?

John Androsavich: I think it depends on who the bio tech company would be. Now mid-stage biotech, are you looking to make drugs or are you actually looking to monetize a model? If you’re looking to make drugs I would advise you to stay on track think about the critical path but be as hungry as possible for data on that critical path, very similar to the ADME story that we talked about before, make as much measurements on your early series as possible and it’s going to benefit you down the road. I would say for the p- the modelers, you have to think about your platform as a, a hamburger in many ways. That is if you’re going to build a hamburger shop you’re probably not going to grind your own meat, bake your own bread grow your own lettuce, you’re going to source a lot of these things. But what the really important aspect is understand what is your special sauce. What are the one or two ingredients that you need to own that makes your hamburger special and really invest in your data there and make sure that it matches your model architecture and is consistent with your entire business strategy. But once you have that I would say that you really need to start investing in a way and again it’s not just quantity of data, it’s the quality and diversity of the data that really matter.

Ross Katz: No as you’re, as you’re talking it strikes me that like all of the success in the protein modeling space is the based on the, based on the quality of the data that was available and the inductive bias that the modeling architectures had available and so this beautiful coupling of the of data and model is what led to the success that we’ve seen from, like from AlphaFold. And so that as a modeling company that’s what you’re searching for is that sort of synergy between the data that you’re generating and the and the type of model that you’re architecting. Do you think I’m thinking about that right?

John Androsavich: That’s exactly right. And also the ecosystem tools that you pull in to help you to give you that uplift as well is that you don’t need to invent everything from scratch.

Ross Katz: That makes sense. Do you have any final thoughts that you would want to leave us with as we, as we prepare to let as we prepare to let you go?

John Androsavich: Yeah this has been a really fun conversation. I want to end it on a high note. I’ve expressed a number of criticism, I talked about gaps and maybe offered some solutions on this. But I’ve never been more excited about where we are in the space. I consider this to be a point in time and it’s I recognize it’s an early point in time. As scientists we are always a little bit more cautious, we’re data-driven and once we see the data then then we can make an informed decisions and move forward. And I think what we’re looking at now and we talked about scaling laws and what type of scaling works, what doesn’t, that’s a barrier but we’re going to solve it. We talked about how do you actually understand what the best models are and how do you empirically evaluate this across the industry. Again, that’s a data problem we’re going to solve it. All these things are solvable and the good news is that us as scientists especially in the in the biotech world we are very well equipped to actually solve these problems that are currently presenting themselves. I think it’s just a matter of time and of course investment. But the investment is only going to continue to increase as we start to see what the return on that investment is and I think that’s already showing up and will continue to do so in the years to come.

Ross Katz: That’s great. Great place to end it on. Thank you, John.

John Androsavich: Thanks, Ross.

Jason: And that’s it for this episode of Data in Biotech. If you enjoyed the episode, please subscribe, rate, or leave a review in your podcast platform of choice. See you next time.

Frequently Asked
Questions

Why does biotech underinvest in generating biological data?
John Androsavich frames it as paralysis of choice rather than free-riding. In a field moving this fast, nobody can say which dataset will still matter in eighteen months, and a public claim from a respected lab that model architecture matters more than data makes the expensive path look naive. Buyers hesitate because being proven wrong by a competitor who got there on existing data is more embarrassing than being slow. Biopharma also has a working model that rewards waiting: more than half of approved drugs originate outside the company that sells them, so pharma can fund early innovation indirectly and buy the winner later.
Do scaling laws hold for biological AI models?
Not uniformly. A 2026 Microsoft Research paper found that single-cell foundation models trained on 20 million cells saturate their learning between 200,000 and 2 million cells, so 90 to 99 percent of the data adds nothing. Ginkgo saw the same ceiling in its own antibody work: a 2,500-sequence training set beat a 250-sequence set, but nowhere near ten times, and the larger set turned out to be less diverse. Protein language models appear to scale more like large language models, so the answer depends on the domain.
What is DRUG-seq, and how does it compare to single-cell RNA sequencing?
DRUG-seq is an arrayed transcriptomic assay invented at Novartis in 2018 that runs one perturbation per well in 384-well plates, measuring roughly 10,000 genes per well at about $10 a well. Single-cell methods pool everything into one tube and sort afterwards, which captures heterogeneity but produces a sparse matrix full of dropouts, and dropout correlates with expression level, so the moderately expressed genes most sensitive to perturbation are the ones most likely to go missing. Androsavich calls single-cell data high in calories and low in nutrition for model training. DRUG-seq gives up whole-transcriptome coverage that bulk RNA-seq provides at $300 to $400 a sample, and Ginkgo positions it as complementary to single-cell rather than a replacement.
How is Ginkgo's Antibody Developability Consortium different from other federated learning consortia?
Most consortia federate historical data, which Androsavich compares to doing a puzzle in the dark: partners can train on the distribution without seeing how each assay was run, and a single company often changed its own methods over the decades the data covers. The Ginkgo and Apheris consortium generates a purpose-built dataset instead. Around five founding members each submit 2,000 antibody sequences and receive the raw results for their own 2,000, then train on all 10,000 inside the federated environment, every measurement produced under one protocol. Ginkgo also trains a central baseline model so a member whose own modeling falls short still has something to fall back on.
What should a mid-stage biotech prioritize in its data strategy?
Androsavich splits the answer by what the company sells. If the product is a drug, stay on the critical path and buy as much data as possible along it, since cheap early assays generate the negative results that later models need. If the product is a model, treat the platform like a hamburger shop: source most of the ingredients and own the one or two that make it distinctive. In both cases he argues the target is diversity and quality of data matched to the model architecture, not volume.

Need a data partner for life sciences?

CorrDyn helps biotech and pharma companies build the data infrastructure that accelerates research and operations.

Book an intro call