Listen on
Overview
The pharmaceutical industry grapples with a fundamental challenge: R&D productivity has declined for decades, while drug development costs and drug prices continue to rise. This isn’t just a research problem; it impacts patient access, competitive advantage, and the industry’s reputation. Traditional drug discovery, often an ‘artisanal craft,’ struggles against the sheer scale and complexity of biological systems.
Mike Nally, CEO of Generate:Biomedicines and a veteran from Merck, recognized these trends as monumental hurdles. He believed a new approach was necessary to reverse them: ‘generative biology.’ This method uses advanced AI to design entirely new proteins, moving beyond simply modifying what nature has already provided. He argues that biology is the ‘original information technology,’ inherently programmable.
In this episode, host Ross Katz speaks with Nally, who explains how Generate:Biomedicines fuses computational design (dry lab) with rapid experimental validation (wet lab), creating a powerful data feedback loop that accelerates drug development. He details their breakthroughs in de novo protein generation, the significance of tools like the open-source diffusion model Chroma, and how this new approach aims to deliver better, more affordable medicines to patients faster and at a lower cost.
Key Takeaways
Biology is fundamentally an information technology, requiring programmatic solutions.
Traditional drug discovery often relies on trial-and-error, treating biology as an enigma. Mike Nally argues biology is ‘the original information technology’—a coding problem where we can now read and program the code. This reframes R&D challenges, shifting from an artisanal craft to engineered systems capable of rapid iteration and precision.
An integrated dry and wet lab creates a self-improving data feedback loop.
Generate:Biomedicines links computational design (dry lab) directly to experimental validation and measurement (wet lab). Every generated protein is built, tested, and the results feed back into the models. This critical loop turns model ‘hallucinations’ into valuable training data, accelerating refinement and drastically improving hit rates for functional molecules.
De novo protein design opens up previously inaccessible therapeutic targets.
The ability to generate entirely new proteins—rather than just modifying existing ones—allows precise control over specificity and function. This includes designing antibodies for ‘immune cryptic domains’ that traditional methods cannot reach. This capability expands the therapeutic landscape, addressing unmet medical needs.
Generative biology directly reduces R&D timelines and drug costs.
Generate’s first clinical program moved from concept to clinic in 17 months, a significant reduction from the typical 3-5 years. Furthermore, advanced binding affinity can extend dosing intervals (e.g., from every 4 weeks to every 6 months), directly cutting patient costs and improving access. This demonstrates a clear path to reversing industry productivity and pricing trends.
Related: CorrDyn helps biotech and life sciences companies develop strong AI strategies and implement efficient data engineering systems. Learn more about how to access data value in biotech manufacturing.
Full Transcript
Jason: Welcome to Data in Biotech, a podcast from CorrDyn where we explore how companies leverage data to drive innovation in life sciences. Every two weeks, we sit down with an expert from the world of biotechnology to understand how they’re using data science to solve technical challenges, streamline operations, and further innovation in their business. This week, we sat down with Michael Nally, CEO of Generate Biomedicines, a company that is revolutionizing the way medicines are created by using a new approach called generative biology. During the interview, we discussed the use of generative AI-based approaches to construct new proteins and solve the challenges of declining research productivity and drug pricing in the pharmaceutical industry. Michael also introduced us to Chroma, an open-source diffusion model that expands the natural universe of proteins, and shared his thoughts around the feedback loop between computation and experimentation in protein design. Finally, we touched on Generate Biomedicines’ culture, one that emphasizes audacious ambition, being intrinsically digital, and patient-centricity. Here we go.
Ross Katz: Michael Nally, welcome to the Data in Biotech podcast.
Michael Nally: Thanks so much, Ross. Great to be here.
Ross Katz: Just to kick us off, would you give us a brief introduction to you and your background and what brings you here?
Michael Nally: So I’m Michael Nally. I’m the CEO of Generate Biomedicines and a CEO partner at Flagship Pioneering. I joined both Generate and Flagship a little over three years ago after having spent 18 years at Merck. Worked across most parts of drug discovery, development, commercialization, and was fortunate enough to see a—what I thought was a transcendent technology a few years ago where Generate was using generative AI-based approaches to construct new proteins that nature hasn’t been able to discover. And that belief has only grown over the past three years.
Ross Katz: What excited you about Generate Biomedicines that made you want to step out of Merck and take on this new challenge?
Michael Nally: It was one of those things where, over the couple of decades that I was in the industry, I stepped back and said, what are the things that are actually holding back progress? One was this troubling trend on research productivity, where if you look at it in the aggregate over the past four decades, we’ve seen declining productivity consistently at an industry-wide level. And secondary, you had this somewhat interrelated but oftentimes independent dynamic where pricing practices from the industry led to inadequate patient access and ultimately undermined the reputation of the industry. A lot of my career has been oriented toward how do you solve these two monumental challenges? I became a clear believer that over time, with the use of technology, we’d be able to understand biology in more profound ways than the human mind alone could understand it. While we’ve had some of the best human minds on the planet studying biology, we still only know probably about 10% of it. And if you think about how humbling that is, we’ve been staring at it for all of our existence and yet we know so little. Could this era of technology help us understand it in more profound ways? I became a big believer that this intersection would be the path to reverse those trends—that you could reverse the productivity trend, and then by reversing the productivity trend, you could rethink the pharmaceutical business model. When I saw the early data at Generate, I almost fell off my chair. I had never seen a technology that held the promise. Now, I will say, when I first met the team here about three and a half years ago, it was still very early days. The data was relatively scarce. But what became very quickly apparent was that if it were to hold true, it would change the way every drug is made in the future. And that’s what’s been so exciting about it for me.
Ross Katz: That makes a lot of sense. Can you tell us more about Generate Biomedicines and why you think this company is well-positioned to solve the R&D productivity problem as well as the drug pricing problem?
Michael Nally: If you take it at first principles, drug discovery has been an artisanal craft. It’s down to an individual genius, a team of geniuses that have a fundamental biological insight. And then basically we borrow from nature and we tinker with this in a trial and error process to try and get a therapeutic. It’s no wonder why this is such a difficult business—it’s easier to put a person on the moon than to discover a drug. And what we also know—and this has been the fascinating journey I think we’ve been on as an industry—is that when human beings are conceived, what’s passed from parent to child is a code, DNA. And yet that code then leads to all biological function. The problem is we’ve never understood the code. This is truly a coding problem. I believe biology is the first information technology, the original information technology. Until about 70 years ago we couldn’t see the code, we could never read the code, let alone program it. With the advent of new tools and technologies, we’re now at a point where programmability starts to become a distinct prospect. We’re very early in that journey. But if you think about the starting point, understanding about five, ten percent of biology—if we could double that over the course of the next decade, the amount of good we could do for humanity, the amount of disease we could cure would be monumental. In almost a simple way, whenever these complex fields become engineerable, industrial revolutions occur. We’ve seen this with chemical systems, we’ve seen this with electrical systems. We saw when Bernoulli distilled the principles of flight down to a mathematical equation, all of a sudden the aerospace revolution takes place. Biology is on that precipice. That’s what’s so exciting about the field, about where we are, about the work that Generate’s trying to do—how do we distill that down as quickly as possible to have that monumental impact.
Ross Katz: That makes a lot of sense and it is very exciting. It sounds like just from the name, Generate, that between the reading of the code and the writing of the code, the differentiating feature of Generate is your ability to write the code. Am I thinking about that right?
Michael Nally: You are. The rate-limiter for us in many respects has been data. What we’ve seen over the past decade is a convergence of three major themes. We saw a massive increase in computing power—and we see even with the news last night of these blowout earnings from NVIDIA, the demand for compute continues to grow pretty massively. The second theme has been advances in machine learning techniques. Some of the core techniques that we use at Generate have basically been invented in the past 10, 15 years. You have those two dynamics at play. But with biology, unlike the advances we can make in machine learning in large language models, we don’t have the internet to digest. Data is much more disparate. Generate was set up as an integrated dry and wet lab—and I think this is a really important piece. Every DNA sequence that our models will spit out, we build into the protein format of choice, we test, and we measure. We then have an informatics layer that sits on top where every data point is captured and fed back to refine the computational approach. Now, when you start a company like ours, you’re reliant on external data. Proteins had two extraordinary data sets: the Protein Data Bank, which gave us about 200,000 high-quality, high-resolution structures of proteins; and the sequencing revolution, which gave us 180 million amino acid sequences across all species. That allowed us to start to understand interrelationships between sequence and structure. But if you’re interested in therapeutics, what you really care about is function. Our wet lab has been our attempt to classify and quantify function in a much more profound way. We’ve now been able to generate millions of proteins that nature hasn’t discovered, we’ve built all those, we’ve tested all those, we’ve learned on all those, and the models are consistently improving, getting us better and better answers. That’s what gives us so much hope that this will be a profound revolution to address that productivity equation. And then if we’re more productive, we don’t have to be reliant on traditional pricing models. We can think differently about how you price things. I can give you some specific examples as we continue the conversation.
Ross Katz: That’s really interesting, and I love the way you characterized the feedback loops between the dry lab and the wet lab. Can you help us think through the layers of modeling that are occurring—how the broad space of proteins gets searched, how you decide which ones you want to experiment with, which data points you want to gather, and how you feed those back into the system as a whole?
Michael Nally: Traditional antibody discovery is done by immunizing a human, a mouse, or a llama. We then cull through plasma, we see what sticks to a target, and then we try and modify it to have drug-like properties. Generate’s models start with structural information. Do we have the structure of a protein binding a target of interest? Or in the most extreme version, do we just know that there’s a target that we want to interact with but we have no prior knowledge of anything ever binding to that target? That second example has been the holy grail for the protein design field—it’s called de novo generation. As far as we were aware, that had never been done before three years ago. On the first example, where you have this co-crystal structure of a protein binding a target, if we ask our models to suggest other DNA sequences that perform as well, if not better, than the reference molecule, our hit rate right now is about 60% straight out of the computer. Six out of every 10 will be highly functional molecules. And what’s been really interesting is the diversity of those molecules is extraordinary. If we look at the sequence that led to the reference molecule and then we look at the sequences that come out of the model, we can change 30, 40, 50, 60% of that code while retaining or enhancing function. Historically, the way we’ve designed or manipulated proteins is either we use random mutagenesis—something called error-prone PCR where we mutate the protein sequence—or we’ve used computational techniques, biophysical-based methods like Rosetta. Using those approaches, you can change about 10% of the binding region before mathematically you find no functional variants. This is because the search space for proteins is almost unimaginably large. If you think about the average protein being, say, 200 amino acids, and there are 20 natural amino acids, the combinatorial possibilities are $20^{200}$. That’s atoms in the universe squared or cubed—just an array of opportunities that are really hard to search. Having good prior models that allow us to search that space efficiently is at the core of what we do. This concept of de novo generation—and this is what gives us the most excitement—we saw our first de novo antibody hit three years ago, structurally confirmed. What de novo generation gives you is control over specificity. If you think about the randomness of relying on the human immune system, a mouse’s immune system, to generate antibodies that bind with specificity to a given location, we have no control over where those antibodies bind. They bind where they bind. We estimate that roughly half of the target landscape is inaccessible—immune-cryptic domains that our bodies don’t generate antibodies to. If we now have tools that can give us those specific answers, that gives us tools to prosecute biology in ways we’ve never had before at our disposal. We focus in on both of those technologies. We focus in on an optimization technology where we can co-optimize things for all of the relevant therapeutic properties in one go rather than relying on sequential processing. And that allows us to get better answers faster than traditional techniques.
Ross Katz: I love that, and maybe this is a place where we lead into Chroma and the open-source model that you’ve released—because what I noticed about Chroma is the level of control you’re given over the input to the system in order to guide the development of the de novo proteins. Would love to hear your introduction to Chroma and how you view the way that model feeds into the Generate model you’ve been discussing.
Michael Nally: Chroma is one of dozens of models that we use. What’s unique about Chroma—it was the first diffusion model publicly released that allows us to expand the natural universe of proteins with extraordinary precision. If you think about the three and a half billion years the planet has been around, nature has basically given us about 20,000 unique protein concepts. We believe that is one drop of water in a bucket of potential proteins. What Chroma—and you can see some of the examples behind me—gives us are mathematically correct proteins that when we put these codes together, perform and fold like the proteins that nature’s discovered. This goes back to what we were talking about earlier: sometimes Mother Nature is an extraordinary technician and evolution is an extraordinarily powerful tool. But given the vastness of this search space, we estimate that evolution has basically touched one drop of water in all the Earth’s oceans of potential proteins. What these tools are giving us is an ability to survey the oceans. And inevitably, when you survey the oceans, you’re going to find some pretty cool stuff. What we do with Chroma at Generate is: Chroma is trained on the publicly available data, then we have a series of conditioners that we apply to Chroma to take these protein concepts and turn them into purpose-built therapeutics for unmet medical needs. It’s this combination of diffusion models like Chroma with a series of other models that allows us to dial in very precise therapeutic attributes that address major biological challenges.
Ross Katz: What I’ve seen is that there’s a kind of scaffolding going on where you’re using the diffusion model to generate the candidates of the proteins you might explore, but then you’ve also got this Bayesian optimization approach that’s applying constraints to what might come out of it, while also providing a methodology from which you learn from the experimental data you’re gathering over time. Am I thinking about that right, or how does it look on the ground?
Michael Nally: 100%. It’s this virtuous cycle between computation and experimentation that’s so powerful. One of the Achilles’ heels of generative-based approaches early on in many other domains are things like hallucinations. With models like Chroma, when you’re using experimental validation, hallucinations become features, not bugs—because it allows us to refine the computation, recognizing the codes that don’t work as much as the codes that do work. When we created this model, we had no idea what percentage of these proteins would express, what percentage would be performant. The rates have been well beyond our wildest dreams. That tells us we’re getting pretty good at understanding the underlying mathematics of what it means to be a protein. And in doing so, it should allow us to come up with new concepts to treat some of the world’s greatest diseases. The nice thing about these protein-based models is therapeutics is just one application. We can think about whether we can manipulate enzymes—can we make a more efficient version of RuBisCO for photosynthesis? Could we think about entirely new surface materials replacing plastics? These things have transformative potential across so many domains. We’ve particularly focused on therapeutics because of the unmet medical need that could be addressed in its early days. But the potential of these approaches is far wider.
Ross Katz: That makes a lot of sense, and it is very exciting. You mentioned data earlier—is data the main limiting factor for what prevents you from exploring even more of the space? How do you think about the potential for expanding into other areas and leveraging these capabilities that you’re developing internally?
Michael Nally: It’s people, Ross, in all honesty. When I started at Generate, we estimated that the number of people skilled at the intersection of machine learning and protein engineering was about 90 people on the planet. Most of those folks were working for large tech companies, interestingly enough. Part of what we’re seeing is this blending of disciplines. Classically you were a computer scientist or you were a biologist. Blending those two was a rarity. What we’re trying to do within Generate is find people who are multilingual—who can understand computation, understand experimentation, understand therapeutic development. These are rather tribal cultures; the clinicians don’t speak the same language as the computational scientist. How you create a new vernacular for these unicorns that can span different domains is actually the biggest challenge, but also the biggest opportunity. Because a lot of times, the greatest innovations that society sees occur in these white spaces between disciplines when you’re able to fill those in. This is a classic case of having to marry three pretty disparate disciplines to get therapeutics that can make a big difference in the world.
Ross Katz: How has your view of the role of machine learning in the R&D ecosystem at a biotech organization evolved? I imagine you’ve had one perspective coming from Merck and then things have expanded over time.
Michael Nally: What plagues the space right now is some people are claiming it to be a panacea for everything. And it’s not. It is an extraordinary tool that every drug company will need to deploy broadly. That requires a lot of careful thought and consideration because you need to have a talent base that can embrace it. You need to have a culture where data is seen as a common asset, not a proprietary asset, which is a major cultural shift. And you need systems that support an end-to-end data strategy. At the same time, while I say it’s not a panacea, I believe pretty firmly that all parts of the drug discovery and development continuum will be positively impacted by these computational approaches. At different rates, at different points in time, but these tools will be deployed across the continuum to improve efficiency. One of the things we’ve been confronted with recently was entering the clinic with our first two molecules. Traditionally, therapeutic discovery would take three to five years for a protein therapeutic to reach clinical development. Our first program was 17 months from concept to clinic—a considerable savings. Then you get to the clinical stage and you dissect the problem: we need to recruit patients into a study—technology can help us do that, patients are identifiable. At the back end of a trial, you need to close a study out—that’s a digitizable endeavor. And then how do you select patients? What are the inclusion/exclusion criteria? Are there biomarkers that can make the study more effective and efficient? All of these things are going to be positively impacted by technology. But we also have to recognize that as you move toward the clinic, you enter a much more highly regulated space, and you’re going to need the concurrence of an important stakeholder like the FDA on some of these things. Being open-minded, being a digital-native biotech company allows us to think about these problems at a foundational level—because ultimately what everyone in the biopharmaceutical industry is oriented toward is how do we get better medicines to patients faster. And if we can use these technologies to do that in a safe and effective way, those are the types of things that have the prospect of improving the reputation of the industry.
Ross Katz: I love a lot of what you said, and it brings to mind two questions. One is where is the human factor most valuable to the process—where does human attention need to be placed? And then also, how do you improve the way that collaboration occurs between the models, the tools, and the people whose attention needs to be on the process?
Michael Nally: It’s an important point, because there’s this overt fear within society that these machines will replace humans. Our people, our scientists at Generate are the entire key to our success. It’s the machine learning scientists who develop the models—having great machine learning scientists that can develop the best models is what makes us distinctive. Having great and innovative experimentalists who can think about these questions through a different lens is what builds the technologies that allow us to go from 100 variants to a million variants in a cycle. And then we need the clinicians who understand unmet medical need, the PhDs who understand biology at a foundational level, to come up with the insights and say, if we’re going to direct this technology, let’s direct it here because this will make the biggest difference in patients’ lives. For us, these are tools that augment brilliant scientists. They force us to rethink workflows, to rethink the how. But the absolutely essential nature of extraordinary scientists is as strong here as anywhere I’ve ever been, and I don’t see that changing in our lifetimes. The long-term vision at Generate is to ultimately be able to instantly generate the optimal therapeutic to address the greatest unmet medical needs. That would be a great day for humanity, for society, for the world. We’re a long way from there. But we need to keep that North Star because that’s what we’re driving toward—how do we make progress in climbing that Mount Everest goal.
Ross Katz: Can you share any examples of how the workflows have evolved as the tools have improved between the brilliant people you have and the tools you’ve created?
Michael Nally: I mentioned earlier some of the work that has been recently done on cell-free expression. Basically, if you have computational models that spit out DNA sequences, the first thing you have to do is build the proteins. Using traditional methodologies, that can be a few-week process. Cell-free takes that down to a two- to three-day window. If you think about data velocity, changing out that technology gives us insight in much more rapid loops. When you have a data strategy, part of it’s the quantum of data, part of it’s the quality of data, and the quantum is directly affected by the size and speed of the experiment. We’re constantly putting in new technologies all along these pathways. Another example: if we’re screening, what sort of screening technologies can we do not in a 96-well plate—the traditional experimental volume—but using microfluidic devices? We can go from 96 experiments simultaneously to a few thousand. All of these things require entirely new hardware, entirely new ingenuity, and as you change out each of these steps, you’re disrupting the traditional workflow. We have to operate almost like a software company where we’re constantly innovating in the techniques because nobody makes these devices. And then where do you swap them in and swap them out as they become more industrialized and productionized? Quality still is of paramount importance.
Ross Katz: What do you think distinguishes Generate? There’s been a lot of AI/ML-driven drug discovery companies started over the last five years or so, and I’m interested in how you view the main differentiators between Generate and the rest of the field. You mentioned small molecules versus de novo proteins, protein engineering earlier, but any other differentiators?
Michael Nally: It’s hard for me to speak on behalf of all the other organizations because I’m sure there are a lot of other great ones. This is an extraordinarily talented group of scientists that are passionate about this specific challenge. I feel extraordinarily fortunate to walk through the doors every day and work with people who are far smarter than me, far brighter than me—ingenious in their designs, fearless in pursuit of creating extraordinary medicines for patients. I know that sounds very soft, but it’s those environments that change the world. I give a ton of credit to the founders of Generate, the team at Flagship, Avak and Molly who have been core members of the leadership team. They created a culture where we embrace six core values. The first was something that—when I read it the first time—most organizations would never say. It was audaciously ambitious. The group of scientists here, what good looks like is changing the world. Being intrinsically digital was another one, because we’re taking on a scalability challenge. Thinking about the how, we are fiercely united as a team. Insatiable curiosity. These were the things that came to the fore. And then as we’ve matured as an organization, being patient-centric. The culture here has been very distinguishing. The world needs a lot of these companies to succeed. We hope to play a small role in that, and we think with the people, with the ambition, with the support that we have, the technology’s going to make a big difference in the world.
Ross Katz: I don’t know if you’ve gotten exposed to the controversy around this, but there was a recent BCG article about the success of AI-discovered drugs, and there was some backlash to the findings from, for example, Derek Lowe poking holes in the methodology that was used to assess that. Since you’re on the inside, what’s your viewpoint on the success or lack thereof, or what this entire conversation gets wrong about the enterprise as a whole?
Michael Nally: I welcome skeptics. The reality is people don’t have to believe. We just look at the data. We’re seeing something different. And the longer people don’t believe, the longer they don’t pursue these approaches. That’s good for business. To judge this—we’re in a long lead-time business. It takes 10 to 15 years to typically create a therapeutic. How are we judging this at this point with any real thoughtfulness? The technologies of today can’t be reflected in some respects in the assets that are in the clinic today. Our core model that we did a lot of the foundational de novo work on was invented in the 2018 period. If you were trying to do that a decade ago, it would have taken 100 years to train. The molecules that we’re judging this on were almost impossible to have created in the very recent past. Just like we’ve seen this change in large language models—what society knows best is the difference between the versions of ChatGPT—we’re on an exponential curve. If people want to say it’s different in biology, that’s fine. We’ll see. We think the answers that we’re putting into the clinic, that we’re starting to see in our pipeline, are simply different.
Ross Katz: Are there any updates you can share on your two leading candidates or any other exciting things you have on the horizon at Generate?
Michael Nally: We saw the early Phase 1 data. The first asset we put into the clinic was a program initiated post-emergence of Omicron. The world was pretty scared, we were in the midst of the pandemic, and all of a sudden we saw this mutant variant that jumped every fence. It led us to ask: is there anything we could do differently than what the world has done in COVID? All of the antibodies that had been made historically targeted the receptor binding domain. This is the immune-dominant region where, if you have natural infection or vaccine-induced immunity, our bodies produce antibodies to this receptor binding domain, and they block the infection potentially. Unfortunately, the virus is a wily foe and mutates away from these blockages. But there are domains on the spike protein that are extraordinarily well conserved. Despite the world’s efforts in finding answers here, nobody as far as we’re aware had put an antibody directed toward the stem helix peptide in the S2 domain of the spike protein into human testing—because the antibodies found through traditional means were not potent and were not broadly neutralizing. We used our technology to create a potent and broadly neutralizing antibody to that domain. The Phase 1 data shows it’s extraordinarily well-tolerated, it’s safe, very low ADA rate, and it neutralizes all variants of concern experienced to date, including other sarbecoviruses like SARS-CoV-1 and bat coronaviruses. This is an answer that was probably not feasible using traditional techniques—if it was, the world would have found something like this before. Our second program is an anti-TSLP antibody for severe asthma. We put that into Phase 1 at the end of last year. The beautiful thing about that antibody is that it is an extraordinarily high-affinity antibody. The current product in that class is dosed every four weeks. Our antibody we believe will be dosed every six months, because it’s about a 20-fold improvement in binding affinity. That 20-fold improvement leads to about a five-fold improvement in potency, and gives the prospect of extending the dosing interval to every six months. We go back to where we started—how do you solve productivity? How do you solve affordability? If you only have to dose something every six months, the cost to the system, the cost to the patient is fundamentally different. That’s what gives you the prospect of ultimately democratizing the availability of these drugs. We’ve gone after areas of de-risked biology in the early days of Generate, but over time we want to give ourselves better answers to validated biology, and then ultimately take on some of the great therapeutic challenges with these unique tools. It’s still early days, but the data is highly suggestive that these tools are going to make a big difference in the world.
Ross Katz: I think that’s a great place to leave it as we bring this to a close. Michael, it’s been a pleasure to have you on. So much insight shared today, and I really thoroughly enjoyed listening to you. Excited for great things to come from Generate. Thanks so much.
Michael Nally: Thanks, Ross.
Jason: And that’s it for this episode of Data in Biotech. If you enjoyed the episode, please subscribe, rate, or leave a review in your podcast platform of choice. See you next time.





