Skip to content
Fred Manby — AI-Driven Drug Discovery with Fred Manby
Data in BiotechEpisode 35

AI-Driven Drug Discovery with Fred Manby

Fred Manby of Iambic Therapeutics explains how AI platforms and uncertainty quantification are reshaping early-stage drug discovery.

38:19Full transcript below
FM

Fred Manby

Co-Founder & CTO at Iambic Therapeutics

Overview

Drug discovery faces a critical challenge: predicting human clinical outcomes early and accurately, despite the inherent scarcity of human trial data and the vastness of chemical space. Traditional approaches often struggle to bridge the gap between laboratory results and patient impact, leading to costly failures and prolonged development cycles. This episode unpacks how a data-first approach can accelerate drug pipelines and drastically improve capital efficiency.

Host Ross Katz speaks with Fred Manby, Co-Founder and CTO of Iambic Therapeutics, who brings a unique perspective as a former quantum chemist now building data-driven models for molecular prediction. He explains how Iambic tackles this problem by treating drug discovery as a massive data generation act. Their multimodal transformer model, Enchant, integrates over 15 distinct data types, from protein structures to biomedical literature, effectively expanding the available data universe to make reliable predictions.

The conversation details how Iambic’s high-throughput experimental platform generates rich, multi-parameter data for thousands of molecules, continuously fine-tuning Enchant. Manby discusses how this closed-loop system, coupled with rigorous uncertainty quantification, not only improves lead optimization but critically enables predictive insights into human pharmacokinetics using minimal clinical data. This approach shifts the competitive landscape, allowing even young biotech companies to de-risk programs far earlier in the development pipeline.

Key Takeaways

Predicting human clinical outcomes requires minimal human data.

Iambic’s Enchant model, a multimodal transformer, demonstrates the ability to predict human pharmacokinetic (PK) properties with meaningful accuracy using preclinical data and information from as few as five human molecules. This capability significantly de-risks early-stage discovery programs by providing critical clinical insights without extensive and costly human trials. The model achieves this by learning meta-patterns across species and data types, effectively filling data gaps.

Calibrated uncertainty quantification directs experimental spend.

Rather than simply making point predictions, Iambic’s models provide well-calibrated uncertainty measures for each molecular property. This allows scientists to convert predictions into probabilities, informing decisions on which experiments offer the highest chance of generating impactful results. This data-driven framework ensures that experimental resources are allocated to maximize learning and program advancement.

Drug discovery is a data generation act; AI improves its capital efficiency.

Fred Manby frames drug discovery as a process of continuous data generation, where high-throughput experimental platforms produce thousands of data points weekly. By integrating this data to fine-tune AI models like Enchant, companies can systematically reduce the capital costs of finding new medicines. This closed-loop system creates a virtuous cycle where each experiment refines the model, making subsequent predictions more accurate and enabling more adventurous exploration of chemical space.

AI effectiveness in the lab hinges on user-centric interface design.

The real-world impact of advanced AI models like Enchant depends heavily on how easily scientists can interact with them. Iambic invests heavily in designing intuitive interfaces that allow researchers to access predictions, evaluate probabilities, and trigger automated experimental validation. This backend engineering and front-end design work ensures a fluid workflow, translating complex AI outputs into actionable steps for drug hunters.

Related: CorrDyn supports biotech and life sciences companies by building machine learning solutions and reliable data engineering pipelines. See how biotech manufacturers gain data value.

Full Transcript

Jason: Hi everyone, this is Jason, producer of Data in Biotech. Before we get started I wanted to let you know about our latest white paper. It’s a comprehensive guide to implementing machine learning models in biotech manufacturing. It’s a complete overview of all the potential problems of ML adoption and more importantly how to solve them. To download it, simply visit connect.corrdyn.com/biotech-ml. We’ve also dropped the link in the show notes of this episode. Okay, let’s get into it. Welcome to Data in Biotech, a podcast from CorrDyn where we explore how companies leverage data to drive innovation in life sciences. Every two weeks we sit down with an expert from the world of biotechnology to understand how they’re using data science to solve technical challenges, streamline operations, and further innovation in their business. In this episode we sit down with Fred Manby, co-founder and CTO of Iambic Therapeutics to discuss the innovative approaches his company is taking in drug discovery. He talks through the importance of leveraging data and technology, particularly through their multi-modal transformer model, Enchant, to address the complexities of drug discovery. We also dive into the role of machine learning to enhance drug discovery processes through fine-tuning and uncertainty qualification. Here we go.

Ross Katz: Fred Manby, welcome to the Data in Biotech podcast.

Fred Manby: Thank you so much, great to be here.

Ross Katz: Well just to kick us off, would you give us a brief introduction to your background and what brought you here today?

Fred Manby: Absolutely. I’m co-founder and chief technology officer at Iambic Therapeutics. Our mission is to best leverage the available data in drug discovery to accelerate and improve capital efficiency for the discovery of drugs.

Ross Katz: Amazing. So how did you end up at Iambic? How did your background lead you there?

Fred Manby: So I’m a chemist by training and I had a long academic spell as a chemist based in the University of Bristol in the UK. But at a certain stage I got the itch to explore more direct ways of delivering impact through the science that I was working on and in the end I was very fortunate to, together with my co-founder Tom Miller, be able to create this great company Iambic that we’re having so much fun with and, yeah, really getting the opportunity to do that direct translation into drug discovery.

Ross Katz: Can you talk a little bit about how Iambic does drug discovery?

Fred Manby: For sure. There are lots of AI for drug discovery companies. They’re all over the place. One thing that we hold as very, very important to us in Iambic is holding our technologies and our innovation to the standard of execution in drug discovery. We create technologies—many companies create AI technologies and try to apply them in various ways—but for us there’s a really hard filter. Does this technology we just created actually help with the real drug discovery program that we’re executing on by ourselves right now? And if it doesn’t deliver, we just go and work on something else and that’s okay. Not everything bites in that way. But that is a key way that we think about the differentiation of Iambic—it’s a company where lots of technology is being built, and we’ll dig into some of those technologies as the conversation unfolds. But the key principle is those technologies have to deliver, and they have to deliver in real discovery programs.

Ross Katz: I think it would be interesting to have you talk at a high level about what are the technologies that you’re bringing to bear that you’ve determined are some of the best technologies to actually have that clinical impact.

Fred Manby: Absolutely. The technology suite at Iambic basically contains two distinct types of technologies. One is machine learning technologies and an early learning for me in this space was that it’s no good to have one technology that solves one problem. Drug discovery is never about solving one problem. It’s about solving 20 different problems all in the same molecule. One really important class of technologies that we’ve developed are machine learning technologies focused in the space of molecular discovery. But the second part is a high-throughput experimental platform that delivers data. And that is completely indispensable from our perspective—other companies do it otherwise, but from our perspective a platform that delivers high-quality curated chemical data is an indispensable component. Because chemical space is really huge. You often hear this number 10 to the 60 as a rough estimate of the number of feasible drug-like molecules. Even if you gathered all of the data in the world that has ever been produced in drug discovery, it would still barely scratch the surface. So as you explore through chemical space in a discovery program, you have to generate new molecular entities and measure biological and chemical properties of those new molecules. That’s the other area where we innovate and we use a lot of laboratory automation and high-throughput technologies to synthesize hundreds or thousands of compounds on weekly cycles and bring them to a whole host of different biological and chemical assays.

Ross Katz: Can you talk a little bit about what you were just mentioning around your experimental platform and the data that you collect? As you’re preparing to explore this broad chemical space, what data are you collecting and how do you work with it?

Fred Manby: I think it’s an important axis for the ways in which different companies think about this problem differently. One choice is to go incredibly wide across chemical space. An example of the type of company that might do that would be a DEL screening company that may be making millions or even billions of different compounds and trying to measure whether they do or do not hit a particular protein target. That’s one approach—you can go wide on molecule. We didn’t choose that approach. The reason loops back to this point that in drug discovery, getting incredibly good at solving one problem does not move the needle. You’ve got to get good at solving 20 different problems. And that has translated itself into the shape of our platform. We don’t go enormously wide in terms of different molecules that we’re creating—we can talk about the way that we compensate for that—but instead we don’t make millions of molecules, we make thousands of molecules, but we measure a large number of different properties on those molecules. And that is really the framework in which we can drive forward high-efficiency multi-parameter optimization projects. In a typical program we’ll be measuring things like biochemical activity, your standard assays, cellular activity assays, but also a whole suite of metabolic assays. We’ll be doing MetID, we’ll be doing Cyp profiling, we’ll be doing a whole suite of different assays. The data from these assays are all being gathered through highly automated data pipelines and then used to fine-tune our machine learning models. And that’s enabling us to drive programs in a multi-parameter optimization framework from the earliest stages. That’s very important for us. We don’t really see hit identification as a separate thing. Why would you only worry about whether the molecules hit the protein when you have machine learning models that can also be informing you on a host of other properties? It’s a very important principle to how we operate in Iambic that we’re multi-parameter out of the gate as soon as we start a program.

Ross Katz: You mentioned some of the assays that you’re measuring in the lab but I’m also assuming that there are other things related to multi-omics data, related to protein structure and function data. Are there other modalities that you bring to bear besides the things that are getting measured in the assays that you mentioned?

Fred Manby: For sure. Complementing the data that’s generated on molecules by our platform, we of course access vast amounts of public-domain and also not-public-domain data across a whole suite of different modalities. And inevitably we’ll end up talking about multi-modality, but just to contextualize that a little bit, this involves everything from protein structure—for example through the PDB—and molecular structures of small molecules, omics data, knowledge graphs, biomedical literature, assay data from Patents, an enormous range of different types of data.

Ross Katz: We have these different modalities that you’ve just mentioned. Can you introduce us to Enchant as a model and how it works with the multi-modal data that you’ve just described?

Fred Manby: We all know that there’s been this enormous revolution in AI and that has come at the intersection of amazing algorithms like the transformer model, incredible GPUs like the GPUs made by Nvidia, and an internet full of data. In drug discovery, you have all those things except the internet full of data. That’s the clear difference between the general AI world and the specific AI in drug discovery world. One strategy for addressing that shortfall in terms of the quantity of available data is to widen what you mean by data. And that is another key motivation for driving hard on multi-modality. By building technologies that are multi-modal, you widen the definition of data, you increase the amount of available data that can be exploited, and that allows you to access larger model scales. For us, this has been a really important epiphany about how we’re going to be able to achieve model scales that really deliver something transformative for drug discovery. Enchant is our effort to do that. Enchant is a multi-modal transformer model, and it’s designed to be scalable with respect to the number of different modalities that are included in the training data. So far we’ve trained models with up to 15 different data modalities, but it’s expandable beyond that. As we encounter new data sources in different shapes, we can just bolt on a new modality and we’ll keep scaling that up.

Ross Katz: I want to dive into that a little bit because, as you mentioned multi-modality and the multi-parameter optimization space, as you have more modalities, the data becomes more and more sparse—you have certain data available about certain molecules going into the pipeline, and the model needs to account for the gaps in the modalities in order to get the most signal out of the model that you’re training. Can you help us understand how Enchant accomplishes that?

Fred Manby: You’re absolutely right, Ross. You’re inevitably going to have data sparsity. This only makes sense as a strategy if you can bridge over the sparsity in one modality due to insights that are extractable from the other modalities. You hear anecdotal examples of that in the general AI space—models that get better at writing sonnets by being trained on more music. There are lots of examples out there. But another thing to say about this is that in drug discovery, there are very specific ways in which you can support the models to make those translations between different modalities. Let me give you an example. You have three-dimensional structures and one of the three-dimensional molecular structures happens to be Aspirin. But then you have a paper which has text in it that’s about Aspirin. And then you have a table of data and one of the SMILES strings happens to be the SMILES string for Aspirin. You can help the whole translation between modalities by doing entity recognition and just marking up all of those three different things with a tag that says ‘this is Aspirin.’ And that supports the understanding in the model that insights are translatable about this specific entity between different modalities.

Ross Katz: How do you think about compressing the information that you have available into the multi-modal model in a way that makes sense to the model, lets it build on the unstructured stuff that’s out there?

Fred Manby: A fundamental challenge for all machine learning models for small molecule drug discovery is you never have enough data. Chemical space is so large and the data that you have on particular molecules is just a tiny, tiny sampling of the possible space of information that’s available there. You inevitably get into debates about the degree to which you can trust the model’s ability to interpolate and particularly extrapolate into new regions of chemical space. Those are all totally legitimate concerns and they’re real things—you do encounter this when you actually try and put these models to work. All I will say is that larger models—this is a very general statement, but broadly, larger models are better capable to do those interpolation and extrapolation tasks than smaller models. If you’re locked in a world of training machine learning models just on one particular assay property of molecules, let’s say you’re making a model to predict how quickly the drug is cleared by hepatocytes—some particular drug discovery related property—you might be lucky and have 2,000 data points. That’s great, but 2,000 is a lot smaller than 10 to the 60. So you’re asking a lot of that model and you do run into difficulties in extrapolability with models of that sort. By making enormous foundation models like Enchant—and again to re-emphasize, they’re enormous because multi-modality allows access to much wider resources of data—you can create models that actually do have demonstrably improved ability to extrapolate and interpolate beyond the training data.

Ross Katz: I know that one of the extrapolations and interpolations that you’re really excited about is in the clinical data space—the ability to bring clinical data into the drug discovery process and feed what clinical data you have into the Enchant models so that it can make that mapping between things like the structure of the molecule and the clinical outcome. Can you introduce us to how that works and what you’ve seen as a result of that?

Fred Manby: This for me has been really the most exciting aspect of the Enchant project. We all understand that when it comes to laboratory science, we can—collectively, humanity—build automated lab automation platforms that generate data, couple those to efficient machine learning models, and make a virtuous cycle out of that that is going to work and deliver efficiencies. I think many companies are proving that hypothesis. I think Iambic is a great exemplar of how to do that well. However, that is not a trick you can pull off in the clinical space. You can’t just scale the number of molecules that you put into humans because there are obvious ethical and regulatory and financial barriers to doing that. So you have to have a different conceptual framework for how to address the data challenge when it comes to predicting clinical properties. Multi-modal transformers, and in particular our Enchant model, are our answer to that challenge. What we’ve demonstrated through Enchant is the ability to make predictions of clinical endpoints—in this case human pharmacokinetics endpoints—where the quality of the prediction improves by being trained on more pre-clinical data, more laboratory data, data of the kind that you can scale production of. This feels like a very important breakthrough to us—it allows us to de-risk our discovery programs further into the development pipeline. And it also has another really important effect: if you’re a very mature big pharma company, you’re sitting on a huge repository of clinical data. It’s absolutely not a huge repository as enumerated by number of different molecules. But there’s a lot of data that’s hidden to a small company—Iambic’s four years old, we have one clinical program, we’re not sitting on an enormous repository of clinical data. But Enchant is undermining the value of that data moat because we can get better at predicting clinical outcomes without having more clinical data. That feels like a very, very important breakthrough for us and a thing that we’re very excited to keep pushing on.

Ross Katz: I just want to make sure I understand the mechanism by which that’s possible. My understanding from looking at some of the things that Iambic has written is that you’ve got these modalities and we’ve already talked about the idea that there’s this imputation, this interpolation that can happen—if you’re missing data about a particular modality, the model as part of its training objective is learning to fill in the blanks of the different modalities. As I understand it, clinical data is another modality there. Even though you don’t have very much of it, the model is learning to leverage as much information as it can from the clinical data that you have to then fill in the blanks of the clinical data that’s missing for the rest of the things that you’re training or inferring on. Am I thinking about that right, or how does it work?

Fred Manby: I think you are. Part of the answer to how does it work is that is an open research question. But when we first started getting into multi-modal transformers as a solution to this, we were very excited about—there’s a paper about the PaLM model from Google which demonstrates that large language models trained on enormous corpuses of text in many different languages but not including Persian suddenly emergently develop the ability to do Persian question and answers in a few-shot context. That is extremely interesting because it tells you that somehow there’s a framework in which these models can learn something more meta about the nature of language in general. My sense is that that is what’s happening here. We’re making models that are trained on enormous amounts of in vitro data from in vitro metabolism assays, for example, but also huge amounts of mouse PK data and other species PK data. Somehow what’s happening is the model is learning more generally about the types of PK properties of molecules and the types of variations in PK properties between species. And that is allowing a different order of extrapolation than would be otherwise achievable. In the blog post that we stuck on our website we show very meaningful predictive capacity for one particular human PK property by training on only data for five distinct molecules. Five examples is enough for that model to get a sense of how PK properties are varying between species and thereby make reasonable predictions of human pharmacokinetics just based on that minuscule amount of data.

Ross Katz: That makes a lot of sense. You mentioned earlier that you come from a chemistry background. Can you share a little bit about how the perspective that you bring to bear alters your paradigm of how modeling should occur in a context like Enchant, or how the outputs of the model should be leveraged and utilized inside an organization like Iambic?

Fred Manby: Having a background in chemistry can mean a lot of different things. In my particular case the area that I focused on was very theoretical—I was a quantum chemist. I worked on quantum mechanics and the creation of software that could use quantum mechanical calculations to predict things about molecules. I spent my 20-year academic career thinking about models that could be used to predict things about molecules. What we’re doing in Iambic is an extension of that where instead of physics-driven models we’re primarily focusing now on data-driven models. That’s one part of the answer. A different version of that answer is we have an incredible head of machine learning and he is able through his leadership to really drive a program of technology creation that encompasses a whole slew of different data-driven algorithms. I wanted to acknowledge that because a key part of our differentiation is having a wonderful drug discovery team—experienced drug hunters—and it takes a large group of people with different backgrounds to actually come together and create technologies and be able to leverage them in discovery programs.

Ross Katz: It sounds like you’ve got a head of machine learning who is developing a variety of models, some of which seem to be potentially more classical, more computationally heavy methods of evaluating quantum mechanics—using what we know about the math of chemistry and of physics—and then you’ve also got this outstanding drug discovery team. Can you talk a little bit about how those pieces fit together to create a drug discovery pipeline that is AI-supported and AI-driven?

Fred Manby: It doesn’t come for free. That means you have a project in the company that is a never-ending project—to work on that relationship and to understand the perspectives in both directions and to build technologies that actually answer the real questions. There’s a temptation to build a technology and then ask what problem does this solve. No, that’s not good enough. You’ve got to actually get into the detail and understand what is the challenge in this particular discovery program and how can we build technologies that address that specific challenge. I’m glad that it’s never ending because it’s actually a lot of fun. That’s an important part of the way to build a truly tech-enabled discovery operation—to work constantly on that inter-weaving of different disciplines.

Ross Katz: Can you help me understand what are the different types of problems that Enchant can help to solve through the fine-tuning lens—fine-tuning it to different tasks that you use internally?

Fred Manby: One of the powers of the Enchant framework is it can be fine-tuned to perform an extremely wide range of different tasks. In the conventional AI world, if you have a transformer that understands both images and text, you can make a chatbot out of that, you can make an image-to-image refinement tool, an image generator from a text prompt, a whole range of different technologies based on that common foundation model. The same is true for Enchant. But right now, the most important thing that we’re doing with Enchant is integrating it into our high-throughput experimental platform so that data coming off through weekly cycles of experimentation can be efficiently used to fine-tune models that then make high-fidelity predictions on those endpoints. We use LoRA for that—partly that’s just an efficiency thing—and we have set up fine-tuning on all of our key experimental endpoints happening automatically on a weekly cycle. That gives access to the latest models—as you come in on Monday you have access to the latest models that have been updated with all of the week’s data, and that is then used to think about and ideate around the next cycle of designs.

Ross Katz: The way that I think about that is that you have a limited throughput—experiments cost a decent amount of money. There’s a limit to the number of physical experiments that you can run, and you’re fine-tuning every week because every bit of data that you get out of the real-world experiments you’re running can be used to improve the model’s ability to predict the outcome of that same experiment. My understanding is that that increases the amount of molecules that you can evaluate with regard to that assay exponentially because you no longer have to run that experiment. Am I thinking about that right?

Fred Manby: Absolutely right. It’s a bit like walking around a dark woodland with a flashlight—you can see a bit of the land in front of you so you don’t bump into too many trees. The more powerful the flashlight, the further you can see. In some ways the framework of creating molecules, measuring properties, and training machine learning models has not changed. But by substituting out our old technology with Enchant, we’re able to see much further and make much more extrapolable predictions even in regimes where there are very few data points. Early in a discovery program, you have a very feeble flashlight because there’s hardly any data to inform what’s going on. Enchant improves that situation. You can see further and explore more adventurously in chemical space.

Ross Katz: Is Enchant in those cases giving you a bound of uncertainty or a confidence interval that helps you understand where you should be running experiments in order to give the model more information about the places where it’s uncertain?

Fred Manby: This is a great point, Ross. Uncertainty quantification is really critical to the way that we utilize machine learning models in the company. If you’re able to reliably read out an uncertainty measurement on any prediction, you can convert predictions to probabilities. You can say, I’ve got a distribution of predictions—what’s the probability that the compound has some activity that exceeds some threshold? This is wonderful because then that can be integrated into an all-encompassing decision-making framework. It’s not just that you do this for one property—you do it for 20 different properties because you’re trying to make a drug. There are going to be lots of different challenges to fix. And you can ask really data-informed questions like, should I do experiment A or should I do experiment B? Let’s calculate the probability that we’re going to get impactful results from those two experiments. And that works in the regime where you have well-calibrated uncertainty quantification. We assess that retrospectively—once we’ve done the experiments, we go back and scrutinize our uncertainty quantification predictions so that we can say whether it predicted the right probability for compounds to exceed some threshold. UQ is really a critical part of the way that we deploy machine learning models in general in the company.

Ross Katz: And I’m assuming it also guides which experiments you choose to run?

Fred Manby: It really does. If you have experiment A and it’s got a 5% chance of achieving a molecule that has some particular set of properties, and you have experiment B that has a 10% chance, it’s a no-brainer—you just do the one with the higher chance. We’re doing that as part of a fully data-driven decision-making framework throughout our programs.

Ross Katz: There’s a lot in the news right now about foundation models for biology, about large language models for biology. How would you compare Enchant to other biological foundation models that use this transformer-based approach in the space?

Fred Manby: As a company we haven’t focused on target selection, for example. Partly the reason is that the co-founders are both chemists—that’s the candid reason. But it’s also because once you’ve identified a target you’re a long way from generating value. That’s a key business reason why we really tightly focus on the challenge of coming up with the right molecule, because that’s the way to in the most tangible way drive value. But that means there’s a whole world of multi-modal biology-facing transformer models that are excellent for the tasks that they’re being deployed on. We scrutinize how we line up with competitor technologies in the molecular space. Right now it’s true to say that there’s nothing competing with the Enchant technology that is in the public domain. People have clearly had the thought—there’s all this text, there’s the SMILES strings, there’s loads of data, we must be able to do something. But the technical challenges of really executing well on that are such that until Enchant, nobody had really made a model that exceeded the state-of-the-art. Whereas Enchant is pretty systematically superior to the state-of-the-art over a very wide range of different molecular tasks.

Ross Katz: Can you talk a little bit about how you think about when new modalities should be brought in? That seems like a reasonably big decision, considering that training these models requires a good amount of investment and you’ve got real work that can be done with the version of the model that you have today. How do you think about that kind of update?

Fred Manby: It’s always a roadmap. It’s never that we trained a massive foundation model and now we’re going to fine-tune it for the rest of our lives. The field is evolving rapidly and the scale at which it’s feasible to train these models is increasing. A key bottleneck to doing this kind of work is just the data processing work—you have to have data scientists and software engineers building out the data pipelines to clean, do entity recognition, do modality extraction, figure out how to do tokenization. There’s an infrastructure layer that you have to build. That is what’s determining the rate at which you can add new modalities. That team in Iambic is a wonderful group of people and they’re working hard on that—that’s also part of our roadmap, enabling us to increase the different types of data and the different data sources that we use. If I had to add one modality right now, I think I would be excited to add image. The phenotypic response of cells to various perturbations, be they genetic or molecular, does provide an interesting different modality and different source of information. That might be an obvious choice for an additional modality that we don’t currently have but would love to include in the future.

Ross Katz: When we spoke earlier, one of the things that you mentioned was that you’re really passionate about thinking through how to design interfaces for experimental scientists to maximize the scientific value that you’re getting out of the data, the machine learning, the AI platform that you’ve developed. Can you talk a little bit about what you’ve learned about designing those interfaces?

Fred Manby: I’ve learned that it’s important. This is not a thing that I brought from my academic experience. But through building a team that could create the technologies to provide access to artificial intelligence to our research scientists, I’ve really learned quite a bit about the importance of that and the things that matter. As an academic scientist, you don’t spend a lot of time talking to UX designers. Design of interfaces is extremely important. It’s given me a deep respect for the back-end engineering that’s needed to make technologies that fulfill expectations. It’s the classic thing where you have a web app and somebody wants to add a button that does something—you provide that—but then they want a button that does that a million times. That is an enormous change of requirement and entails a huge technology lift on the back end. We have a great software group in Iambic and we love them and celebrate them, but sometimes that aspect goes a little bit under-celebrated—they’re the unsung heroes of the whole thing. It’s a huge undertaking to build platform technologies that can scale to fulfill requirements that give a fluid experience for scientists as they go through drug discovery activities. Front-end design is a critical element of that as well. We built a huge amount of tech to enable all of that kind of interaction and that’s an ongoing activity in the company.

Ross Katz: Can you give us any insight into what the workflow of the scientist is that you’re enabling with those interfaces and how the interfaces support that?

Fred Manby: Absolutely. The fundamental elements are—in small molecule drug discovery, the thing you want to be able to do is have a list of compounds. You might have generated that list using some technology or you might have drawn them with pen and paper and typed them in—there’s a whole range of ways that list might have come into existence. But that’s a fundamental unit: a list of compounds. And then you’re going to want to have access to all of the data that you’ve ever measured on those compounds. You’re going to want click-button access to any prediction that we can make. You’re going to want to be able to make statements about probabilities around those compounds. Having chosen an exciting set of compounds on the basis of those probabilities, you’re going to want to be able to trigger the experimental validation of those expectations. You have to be able to carry that list of compounds through to an ELN that allows you to generate robot instructions for executing that particular synthesis and later conducting those compounds through to the relevant assays. There are a lot of components that have to be built there. It’s been a great example of collaboration between technologists and laboratory drug discovery folks. You can’t just design this stuff—you have to sit down with people who are experts in the industry and really understand what has to happen, what the workflows are, what steps have to be taken account of. That has been a big collaboration. We’ve been doing this for a few years now and we have a really advanced suite of technologies that allow access to all of our AI technologies but also to our laboratory technologies as well.

Ross Katz: That makes a lot of sense. To use the puzzle analogy—when you’re doing a puzzle, you put all the pieces face up on the table so that you can see them, organize them, and understand where each piece fits. It sounds like the interface is about understanding what is the puzzle we’re trying to create, what are the different pieces that we would need in order to validate that this new drug candidate is something that we should pursue into the next step, and then understanding all of the different measurements, all of the different assays, all of the different aspects of each molecule that could lead you along that analysis path. As we draw toward a close, as someone who’s deeply involved in turning data in biotech into actionable insights, what excites you most about the future of data in biotech?

Fred Manby: My perspective on this—developed over the few years of working in Iambic—is that drug discovery itself is a massive act of data generation. It’s so exciting to be part of the realization that using AI to leverage the value of the data that we’re creating anyway, there’s just this enormous opportunity to improve capital efficiency. It’s great to have nice blog posts, it’s wonderful to have nice papers, but fundamentally what we’re doing here is getting better at leveraging the data that’s produced in drug discovery to reduce the capital cost of finding new medicines. There’s a lot of work to do, but that’s what excites me and keeps me going.

Ross Katz: Can you share any information about the progress that Iambic has made with your pipeline?

Fred Manby: As I mentioned, as soon as we had a lab we started discovery and our lead program is an oncology drug—it targets the HER2 oncogene. It’s an incredibly exciting compound. It not only hits the wild type of that gene but also all of the key cancer-driving mutations. It has exceptional selectivity, it’s brain penetrant, we’re super excited about that compound. We advanced that program from launch to IND filing in two years flat. It was a very fast process to start clinical science on that drug. Right now it’s in a Phase 1 trial. The rest of the pipeline covers a range of other targets in oncology—the next program is a dual CDK2/4 inhibitor program which we’re very excited about. The program after that is a kinesin motor protein, KIF18A, a really, really exciting and new target in the space of cancer therapeutics. And we’re branching out into other therapeutic areas as well. We have an undisclosed GPCR project and we also have a partnership with Lundbeck in a neurological indication.

Ross Katz: And where can people go to find more about the work you do at Iambic?

Fred Manby: The website is a great resource for all kinds of information—you can see our publications, our news stories and our blog posts there. I would highlight—we didn’t talk about NeuralPlexo—but just to highlight that we just very recently released a paper on our NeuralPlexo3 model which is a protein-ligand structure prediction technology that surpasses AlphaFold3 in performance on key metrics. We also released as part of that drop our benchmarking suite as an open source resource that people can use. That’s part of our effort to help the community more broadly to understand the performance of these prediction technologies, structure prediction technologies, and to put that onto a uniform playing field so that we can make reliable comparisons between models.

Ross Katz: Fred, thank you so much for joining. It’s been a pleasure talking with you and look forward to connecting down the line.

Fred Manby: Really enjoyed it, Ross. Thank you so much for the conversation.

Ross Katz: Take care.

Jason: And that’s it for this episode of Data in Biotech. If you enjoyed the episode please subscribe, rate, or leave a review in your podcast platform of choice. See you next time.

Frequently Asked
Questions

How does Iambic's approach improve drug development speed and capital efficiency?
By integrating high-throughput experimental data with AI models like Enchant, Iambic accelerates lead optimization and identifies promising drug candidates faster. The model's ability to predict clinical outcomes from preclinical data de-risks programs early, reducing costly late-stage failures and shortening the path to IND filing, as seen with their 2-year IND achievement for a lead oncology program.
What infrastructure supports a multimodal AI model for drug discovery?
Supporting a multimodal model requires robust data pipelines for cleaning, entity recognition, modality extraction, and tokenization across diverse data types (e.g., protein structures, omics, literature, assay data). This infrastructure is key to continuously integrating new data sources and scaling model training, requiring significant software engineering and data science investment.
How does Iambic prioritize experiments to maximize discovery impact?
Iambic uses uncertainty quantification to convert model predictions into probabilities for various molecular properties. This data-driven framework allows them to compare potential experiments and select those with the highest probability of yielding impactful results, ensuring efficient allocation of laboratory resources and focused progression of drug candidates.

Need a data partner for life sciences?

CorrDyn helps biotech and pharma companies build the data infrastructure that accelerates research and operations.

Book an intro call