Listen on
Overview
Drug discovery is slow and expensive, often hindered by the sheer volume of data needed to train effective AI models—a resource smaller biotechs rarely possess. Eli Lilly’s TuneLab directly addresses this by offering a novel “give-get” model: external companies gain access to Lilly’s proprietary AI models, trained on decades of internal data and representing over $1 billion in development. In return, these partners contribute their own data via federated learning, improving model generalizability without compromising data privacy or IP.
This approach accelerates preclinical drug development for biotechs, reducing the need for extensive wet lab experimentation. For Lilly, it broadens the applicability and performance of their models across diverse chemical spaces. Dr. Aliza Apple, VP of Catalyze360 AI/ML and Global Head of TuneLab at Eli Lilly, brings deep expertise as a biomedical engineer and former McKinsey consultant to explain the platform’s architecture, its current applications in small molecule AdMet and antibody developability, and its potential to reshape the biotech landscape by making advanced AI tools accessible.
Ross Katz discusses with Dr. Apple the technical intricacies of federated learning, the critical role of third-party security, and how TuneLab is poised to drive the next generation of in vivo predictive models. The conversation reveals a strategic shift in data collaboration, demonstrating how shared infrastructure can overcome individual data limitations to advance collective scientific goals.
Key Takeaways
Federated Learning Overcomes Biotech’s Data Scarcity
Many biotechs lack the extensive historical data needed to train reliable AI models for drug discovery. TuneLab solves this by providing access to Lilly’s proprietary models, while participants contribute their data via federated learning. This allows models to improve on a wider dataset without any company’s raw data leaving its secure environment, offering a practical path to advanced AI for smaller firms.
Model Generalizability, Not Just Accuracy, Drives Value
While Lilly’s models are already accurate in known chemical space, TuneLab’s primary goal is to enhance model generalizability. By incorporating diverse data from biotech partners, the models perform better in novel or underexplored chemical spaces. This broadens their utility for all participants, enabling more effective predictions for new therapeutic programs.
Third-Party Hosted Federated Learning Builds Trust
The success of a collaborative platform like TuneLab depends on absolute data security and privacy. Lilly intentionally uses a third-party host for its federated computing environment, ensuring that participant data remains under their control and is never directly accessed by Lilly. This architectural decision is crucial for earning the trust required for sensitive data contributions.
In Vivo Predictive Models are the Next Frontier in Discovery
TuneLab is actively developing in vivo predictive models, aiming to accurately forecast a molecule’s behavior in an animal from a simple chemical structure. This capability, still at the frontier even for large pharma, promises to accelerate discovery, drastically reduce animal testing, and cut costs. TuneLab’s collective data model is designed to overcome the significant data scarcity challenge inherent in building such complex models.
Related: CorrDyn provides AI strategy and data engineering expertise to clients in biotech and life sciences. Learn how to help biotech manufacturers access data value.
Full Transcript
Jason: Welcome to Data and Biotech, a podcast from CorrDyn where we explore how companies leverage data to drive innovation in life sciences. Every 2 weeks, we sit down with an expert from the world of biotechnology to understand how they’re using data science to solve technical challenges, streamline operations, and further innovation in their business. Here we go.
Ross Katz: Aliza Apple, welcome to the Data and Biotech podcast.
Aliza Apple: Thanks for having me.
Ross Katz: To kick us off, would you mind introducing us to you and your background, and what brings you here today?
Aliza Apple: Aliza Apple, I’m the VP of Catalyze 360 AI, and I’m now allowed to share that I’m also the global head of TuneLab, a new offering that we’ve just launched. I’m a scientist by training. I trained as a biomedical engineer at UCSF and UC Berkeley, postdoc’d at Cornell Medical College, spent about a decade at McKinsey, and a couple stints in biotech before joining Lilly.
Ross Katz: I know we’re here to talk about TuneLab today, but before we get into TuneLab, would you talk about, in the context of Lilly as an organization, what Catalyze 360 is and what it encompasses?
Aliza Apple: Lilly Catalyze 360 is rooted in the goal of finding the most exciting science wherever it lives. We focus on supporting the associated biotechs by meeting them where they are. We’ve got 4 major offerings that we call pillars. The first is Lilly Ventures, where we make investments focused on novel areas of science at their earliest stages. The second is Lilly Gateway Labs, our innovation hubs where we offer state-of-the-art lab facilities and scientific engagement with the companies themselves. The third is Lilly Explore R&D, where we offer R&D services all the way from discovery through clinical proof of concept. The fourth is our newly baptized Lilly TuneLab, our collaborative AI/ML platform for drug discovery. We’re focusing on providing a 360 suite of offerings. That’s where the 360 comes from. We get that question a lot. It’s not 360 companies; it’s 360 degrees.
Ross Katz: I should think of Catalyze 360 as external innovation, how Lilly participates in the external biotech environment to foster innovation so that the organization can stay abreast of where innovation applications of biotech are coming from. Am I thinking about that right?
Aliza Apple: That’s exactly right. Many of our peers call a comparable organization their external innovation team.
Ross Katz: We’re here to talk about TuneLab, which you just launched. Will you give us an overview of TuneLab and what it offers?
Aliza Apple: TuneLab is a first-of-its-kind collaborative AI drug discovery platform with a unique give-get value proposition. We give biotech companies access to powerful Lilly-trained proprietary models to accelerate their own therapeutic programs. These are the same models that Lilly scientists use every day, trained on decades’ worth of our own historical data. In terms of what we get, in return for access, our participants contribute training data via federated learning, which I’m happy to share more about. This flywheel fuels continuous model improvement while maintaining full data privacy and security for all participants. To be clear, that is all that we get. Companies retain full control of their proprietary information, ownership of their IP. It’s a secure way of mutually benefiting the models and driving science forward together.
Ross Katz: That’s awesome. I’ve never heard of anything like this before. As I understand it, Lilly is putting up its data accumulated over decades and the models trained on that data, offering the opportunity for these external companies to utilize those models for their own purposes. What they contribute to participate is their own internal data, which gets incorporated into the overall system using federated learning. Am I thinking about that right?
Aliza Apple: That’s exactly right. If you think about this longitudinally, once you’ve identified a target and started generating hits in preclinical drug discovery, these models focus on the phase from hit to lead to optimization, all the way to development candidate nomination. In the old world, pre-modeling – think back 15 years – you would have to synthesize all molecules, test them in the wet lab, then look at all those results and decide what to move forward. Here, we’re making models available that companies can use much higher upstream in their program development to predict how a molecule would likely perform before they synthesize it, and certainly before they do any wet lab work. This is a way of making the scientific discovery paradigm faster and cheaper. Pharma has been working this way for many years; we’ve had models that look like this as far back as 10 years. They’ve improved dramatically over that timeframe, but most biotechs lack the historical data, particularly in these phases of development, to train their own models. What we’re offering them is day one access to extremely high-performing models that have cost us over $1 billion to generate – I kid you not, it’s with a B. The idea is that as we collectively contribute to the model’s improvements, both in performance and generalizability, everyone wins.
Ross Katz: Can you give us an idea of what models are being exposed and the different purposes to which they’re being applied? Also, a little bit about the datasets that feed those models and the type of data that external biotech organizations are expected to bring to the table to utilize them?
Aliza Apple: At launch, we’re focused on two major use cases, firmly in that sweet spot of post-hit and leading up to development candidate nomination. The first is small molecule AdMet. The second is antibody developability. On the small molecule AdMet side, we’re making our full suite of in vitro proprietary property predictive models available across absorption, distribution, metabolism, excretion, and toxicity. These models have over 500,000 unique molecules in the training set, collected over several decades. I’m excited to share that we’ll be making in vivo predictive models available. They’re actively being built in a collaboration announced in parallel to our launch with insitro, a company led by the prolific CEO and founder, Daphne Koller. These in vivo models will predict species-specific in vivo PK endpoints and 4-day rodent safety endpoints, plus some additional in vitro safety pharmacology endpoints, expanding that suite and building these models for the first time, making them available to Lilly and TuneLab participants simultaneously. On the antibody developability front, we’re focused on 6 of the ‘horsemen of the apocalypse’ for antibody programs: aggregation risk, thermostability, nonspecific binding, solubility, chemical liability and deamidation (which assesses the likelihood of a spontaneous chemical change in an antibody that can affect its potency and stability), and viscosity. We’re looking to make a fast follow to add models to both of those suites as we go.
Ross Katz: You’ve got these two classes of models that you’re making available. These are the two main use cases that TuneLab is focusing on in the immediate term. Are there external innovation partners or companies doing work in the biotech ecosystem that you think will benefit particularly from these – companies at a specific phase of discovery development or operating in certain therapeutic categories? How would you classify the segments of the ecosystem that would be most interested in the models you’re exposing?
Aliza Apple: It is designed to be of use to any company working on preclinical programs. It’s more of a modality-focused criteria: small molecule and antibody as a starting point. We have the ambition, as we expand both the small molecule offering and the antibody offering, to add additional modalities. Outside of the modality focus and the preclinical stage more broadly, these models are completely therapeutic area agnostic. We have been focused on bringing on board our early-stage biotech partners. I can tell you more about who’s already working with us, but in general, we would welcome participation from any company aligned philosophically with the give-get model and that could make use of these models in their own preclinical programs.
Ross Katz: I would love it if you would share some of the companies that are already working with you. That would be great. Then some information about what they needed to bring to the table to participate, and what that process of joining up and integrating looks like.
Aliza Apple: Ahead of the public launch, we created an early access program in stealth mode, and we’ve had over a dozen companies already join us and participate. We’re particularly grateful to Firefly Bio, Superluminal, Sysonq, and Interdict. They have all been helpful in these formative stages, giving us feedback, helping us understand the user experience. Supraluminal even conducted some of our first federated learning with participant data. We’ve had active participation from some of these early companies that we know well and are very grateful to. We anticipate a large ramp-up over the first year. The idea is to go from double-digit users to triple-digit users as quickly as the industry is excited to get on board.
Ross Katz: What were some of the user experience challenges that your early partners faced that you had to resolve as you were developing this federated learning platform for them?
Aliza Apple: Naturally with any SaaS platform, getting user feedback and making sure everything is intuitive, well explained, and even collecting those frequently asked questions, has been tremendously helpful. One area where I’ve been pleasantly surprised is the reception of the concept of federated learning. While there’s certainly an important quick educational piece that we fully anticipated, we expected more skepticism around the concept than what we’ve actually experienced. It’s very encouraging. The good news is most of the folks we’re working with are data-driven scientific individuals who need to kick the tires. Once they understand the concept and can validate it for themselves, they can get over what is probably the biggest mental hurdle historically for something like this to come to life: the contribution and ensuring it does protect their IP. They can move forward with their programs with full confidence that their data is secure. That’s been incredibly powerful.
Ross Katz: That’s really interesting. It makes sense that from the perspective of external partners, the fact that Lilly is putting up this vast quantity of data that was expensive to produce, and that you’ve selected a third party that meets your internal security standards, means you’re putting a stake in the ground that all this architecture is of the highest possible quality—so much so that even an organization like Lilly is willing to participate in the sharing. Were there any hard questions you got from external partners about Rhino Federated Computing or about the federated learning framework that required a lot of research or providing a lot of information? I’m interested in how those conversations go about the sharing of data back into a system like this.
Aliza Apple: Some of it is understanding the architecture. We’re always happy to walk through that, and we put our participants in direct contact with Rhino, where they have a contractual agreement for their data security. We’ve been very deliberate about this not being a server that Lilly hosts; it is deliberately hosted by a third party. Most technical questions we get focus on probably two major things. One is, “How do we know the model is getting better?” Simply put, one of the key performance indicators is typically to rerun the test set. While we have advanced machine learning scientists on our team who have developed many ways to do this, our ultimately simple yet elegant answer is to actively measure the model’s performance before and after training, and commit to only making the best-performing model available as the global model. Should training result in a decrease in performance, that is simply not the global model we will serve. One of the key performance indicators we are most excited about, in terms of how the model is getting better, is the improvement over the chemical space. Candidly, our scientists, especially in the two use cases we’ve started with, are generally happy with the performance of many of our models. We don’t expect a lot of improvement in accuracy from the TuneLab community. Where we’re most excited, however, is in the generalizability of the model, because the model performs best in chemical space it has seen. The opportunity from our vantage point, which we think is genuinely exciting, is to make the models more generalizable to chemical space not contemplated in our historical data. Everyone will benefit from this collective broadening. That’s the key from our vantage point: less about accuracy and more about generalizability. The second piece, which we are still actively working on, is finding more simple and elegant ways to quantify the uncertainty of the underlying model predictions. As you can imagine, our partners are interested in how much they can rely on that data for decision-making, especially if this is going to be as far upstream as “what do I synthesize.” We’re actively exploring ways to apply techniques called conformal learning, which would help measure the confidence. This is a bit trickier to do in a federated learning context, so it is something we’re actively experimenting with but hope to roll out soon.
Ross Katz: Those questions make a lot of sense. It brings up another question: if external third parties are building on the high-quality dataset you all have developed internally, how do you determine, curate, or evaluate the quality of the data being brought to you in return for participating in the model? I could imagine there might be certain datasets where, no matter how many model runs you try, their addition won’t lead to improvements in the model’s quality.
Aliza Apple: We employ a number of different strategies, including data harmonization aspects. This involves understanding how the data is generated, which protocol is used, and better understanding differences that may arise from those pieces. But again, we have to come back to what the test set tells us before and after. It’s the most simple ultimate north star for how the model is performing, and it’s our primary metric.
Ross Katz: Early on, is there a way to determine whether the data being brought is likely of value to the model? Or if you’re a biotech with an existing dataset, how do you determine if your developed data is worthy of being incorporated into the modeling environment?
Aliza Apple: It’s a great question. In our early days, at launch right now, we are being a bit more relaxed about experimenting and better understanding how much data harmonization we can tolerate and still see gains in model performance. Where we envision quickly advancing with additional partners is creating a data generation “easy button,” so to speak. We actively share our protocols with participants should they choose to generate their own data, and they can use those protocols for that effort. We can also start to bring on board CRO partners, and there are advantages in thinking about how we can all select the same catalog number to generate the same data using the same assay. Those are pieces we’re actively building as we go. Our test set, while we can continue to evolve it and test it with historical versions of the model, it will be hard to capture the breadth of the experience expanding chemical space. As is true to the model, we don’t know the chemical space being added on in a way we could incorporate into a test set. A company might tell us they’re a macrocycle company, but that doesn’t really help us think about how to change the test set. We encourage participants with significant datasets to hold out the traditional 10 to 20% to create their own test set. We also offer fine-tuning for those companies to take the global model, tune it, and then apply it to their specific chemical space. Across a number of those different arenas, should they have enough data to hold out a test set, we also encourage them to share the accuracy results with us. This is something we’ll learn more about over time, and it will be heavily informed by whether we get very early-stage small biotechs and how often they are able to contribute a meaningful amount of data while still holding out a test set.
Ross Katz: That’s really interesting. It also brought up another question for me. You’re creating, in essence, foundation models that can then be fine-tuned. The models you’re developing, every model is a simplification of the world. It’s based on domain knowledge and your understanding of the space. I’m curious how much information you provide about the architecture of the models you’re creating, the different assumptions in it, the way you think about the chemical space you’re operating in?
Aliza Apple: We share detailed model catalogs with all our participants so they can understand the model architecture, specifically the breadth of the training set, and the current, most up-to-date global model performance. We share all those parameters openly, as well as the protocols used to generate the data.
Ross Katz: You mentioned earlier that there are these in vivo predictive models, and that this is something unique you’re doing here. I’m interested in your perspective on what makes building these in vivo predictive models difficult, and what makes this offering different from what else is out there in the open source or proprietary ecosystems?
Aliza Apple: This is cutting edge even for us, exciting territory. We couldn’t have found a better partner than the insitro team, who have deep expertise in machine learning, creativity, and biotech scrappiness in their approach. It’s been an exciting collaboration to kick off. To answer your question from a 10,000-foot view, imagine a world where you can input a SMILES string and accurately predict what that molecule might do in an animal. That would be truly game-changing, both in speed and in a significant reduction in the use of animals in our scientific development. We are excited about that vision and achieving that goal. As we all know, that’s also been a widely publicized macro goal for our industry. The biggest challenges in getting there are simply having enough training data. It’s that simple and that hard, especially when you think about larger species of animals. Some of the models we’ll be building here will incorporate larger species as well. We’re thinking creatively and thoughtfully about how we can train a model across the breadth of our in vitro datasets and the various species of in vivo datasets we have, to create as accurate a model as we possibly can from the outset. Using a forum like TuneLab for other companies to contribute that data, we firmly believe, will be the fastest way to get there. We’re excited to get the train on the tracks and see what it can do.
Ross Katz: It strikes me that this is an inversion of how the marketplace works right now, where part of a biotech organization’s story is building up that dataset to have a proprietary lens on whatever modality or therapeutic area it operates in. How do you think TuneLab has the potential to change how the biotech ecosystem works, or how early-stage biotech organizations grow and evolve?
Aliza Apple: It’s a great question. It was at the core of which use cases we thought about tackling, especially what we wanted to tackle first. I work with a brilliant woman named Nisha Nanda, who heads up all of Catalyze 360. She had been mulling over this twofold question: within Lilly Ventures, Explore R&D, and Lilly Gateway Labs, we see many companies with incredible differentiation in certain areas of their discovery program, but few have thought about actually developing a molecule. This is naturally not where companies of this size would have built up a dataset. There are tried and true methods, and big players like Lilly who have the historical data. There was an interesting paradigm shift in her thinking about how we can both accelerate these companies, and also use the collective and Lilly as an underpinning and convener to realize the promise of what all this data, if leveraged together, could do. I can tell you a funny story about how this whole thing got started, but essentially that was at the crux of it: what is a classic biotech’s true right to win from a differentiation perspective? We can all agree that AdMet is probably not it. Thinking about how we can make those tools available and create this give-get model stemmed from that idea. They’re going to have exciting novel biology insights or very interesting specific chemistry insights, but unlikely toxicity insights.
Ross Katz: For sure. If there’s time, I would love to hear the whole story underlying that. The other question that came to mind as you were talking is, what are the biggest benefits to Lilly as an organization from creating TuneLab and exposing this federated learning model built on top of your data to the ecosystem at large?
Aliza Apple: Simply, it comes back to our mission as a whole for Catalyze 360: ensuring that the best science sees the light of day and has a fighting chance to make the impact it could have. From our vantage point, we certainly think that with the give-get model, we’re excited about what we can get in terms of the generalizability of the models. But we’re equally excited to give because truly, seeing the biotech industry – which, let’s be honest, has faced a difficult time – and being able to give back and do our part in ensuring that companies can realize some of these efficiencies and come through this difficult time with the promise of AI in its tailwinds, would be a noble goal that we’d be excited to have played a small part in.
Ross Katz: There are two questions I think about when I look at a leading biotech pharma organization like Lilly doing this. First, would any of the other large pharma organizations add their data to TuneLab? Would there be a benefit to doing that? Or would they look at what you’re doing with TuneLab as an example of what they can be doing? I’m curious what you see in the landscape of large pharma, what you might expect to see.
Aliza Apple: I would be excited for any of our peers to be interested in participating. We built TuneLab with a focus on our biotech ecosystem and those partners, but nothing would preclude any larger potential partners from participating. We would certainly be open to that. Whether they will follow suit or not, I am in no position to predict what their inner workings might look like. Our peers have not shown signs of wanting to externalize their models. This has been an exciting but not a simple undertaking, and I’ll be curious to see how the broader industry reacts.
Ross Katz: How do you, Nisha, and Lilly as a whole evaluate the success of TuneLab? What would you hope to see from TuneLab over the next year or 5 years to determine whether it’s accomplishing the organization’s goals?
Aliza Apple: From Lilly’s vantage point, there’s the quantitative element of how many biotechs we’ve helped advance, and how many successful drug candidates have entered the clinic as a result. Those are some of the metrics we’re actively thinking about and would be excited to collect and share. We also want to ensure we continue to evolve the offering and actively learn from collaborations such as the one we’re building with insitro. These are concrete goals where we want to bring cutting-edge models to all of TuneLab, and I would include Lilly in that as well.
Ross Katz: As we head toward the end of our time together and zoom out a bit, I’m interested in what TuneLab says about how the emergence of ML and AI changes the drug discovery and development process in the years to come?
Aliza Apple: We’re most excited about seeing the small molecule and antibody programs advance at faster rates, both in how many molecules need to be synthesized and how many experiments need to be run. As these in vivo models come online, the ability to rapidly accelerate those programs is certainly something we’re keeping a close eye on. Looking further afield, as I mentioned briefly, we’ll also be interested to see what we can unlock across other modalities and continue to expand the breadth of the offering.
Ross Katz: As you were talking, one of the things that occurred to me is that as new datasets are added here, there’s an opportunity for Lilly and all your participating organizations to see the relationship between the data introduced to these models and the quality of the results. There’s a lot of talk about scaling curves as it pertains to language models, but these alternative foundation models you’re creating will have scaling curves of their own. There’s an opportunity for all these organizations to hitch a ride on the curve and see how the new data leads to even better outcomes for each participant.
Aliza Apple: That’s exactly right. As we start to see further advances in de novo design, thinking about moving from inferencing as a service to workflows as a service, and pushing into coupling de novo design with these property predictive models, you can imagine rapidly getting from an idea for a target to a molecule that will work in an animal, before you’ve ever done anything in a lab, which is exciting.
Ross Katz: As you look toward the next few years, are there any other hurdles on the horizon for TuneLab that you think we’ll know we’re starting to make it when we overcome a particular type of challenge?
Aliza Apple: Great question. At launch, we’re still humble, so I’m focused on moving from unknown to known, ensuring we transition from skepticism around data contribution to a clear understanding that federated learning is safe. Those are the two biggest pieces we’ll be focusing on in the immediate term. This collaboration with insitro is an exciting mandate for TuneLab to continue pushing the envelope, to figure out where cutting-edge modeling can truly serve the ecosystem. How we can finally leverage the collective to create the right training set to make it possible. Continuing to do a couple of those a year will hopefully move the needle fast.
Ross Katz: As we head toward the end, where should interested biotech organizations start if they want to explore how TuneLab works and how they can participate?
Aliza Apple: Please come visit us. Our website is lillytunelab.com, and there’s a contact us page where you can easily get in touch with our team. We’d love to tell you more about it.
Ross Katz: Aliza, it’s been an absolute pleasure to have you on. The model that TuneLab is establishing strikes me as one that the entire biotech, pharma, and healthcare ecosystem can learn from, because many of the problems we have in resolving health issues, which are a public issue at large, involve getting all the data needed to gain complete visibility. Thank you so much for joining and sharing the work you’re doing at Lilly.
Aliza Apple: Thank you. Thanks so much for having me. I appreciate it.
Jason: That’s it for this episode of Data in Biotech. If you enjoyed the episode, please subscribe, rate, or leave a review in your podcast platform of choice. See you next time.






