Skip to content
Robert Abel — Physics, Free Energy, and Computational Drug Discovery
Data in BiotechEpisode 68

Physics, Free Energy, and Computational Drug Discovery

Robert Abel of Schrödinger on why ML alone fails in 10^60 chemical space and how physics-based simulation reaches near-experimental accuracy.

57:31Full transcript below
RA

Robert Abel

Chief Scientific Officer, Platform at Schrödinger

Overview

Developing new drugs is slow and expensive. The potential chemical space for drug molecules is astronomical—roughly 10^60 options—yet traditional lab testing explores only a minuscule fraction. Machine learning models, when trained on such sparse real-world data, struggle to extrapolate effectively, leading to unreliable predictions and costly dead ends. This fundamental challenge directly impacts R&D ROI and competitive advantage for biotech and pharma companies.

Robert Abel, Chief Scientific Officer Platform at Schrödinger, has spent over 15 years developing computational methods that redefine drug discovery. He explains how Schrödinger’s hybrid approach combines rigorous physics-based simulations with targeted machine learning to accelerate hit identification and lead optimization. This episode details how their platform delivers high-fidelity predictions, drastically reduces project timelines, and increases the reliability of drug candidates before expensive wet lab synthesis. We also explore their unique business model and user-centric software design that makes advanced computational tools accessible to a broader range of scientific users.

Key Takeaways

Physics-Based Simulations Generate Actionable Data Where Lab Work Cannot.

The chemical space for drug discovery is too vast for machine learning models to effectively learn from the limited existing experimental data. Schrödinger addresses this by using computationally intensive, physics-based free energy calculations (FEP+) to generate high-quality, project-specific “virtual data.” This allows ML models to triage billions of molecules in relevant sub-regions of chemical space, significantly accelerating the discovery process from months to overnight.

Quantitative Accuracy is a Decisive Factor in Computational Drug Discovery.

Schrödinger prioritizes the accuracy and reliability of its computational methods. Their free energy calculation technologies achieve an average root mean square error of 1.2 kilocalories per mole in comprehensive testing, approaching the intrinsic noise level of experimental measurements (0.9 kilocalories per mole). This precision allows their platform to consistently prioritize drug candidates with high confidence, reducing costly failures in later development stages.

A Three-Pillar Business Model Strengthens Technology and User Value.

Schrödinger operates with software licensing, collaborative partnerships, and proprietary clinical programs. This structure creates a powerful feedback loop: challenges in internal drug discovery programs drive software improvements, and software customers’ diverse needs ensure broad applicability and robust user interfaces. This constant validation and refinement across multiple contexts leads to a more reliable and broadly effective technology platform for all users.

Making Advanced Tools Accessible Accelerates Iteration and Adoption.

Complex computational tools are often inaccessible to critical users like medicinal chemists. Schrödinger’s Live Design system allows computational experts to embed models with predefined guard rails, letting chemists run sophisticated analyses with a single click. This ease of use fosters rapid in-silico iteration, making advanced drug discovery capabilities directly available to those who need them most without requiring deep computational expertise.

Related: CorrDyn helps companies in biotech and life sciences build robust data engineering platforms. We provide data quality and data reliability services, ensuring the high-fidelity computational pipelines necessary for accelerated drug discovery.

Full Transcript

Jason: Welcome to Data in Biotech, a podcast from CorrDyn where we explore how companies leverage data to drive innovation in life sciences. Every two weeks, we sit down with an expert from the world of biotechnology to understand how they’re using data science to solve technical challenges, streamline operations, and further innovation in their business. Here we go.

Ross Katz: Robert Abel, welcome to the Data in Biotech podcast.

Robert Abel: Well, thanks so much for having me.

Ross Katz: Awesome. Well, just to kick us off, would you mind walking us through your career and what brought you to Schrodinger?

Robert Abel: Sure. Trained as chemical physicist, co-developed a technology that eventually became the WaterMap technology over the course of my PhD. Then, as I was completing my PhD work, Schrodinger expressed an interest in licensing that technology. The last year of my PhD, I ended up collaborating very closely with the company on the productization of that technology, and that led to me joining the company in 2009 to further develop that technology. Post-joining the company, I ended up moving through a number of different roles until eventually I became chief scientific officer platform of the company, and now oversee all the computational methods development at Schrodinger.

Ross Katz: Fantastic. Excited to talk to you about Schrodinger and about the different computational methods. You’ve been at Schrodinger through the development of all these versions of the tooling that you’ve come up with. I would love to hear your story at the organization and how it aligns with how the tooling has developed over time, going from lead WaterMap developer to now chief scientific officer of the platform, and just hearing how the tools have evolved over time, but also how your role in those tools entered into that.

Robert Abel: Sure. One of the things that motivated me to join the company as the last year I was doing my PhD was working on the productization. I was hearing rumors sometimes about what was called at the time Project Troubled Waters, and no one would explain to me exactly what that was. I got some sense that there was an interest in using the technology to fully advance drug discovery. I was interested to learn more. I ended up joining the company, and Project Troubled Waters ended up being the code name for Nimbus Discovery, Nimbus Pharmaceuticals, a joint venture that eventually became a quite successful biotech company. That gave me my first opportunity to start to contribute directly to the use of the technologies to advance drug discovery programs. I learned a whole bunch supervising the use of or giving input to the use of WaterMap on some of those programs. That also led me to have an improving understanding of how the technology needed to improve in various ways to further impact drug discovery. I found Schrodinger to be an environment where any time any employee had an interesting idea regarding how to further improve the technology, there was a lot of license at the company and a lot of willingness on the part of all the employees to work together and see if those ideas had merit and could be developed in a way that would further improve the technology stack such that it could have greater applicability to various parts of the drug discovery process. For instance, I think it was within the first year after I joined, I started to see some interesting opportunities to contribute to force field development at the company and was able to make some contributions to improve the protein force fields. That also allowed me to build some personal connections with the force field development group and have some ongoing interactions to further help out there. WaterMap uses a molecular dynamics technology, so I also got very involved with some of the MD development and eventually the free energy calculation methods development. From 2009 to roughly around 2012, it became apparent we needed greater quantitative accuracy and rigor regarding our ability to predict potency accurately and reliably. That led to the development of the FEP+ technology, which is our free energy calculation technology. This was a major step forward in terms of our ability to calculate affinities — drug-ligand binding, protein-ligand binding affinities — with high quantitative accuracy. That improvement was led in many ways by other force field development work we had been pursuing along the way, because you need a force field with broad and accurate coverage of all chemical space before developing technology like FEP+ would actually be tractable. That was a major step along the way. Another important step was the improvement in our protein structure prediction and protein ligand binding mode refinement algorithms. For example, IFD-MD greatly expanded the scope of systems we could effectively model with free energy calculations. Another important development, of course, was all the innovations that have taken place in the machine learning community. Now, rather than needing to calculate all properties with explicit physics-based calculations, which can be computationally very intensive, we can use those physics-based calculations to generate virtual data to train machine learning models. That allows us to operate at a much greater scale and throughput than would be possible if we only used free energy calculations to guide what we were doing. Some major moments, if you want me to try to put a year on every technology: WaterMap would have been around 2007-2009 was the arc of when it was fully productized. Free energy calculations (FEP+) would have been between 2012 and 2014-15, depending on the continuum where the technologies are further improving. IFD-MD, which is our induced fit docking where you can dock a ligand into a flexible protein receptor, saw major innovations between 2017 and 2020. Machine learning innovations probably would have been fully baked and integrated between 2018 and 2022. There’s a lot of things we’re actively working on now concerning full integration of generative machine learning methods and Generative AI. The story of technology development just goes on and on. Lots of stuff we’re working on right now and hoping to keep that momentum going in the future. There’s various important technologies I probably missed in that overview. I just thought of major contributions made to machine learning-based interaction potentials—machine learning-based force fields—and those have been a major story roughly within the last three to five years. Again, I’m missing a bunch of technologies I’m going to think of, I’m sure in the next five minutes. But those are just a few that were really important that happened along the way.

Ross Katz: It just strikes me that you’re coming from this background in chemical physics and you’re coming into the drug discovery world. That physics-based view on the world gives you a unique approach and a unique thought process that is also embedded in Schrodinger’s approach, as I understand it, to drug discovery and to software development for drug discovery. I would love to hear from you how that lens that you bring, in your opinion, differs from the way that other companies in the space might approach drug discovery?

Robert Abel: It’s an interesting question. Other companies would really depend, of course, on the specific company. What I can say is where Schrodinger tends to really focus is whether we can approach something accurately and rigorously with a full understanding of the molecular detail of the molecular process that’s ongoing. The challenge of running atomistic simulations is something we run toward, where other companies that don’t come from that background would probably be more hesitant to try to break new ground in enhancing sampling methods or doing ultra large-scale simulations. That’s definitely our background and our approach, and we like to take that approach whenever we think it’s going to be tractable. It also probably leads to us avoiding certain problems as things we don’t feel are completely tractable at the present state of the technology to take on from a position of confidence. For example, for the benefit of our drug discovery projects and our software users, we’d love to develop a highly reliable model to predict rate of clearance of drug molecules. That’s a really tricky kinetic problem mediated by a great many enzymes. It’s a very complicated biological process by which drug molecules are cleared. That’s an example where a fully atomistic approach right now may not be feasible. That’s where we might be a little more hesitant to dive into model development there. If you come at it from a computer science pure machine learning background, it’s an input and output; you should be able to build models. You can attempt to do that, but we would have concerns about the reliability of those models when they’re used prospectively.

Ross Katz: I want to touch on reliability when we get into the physics-based methods a little later on in our conversation, but before we do, I want to take a little bit of a detour into Schrodinger as a company. You’re operating, as I understand it, across three different business models: software licensing, collaborative partnerships, and proprietary clinical programs. This strikes me as an interesting model for an organization to take. From your perspective, how do these different business models feed back into each other? How do they synergize to make each leg stronger?

Robert Abel: I view us more than anything else as a technology company. The software business, the collaborative programs business, and the internal proprietary programs business are just different ways of monetizing and realizing value of that core technology asset. What is really exciting about having all three of those lines of business operating in parallel is they benefit each other because they all lead to different systematic improvements in the underlying technology that can then be used by the other line of business. For instance, if one of our collaborative or proprietary programs runs into a scientific challenge, we’re very likely to have something on the shelf that can be brought to bear to address that problem because we’ve had to support all our software users running into all sorts of different scientific challenges along the way. Likewise, because the software is used directly by us to attempt to advance collaborative drug discovery projects or internal drug discovery projects, we can methodologically improve the underlying technology based on those experiences in a way where our software customers benefit because those improvements are later on distributed to them through the distributed software. We think all three lines of business benefit from the other lines of business actively being pursued.

Ross Katz: That’s interesting, and it reminds me of other technology assets that have been developed, for example, by the hyperscalers, where you have a primary customer for the technology that you’re developing and exposing to the outside world, selling the technology directly. If you only develop the technology and were only talking to your customers, then the feedback would be limited because they’re only sharing with you the specific bits of feedback that they think absolutely need to be shared back with you. Whereas if you were only developing internally, the scope of the applications of the technology to drug discovery in this case would be somewhat narrower. You can get that broad exposure that allows the technology platform to become more reliable, more effective, get exposed to different problem statements and problem approaches. Am I thinking about that right, or how would you update that?

Robert Abel: You’re absolutely right. I would also say the software business forces us to give more thought into the quality of the user interfaces and the quality of the automation. When it’s just for internal use, it’s not so bad to communicate to an internal employee, “just do it by command line,” or “you’ve got to go through these 15 steps.” It’s just for you to justify all the investment to make it really clean and fully automated. The software business forces us to do that, and that also benefits our internal use and collaborative use. We benefit from multiple different angles. I think you’re thinking about it completely correctly.

Ross Katz: That approach to scientific technology development, where every tool is a tool that’s for internal use, and even if it’s open source, the internal use version has been exposed. Obviously, that’s a major differentiator when you talk about scientific tooling that you’re paying attention to the interface and the collaboration points. Heading back toward the physics-based approach and what makes Schrodinger’s technology platform distinctive: Would you explain why the integration of the physics-based calculations you were talking about earlier into the drug discovery platform solves some of the major problems you would face if you were more focused on machine learning model development only or on wet lab validation only? How does the Schrodinger approach make drug discovery happen more efficiently or more effectively?

Robert Abel: The place to start there that often surprises non-experts is just how big chemical space is — how many drug molecules could hypothetically be synthesized versus how many have actually been synthesized and experimentally characterized to any degree over the course of all human history. A back-of-the-envelope estimate for just how many hypothetical drug molecules could be synthesized is on the order of 10 to the 60th. That’s vastly more possible drug molecules than exist molecules of water in Earth’s oceans. Doing this from memory, I think Earth’s oceans have somewhere around 10 to the 47th water molecules, which is a way, way smaller number than 10 to the 60th. Over all human history, we’ve only synthesized and experimentally characterized in any serious degree of detail far fewer than about a trillion distinct drug-like molecules, where there’s way more than a trillion water molecules in a single drop of water. When you try to use a machine learning method to learn from this set of less than a trillion and then extrapolate out to the space of 10 to the 60th, it’s quite literally the conceptual equivalent of trying to learn about the entirety of the Earth’s oceans from an analysis of much less than one drop of water. You might conclude, for example, that there are no fish in the ocean because there are no fish in my single small aerosol-sized droplet of water. It is not tractable for many properties to try to learn a global model of that property from such a small training set. Where physics-based simulations solve that very real problem, in our view, is that they give you an ability to generate high-quality, reliable data that is relevant to the drug discovery project. For a given drug discovery project, you don’t need to understand the full 10 to the 60th space; you just need to understand the sub-region of that chemical space that’s relevant to that particular drug discovery project. We can use the physics-based samplings to generate representative data covering the relevant space, and then you can train a machine learning model to effectively triage within that understood space. It’s a way to take the enormous global data challenge that exists in drug discovery, make it a local challenge for a particular project, solve that local data challenge, build the predictive machine learning models that can then be used to triage. When you have these approximate predictions by the machine learning model, you can then use the physics-based simulations to confirm the validity of the machine learning predictions before advancing to wet lab synthesis and assay. We think that’s the right conceptual approach for working in this field, and I think we’ve been winning that argument globally, as more and more companies now seem to be adopting both our paradigm and messaging along those lines. Good ideas have lags. We think that’ll be an idea that continues to be used more and more in this field.

Ross Katz: The reason these physics-based calculations are able to be used in this way is because you’re starting from first principles about the physical system that’s happening in the chemical space that you’re able to model. These physical principles can be mathematically modeled and in some cases, mathematically optimized in ways where you can do these computations relatively efficiently over the size of the space that’s out there. Also, the accuracy and reliability of those calculations has been proven to be accurate enough that if you’re able to use these physics-based approaches, it’s going to give you the ability to extrapolate to parts of the space that, if you just trained that machine learning model on that drop of the ocean, you would not be able to extrapolate there. Am I thinking about that right, or how would you update that?

Robert Abel: You’re thinking about it exactly correctly. The physics-based models are not dependent on input training set data the same way the machine learning models are. The reason they’re not dependent on that input training set data is the physics-based modeling uses detailed molecular understanding of the actual property of interest in order to calculate that property, rather than inferring it from previous relevant data. Once you have that atomistic understanding where you can calculate it directly — although that’s very computationally expensive, on the order of a GPU day per affinity calculation you’re doing — you can then treat that for machine learning model building the same way you would treat experimental data. That’s where there’s this greater power and utility made available by integrating the two techniques.

Ross Katz: Right. The physics-based calculations aren’t perfectly representative of what’s happening in the world. It is still a model, but it is a different type of model that has different strengths and weaknesses relative to the machine learning models being trained off the internal data. What I just heard from you there is also that, from the perspective of someone designing a drug discovery workflow, there are times when you can use the physics-based calculations basically as the generation of your training data or the validation of your predictions that your machine learning model is creating. Then there are other times when you would choose to validate in the wet lab, making sure the outputs in the real world are what you would expect based on the combination of machine learning and physics-based calculations.

Robert Abel: It’s not that the combined physics and machine learning calculations are validating the wet lab data. It’s that you’ll end up generating a much more relevant and valuable set of wet lab data if you follow the guidance of the physics-based and machine learning. You can synthesize whatever molecule you want potentially in the wet lab, but you want to synthesize the most valuable molecules you possibly can. As you mentioned, the physics-based calculations are not one-to-one with reality yet. There will be occasional mispredictions. Based on extensive retrospective and prospective profiling, we would expect about eight out of 10 molecules prioritized this way to look as good as we would expect in the wet lab, which compares very favorably to synthesizing based solely on human intuition or random sampling. We’ve actually published some head-to-head comparisons if people are interested; they can find the relevant publications. When the predictions disagree with experiment, they often turn up some real issues with the running of the calculation. There are two big knocks against these physics-based calculations. The first we’ve already discussed, which is they’re computationally very expensive. The second issue is all the details have to be correct. They need to, with high fidelity, recapitulate the actual environment the protein is in, all the details of the underlying molecular processes. Where there’s disagreement between the wet lab experiment and the physics-based calculation, that often gives us a clue as to how to further optimize the physics-based calculations such that we have a higher degree of fidelity between what’s happening in the simulation and what’s happening in the test tube, or however the experiment is being done. There are also instances where the wet lab ran into some issues. We’ve had some instances where there have been mispredictions where we have to do things like check the purity of the compound or whether the synthesis worked out how we thought it was going to work out. We’ve had cases where, for instance, the wrong compound was synthesized, and we actually had to confirm by 2D NMR to get that confirmed. Not every case; I don’t want to exaggerate how often that’s the case, but it is certainly the case that the details can be off both on the atomistic simulation and on the wet lab experimental data. This also raises an important question that can come up about machine learning methods. In machine learning methods, you’re just training it to the data. You don’t have this natural mechanism that exists with physics-based simulations to try to interrogate, “Okay, is the problem on the cleanliness and quality of the data, or is the problem on the details of the simulation?” That’s an important bit of information that can be disentangled when you combine physics-based methods with machine learning-based approaches.

Ross Katz: This is great. I think where I would like to go is if we could take a mock example of a drug discovery process and start from the beginning. How do Schrodinger’s software platforms fit into that drug discovery process? What are the physics-based methods being applied? What are the machine learning methods being applied? What are the inferences, stage gates, or milestones you’re trying to get to along the way?

Robert Abel: Great. Schrodinger distributes a lot of different types of software to address different challenges that come up in drug discovery. For instance, one challenge that comes up early on in the discovery process is just how to organize all the activities associated with the drug discovery process. You have to do that before you even get into the predictive modeling side of things. We have developed an enterprise informatics solution, Live Design, that allows all the experimental and computational data being generated on the project to be captured in one place in a web-accessible platform. It allows different people to have interactive discussions and planning as to what should happen next. The very first thing we would do is try to get our arms around the full scope of the data of the drug discovery project to motivate follow-on activities. The next activity we would pursue depends on where we’re starting to work with a collaborator on the drug discovery process. There are some drug discovery collaborations where we start, and we’re involved all the way from target selection. There you have to wrangle all the data, and then you have to do some testing and prospective validation and understanding of whether a structure-based drug discovery campaign against that target is feasible. That might involve looking to see if there are public structures available, if there is published data or patents that have been released about those where we can confirm our ability to quantitatively and accurately model the binding affinities of molecules to that particular protein. Ask questions about whether there’s a pocket known to exist in the protein that’s known to bind small molecules. There’s a whole arsenal of questions you need to start to wrestle with even at the target selection phase.

Ross Katz: On the target selection side, are there particular pieces of software or assets at Schrodinger that you’re integrating into it, or is that still part of Live Design?

Robert Abel: There’s all sorts of different stuff that we’ll do. For instance, the first thing we’ll do is look to see if there are any publicly available structures of the protein and whether any small molecules are known. Once we have those public structures, we’ll often re-refine them with Prime X, for example, a software we’ll use where we’ll attempt to improve the apparent resolution of those structures so we have greater confidence in the atomic details of those structures. Then we’ll use methods such as SiteMap and Mix MD to better understand the likely druggability of the various pockets that are found inside those proteins. Once we have determined whether small molecules are known to bind to those targets, we’ll then run free energy calculations to see the extent to which we can recapitulate the binding affinities of those small molecules to those targets. If it’s a target where functional receptor response is as important or more important than binding, we’ll also check whether we can adapt the free energy calculations to recapitulate functional response directly. That’s a relatively recent capability we published, I think, in 2023, about our ability to do that. But for some targets, GPCRs for example, receptor functional response is arguably more important than the binding affinity itself. We want to have as much understanding of our ability to model the target as possible when we’re ahead of running the virtual screen, if we can. Once we have that understanding, then we would run the virtual screen, using both docking and machine learning-enhanced docking. That would be called Denovo VS, which is our ability to use generative modeling and generative machine learning to accelerate the docking process with these super large libraries that are now available. That’s done mostly with the various flavors of Glide, Glide SP and Glide WS. We would then advance the molecules that look best in that docking to absolute binding free energy calculations, and then hopefully assemble a purchase list. We’ll also use ligand-based screening methods that we see as complementary to the physics-based or empirical docking scoring function methods. We try to think methodologically very broadly when trying to find hits, just because the different methods have different strengths and weaknesses. Once we have a hit molecule or a lead molecule, we will both work with medicinal chemistry teams that are ideating around those hit molecules or lead molecules, and use the predictive methods to better inform the synthesis decisions. Then we’ll also do Denovo ideation around those hit molecules and lead molecules, and also do property screening of the ideated molecules. We can now Denovo ideate billions of molecules and triage them to come up with idea sets that are complementary to what the human ideators are coming up with. A recent innovation has been to develop a retrosynthetic analysis solution, which we found has actually been a really important step in making our Denovo design solutions more useful. Many Denovo design solutions will ideate molecules that are in principle synthetically tractable, but some of them may be a “land war” to get them actually synthesized in the wet lab. The integration of a retrosynthetic analysis method that can also give a clear read on synthetic difficulty has been a major step forward to allow our Denovo design technologies to have even more impact on the direction of chemistry on active projects. Hopefully, that was a good smattering of the different things we’ll do. I missed a lot of stuff there. For different properties, it depends on the interest of the project. We have methods, for example, that forecast brain-blood barrier permeability. If that’s important on the project, we’ll do a lot of that. If it’s less important, obviously, it won’t play a major role.

Ross Katz: In terms of what differentiates what Schrodinger brings to the table from the other tools available, either from other commercial vendors or from the open-source options available: Where do you think the highest value arms of the Schrodinger approach are?

Robert Abel: It’s interesting. The part where I spend most of my energy is ensuring the best-in-class accuracy and reliability of the computational methods. For instance, with our free energy calculation technologies, the root mean square error we have in comprehensive testing is about 1.2 kilocalories. You become indistinguishable from the experimental data itself when you’re at about 0.9 kilocalories. That number comes from taking two benchmark methods and measuring the affinities of a set of compounds, let’s say SPR versus ITC; that’s actually the rate of disagreement you see between the two benchmark methods. When you hit that level, it’s not really possible in any easy way to optimize further. We put in a lot of work, energy, and effort to get to that degree of reliability. Competitors will occasionally publish large-scale accuracy benchmarks, including open-source competitors. Their root mean square errors are quite substantially worse. It’s reported when they’re run on proprietary data sets in large pharmaceutical companies, and our software is also used in that context. The difference between the two methods seems to be even greater in terms of accuracy. That’s because we’ve spent a decade doing all this work extending the force field to have broad coverage of chemical space, and also battle-hardening the technology itself to be able to cover really complicated transformations—things like scaffold hopping, macrocyclization, and charge modification. These are things that can be done reliably with our software that are very challenging to do accurately with many pieces of competitor software, and it may be impossible to do with other packages. That’s where we’ve carved out our niche, and that’s where we’re really focused on the accuracy and reliability of the methods. Side by side with those large investments in the accuracy and reliability of the methods, we’ve also put a really big emphasis on automation and ease of use. The more automated and easier to use you can make the software to run it the right way, the more you minimize human error in the modeling process. That also is a difficult-to-disentangle component of the observed error when you’re profiling these technologies. That’s another area where we’ve made substantial investments. We feel quite good about our software offerings and work very hard to ensure that our software customers feel the same way.

Ross Katz: The error that comes from the usability of the software is a really interesting one, especially when you talk about tools that are highly refined and have thousands of possible parameters, and billions or trillions of possible combinations of parameters that you could use. It makes sense that making those tools more usable, making it easier for people to get to the best configuration of the tool to accomplish their goal, is going to reduce a lot of error and also minimize a lot of wasted time and energy trying to get the right result from the tool. If it’s okay with you, I’d like to spend a little bit of time on the free energy calculations, first, explaining what they are for people not exposed to this aspect of chemical physics, how the modeling works, and what the path to optimizing these so they’re almost as accurate as what you can measure using the physical measurement tools we have available can be.

Robert Abel: Sure, absolutely. Free energy calculations are an atomistic technique. If, for example, it’s a free energy of protein-ligand binding, we will actually explicitly model all the atoms of the protein, all the atoms of the ligand, and all the water molecules surrounding both entities. You explicitly model the ligand binding to the protein. It’s done through an alchemical process. For instance, if it’s a relative binding free energy calculation, let’s say I want to add an aryl halide to a ligand. I will alchemically add the halide, which means that rather than, taking translating it through space and attaching it to the ligand, I will slowly — usually we play a movie of this because it’s easier to visualize in this context — you step through space. People are used to diffusion models. You can imagine a hydrogen atom just slowly changing, like imagining the hydrogen atom slowly morphing into a chlorine atom. If you keep track of the fluctuations of the energy of the system as you slowly morph the hydrogen atom to the chlorine atom, both in the field of the protein and then in the solvent field when it’s not bound to the protein, you can take the difference. Because the free energy is a state function, that’s rigorously the difference in the binding free energy of the molecule with the aryl halide versus the molecule without the aryl halide. That’s very interesting information for the discovery team if they’re considering synthesizing the molecule with the aryl halide at that position because they’re usually doing that either to improve the binding potency or to maintain potency while improving some other property. It’s actionable information for the discovery team, the same way that the experimental binding affinity of that molecule would be actionable information for the discovery team. That’s a relative binding free energy calculation. For an absolute binding free energy calculation, instead of turning the hydrogen into an aryl halide, you take the ligand and you just fade it out of reality. It was in the binding site, and you’re going to slowly fade it out of the binding site. If you keep track of the energy of the system as the water molecules begin to penetrate into the binding site because you’re ghosting away the ligand, and you do that same process also in the solvent and take the difference, that gives you the absolute binding free energy of the molecule. That’s very useful if, let’s say, you’re in a hit discovery situation. What is particularly interesting about free energy calculations is that they are just a general machinery to get at in principle any quantity that can be related to a free energy. We can use them not just for binding affinity, but you can use them for binding selectivity, solvation free energy, and solubility. We’ve also been able to adapt these types of calculations to attempt to forecast mutational resistance in the clinic for mutations that destabilize the drug molecule binding to the binding site. Protein thermal stability can also be interrogated with this type of approach. Those are very general machinery, and that’s why we spend so much effort building them and optimizing their accuracy, just because they’re broadly applicable to many different questions that come up both in small molecule drug discovery and biopharmaceutical discovery.

Ross Katz: If I can try to zoom out and summarize what it is and why it’s important: If you can understand the thermodynamics—the transfer of energy within the system—then you’re able to make inferences about what molecules are going to bind, how strongly they’re going to bind, whether making a small change to a molecule will make it bind even more strongly or less strongly. Or, if you look at different conformations of the protein, how strong will it bind in each of these different conformations? It’s a universal, relatively accurate way of exploring chemical space and understanding what configuration of the drug is going to have all the properties you want it to have. By nature of it being calculable and reliable, you can get through space much faster than you would if you were using any of the other methods on the table. Am I thinking about that right?

Robert Abel: Exactly right. You’re using these methods to try to find the molecule that will exhibit the properties you want much more quickly than would be possible to identify that molecule in the wet lab. Again, for non-experts, it’s sometimes surprising to learn that an active drug discovery project that’s quite well resourced, because the synthesis of molecules can often be so challenging, will often only interrogate on the order of about a thousand intentionally designed molecules over the course of a year, at typical medicinal chemistry staffing levels. With these computational methods, we can evaluate a few thousand molecules overnight if we need to. If you combine that with the machine learning methods, then you’re able to triage from those training sets you develop of the few thousand. You can evaluate billions to trillions of molecules — oh, billions, trillions is still tough. But unless you’re using generative methods, then it gets a little tough to figure out the exact number being triaged. But you can certainly triage vastly more than would be possible using wet lab techniques.

Ross Katz: That’s awesome. I think it leads naturally into where machine learning fits into the picture. We’ve outlined what a drug discovery pipeline looks like. We’ve talked a little bit about free energy calculations and how and why they are useful and reliable. I’m interested in where you see machine learning methods fitting into the picture. Then, since you mentioned it, how do the more recent advances in generative machine learning add to the capability of a drug discovery pipeline, like the one we discussed earlier?

Robert Abel: Absolutely. The computational expense of the free energy methods means it’s usually only feasible to evaluate on the order of tens of thousands of molecules explicitly with free energy calculation methods. That’s way more than would be possible in the wet lab. But even though it’s still a big number, it’s a small number compared to the many trillions of molecules that could be of interest to the drug discovery project. We can then use the physics-based calculations to generate training data, and then use those machine learning models to triage that full chemical space of interest. Then, we advance, say, the top 10,000 into follow-on free energy calculations to confirm. Once you have the best molecules from the millions or billions you’ve triaged with the machine learning model trained to the physics-based calculations, you can send your medicinal chemists a list of 20, 30, 40, 50 molecules that are high priority for them to consider for synthesis. If those look good, the drug discovery project has taken a major next step toward its goals. That’s one place the machine learning methods are integrated very tightly; we call it active learning FEP. You’re actively learning to the free energy calculations as they’re being run. Another place it’s important to use machine learning methods is when you’re trying to manage properties that are not directly amenable to physics-based simulations. We discussed rate of clearance earlier as an example. There are still some properties where the molecular processes are either too complex or not sufficiently understood to enable fully predictive atomistic modeling. Then, we, as would any other organization, need to try to use machine learning or some other more black-box type approach to manage those properties. Another place where machine learning methods are being increasingly utilized is actually points of tighter possible integration with the physics-based methods. For instance, it’s become easier to generate very large datasets of quantum chemistry data. That has led to people building machine learning models of molecular quantum energies. The force field is also a model of molecular quantum energies. It turns out you can actually use some of these machine learning interaction potentials that are trained to an enormous body of quantum chemistry data to run statistical physics simulations as the underlying force field for these types of simulations. That’s an exciting direction, something we’re actively working on. We think there’s going to be a tighter and tighter interweaving of the two techniques over time, leading to more accurate, reliable, and computationally efficient predictive modeling methods for the whole industry, for a whole bunch of different properties. That’s where I think we’re still on the part of the integration curve where a lot of clever ideas are coming up about different ways these technologies can be fused together to produce predictive modeling strategies that are more powerful than either technique being used independently.

Ross Katz: You want the models to be functionally coming at the same problem from different directions so that when they agree with each other, you can be relatively confident in the agreement. There’s also the difference in the computational intensity of the models and the quantity of molecules you can screen with them. You’re selective about when you want the accuracy and reliability of the physics-based free energy calculations, and also when you want the scale of the available machine learning models. But then there’s also the opportunity to use machine learning for the physics-based calculations directly, so there can be a check at a lower level.

Robert Abel: Exactly. You can use them in different ways. With active learning FEP, it’s exactly as you’re describing: you’re generating the training set with physics-based data, training the machine learning model to triage based on that information, and then checking with the physics-based calculations. That’s one way to integrate. Another really interesting way to integrate is you can train the machine learning model to the quantum chemical data that’s typically used to parameterize force fields. If you train the machine learning model to recapitulate the quantum chemical data, you can actually run the molecular simulations using the molecular forces coming out of the machine learning model. That’s another level of integration. There’s another level of integration we’re also exploring: to run physics-based simulations, sometimes you need to infer order parameters of the system. Historically, that was done by human experts, but people are exploring whether machine learning algorithms can be used to infer the order parameters for physics-based simulation. I think there’s going to be more and more integration of the two techniques over time, leading to more accurate, reliable, and computationally efficient predictive modeling methods for the whole industry, for a whole bunch of different properties. That’s where I think we’re still on the part of the integration curve where a lot of clever ideas are coming up about different ways these technologies can be fused together to produce predictive modeling strategies that are more powerful than either technique being used independently.

Ross Katz: Where do you view the next biggest opportunities for the expansion and improvement of the Schrodinger platform in the year or two to come?

Robert Abel: Two things we’re working on really hard right now are to expand and extend the physics-based simulations to be able to interrogate not just the on-target protein you’re interested in, and maybe a handful of homologous off-target proteins, but to routinely be able to interrogate all the different anti-target proteins that you want to avoid binding in the drug discovery process. That’s something we’re working on really hard right now. We actually have a beta version of the technology we’re just starting to let people try out. Early feedback has been very positive. There are all these different proteins that you know you shouldn’t touch them. hERG, for example, is a well-known example. If you’re running the free energy calculation for the on-target protein anyway, could we make it really easy for you to also run the free energy calculations for all the anti-targets — think of proteins that would usually show up in a safety panel, just as an example? You can run the free energy calculations for all of those. So, you’d have a much better, earlier read on the safety of the chemical matter under consideration without needing to pay for a quite expensive safety panel. Also, the shipping can take about a month once you decide you want that data before you have it in hand. To make that really easy for people to get overnight when further optimizing molecules, we think that could be really handy. The other nice thing about that is you actually get a binding mode of that molecule in the different off-target proteins just by virtue of running the calculation. If the computational panel detects that you might have that off-target liability, you actually have some actionable information about how you might be able to resolve the liability from utilizing that structure to run follow-on free energy calculations. That’s an area we’re very excited about. As I mentioned earlier, the incorporation of retrosynthetic analysis into our Denovo design systems is also relatively new, and that’s something we really want people to try out. We’ve had cases already where, from running the Denovo design calculation to getting the molecule made in the assay, for even tough stuff like scaffold hopping, that’s been less than two weeks. The reduction of the timeline for us being able to do that is really a consequence that, in addition to being able to calculate all the relevant properties and suggest it’s a good molecule to make, we also ensure that molecule is really easy to synthesize and aligned with the synthetic capabilities of the chemistry team working on the project. That’s also been a really big step forward. So, those are two areas where we’re very excited. We’re also doing a bunch of stuff in biologics that we haven’t had a chance to talk about. If there’s ever interest in part two, we could be more biologics-focused in the next conversation.

Ross Katz: I can tell you I am very interested in part two. We’ll have to think about getting that on the books. Before I let you go, you mentioned that the interface and usability of the software systems you’re producing is one of the key differentiators of what Schrodinger brings to the table. Can you talk a little bit about how you or your team thinks about making computational tools more accessible to someone like a medicinal chemist who might not have a deep computational background but needs to interact with the software?

Robert Abel: It’s a really interesting question. One of the things that’s been really important there has been our Live Design system. It allows the computational chemist, who is the expert, to upload that model for medicinal chemist use with guard rails around its utilization. The computational chemist will understand, “Okay, it can only work for molecules that are within a certain region of chemical space or only molecules that satisfy certain characteristics.” They can apply that as a restriction on the model, post it, and then for the med chem to run the model, they just need to click a button on a web page that says “run.” If you make it easy enough where they just have to click a button on a web page that says “run”—and by the way, that shouldn’t be interpreted as anything negative about the medicinal chemist, they do a lot of other challenging things; it’s just that this part of their daily activity should not be made challenging—when it’s that easy, we have had examples of medicinal chemists where it sort of gamifies the system. They can just go ahead and sketch molecules, keep clicking “run,” getting instant or semi-instant feedback depending on how involved the calculation is going to kick off. Then they can iterate in silico a number of different iteration cycles before actually triggering wet lab synthesis. That’s been an example where, once it’s made that easy and straightforward for them to utilize the technology, we’ve seen a lot of adoption and excitement on the part of the medicinal chemistry teams to use these methods in a bigger way because they can use them directly with their own two hands.

Ross Katz: That’s awesome. It just comes back to showing them the magic as quickly and efficiently as possible and getting them addicted to what the magic can deliver for them. You’ve talked about some of the next things on the horizon for the technology platform that Schrodinger brings to bear. Are there any unsolved problems in computational drug discovery that you think are some of the most exciting things that might be solved by Schrodinger or the community at large over the next five to 10 years?

Robert Abel: Five to 10 years is a really interesting time frame. I don’t want to get too speculative because I think there’s always a danger in these kinds of questions to go a little off the deep end. What I would say is it has been remarkable over the last 15 years to see how extensively both the accuracy with which we are able to model systems has improved and also the efficiency of our sampling techniques. There are a number of problems where right now it appears intractable, both from computational requirements and the degree of accuracy that would be required. Those problems will probably resolve themselves over time. It would not surprise me at all in a five to 10-year time frame if we actually can come up with tractable approaches to get at things like rate of clearance, more comprehensive abilities to evaluate toxicity in ways that could help de-risk molecules substantially before they go into the clinic. We also have substantial investments into things like being able to model battery electrolytes and other energy-related applications. I think over time, all those application areas will mature and materialize. I think it’s very, very hard though to put a five-year time frame versus a 10-year time frame versus a 15-year time frame. I would be a little hesitant to make predictions like that with high quantitative precision. But I do think there will be big advances materializing over the next five to 10 or 15 years.

Ross Katz: If you can design a system that can get to one kilocalorie per mole accuracy on free energy, I can imagine how trying to predict the next five to 10 years in technology development might seem relatively inaccurate and maybe not a worthwhile exercise. Are there any applications of generative AI or Generative AI that you’re really excited about, either in the works or already being applied in the work that Schrodinger does?

Robert Abel: Molecule ideation is a pure win. Also, large library screening. For instance, the vendor libraries have become really, really big. Even with very fast machine learning methods, trying to push the full vendor library through the inference method is pretty painful. It’s much more efficient to train a generative method to generate the library directly and then bias the generation of the library just to the molecules in the library that are best on some property dimension. That’s a clear place where generative machine learning has been a pure win. It’s similar for molecule ideation, for Denovo design. For Generative AI, we have seen really big wins on some of the easier but higher-value things agents can do for people. For instance, it was relatively easy to create an agent trained to, for example, all of our documentation. Rather than somebody needing to ping [email protected] to accomplish some basic task, the agent, Maestro Assistant in this instance, could help the person figure out how to do that task. You can even ask Maestro Assistant to do basic activities itself. We expect those agents are going to get ever more powerful over time, and then it loops back to your comments about how to support the medicinal chemist. You can imagine over time more and more of the work the computational chemist is doing, regularizing and restricting the model, can be done by the computational chemist just asking the agent to regularize and restrict the model in the appropriate way. Or potentially, in some future, large swaths of the more mundane work that the computational chemist is doing could actually be done by the agent. This requires agents to further improve substantially. There’s a ton of huge companies working on making that happen. So, I’m very bullish on what agents will be able to accomplish over time. I think we’re all just going to have to wait and see as an industry how quickly the Generative technologies mature to be able to take on more challenging tasks. But it’s definitely looking like a very promising direction.

Ross Katz: Awesome. Robert, it’s been a real pleasure to have you on the podcast. I really appreciate it. If there are any places listeners want to go to learn more about Schrodinger’s methods and platform, or if they want to follow you and your work, where should they go?

Robert Abel: We have a big company website, Schrodinger.com. People should check it out. We have a big education group too. If people are interested in learning about the technology, sign up for a course. They can learn about all the details of all our molecular modeling technologies directly.

Ross Katz: Is there a way for people who don’t have licenses with Schrodinger right now to be able to test out the tools while they’re taking courses like that?

Robert Abel: Yes. The best way to explore that would be to sign up for the courses directly. There is some software access that’s given as part of the courses because that lets people try them out and do things with them in a way to facilitate learning.

Ross Katz: Robert, once again, thank you for coming on the podcast. Really appreciate it and look forward to connecting down the line.

Robert Abel: Yeah, thanks so much.

Jason: And that’s it for this episode of Data in Biotech. If you enjoyed the episode, please subscribe, rate, or leave a review in your podcast platform of choice. See you next time.

Frequently Asked
Questions

How does Schrödinger ensure the reliability of its computational predictions given the vast chemical space?
Schrödinger integrates physics-based free energy calculations with machine learning. The physics-based models generate high-fidelity, virtual training data specific to a project's chemical region, which is then used to train and validate machine learning models for broader screening. This hybrid approach allows accurate extrapolation beyond sparse experimental data, ensuring reliable predictions.
What specific operational efficiencies or cost reductions does this approach deliver in drug discovery?
This approach significantly accelerates lead identification and optimization. Traditional wet lab synthesis and assay can evaluate around a thousand molecules per year; Schrödinger’s platform can evaluate thousands overnight using physics-based methods, and millions/billions with subsequent ML triage. This dramatically reduces the time and cost associated with synthesizing and testing less promising compounds.
How does Schrödinger make complex scientific software usable for non-computational experts like medicinal chemists?
Through their Live Design platform, computational chemists can configure advanced models with "guard rails" for specific applications. Medicinal chemists can then run these complex simulations with a single click, receiving immediate feedback. This user-friendly interface supports direct engagement from non-experts, fostering rapid in-silico iteration cycles.

Need a data partner for life sciences?

CorrDyn helps biotech and pharma companies build the data infrastructure that accelerates research and operations.

Book an intro call