Skip to content
James Yoder — Success-Driven Drug Discovery with OpenBench CEO James Yoder
Data in BiotechEpisode 65

Success-Driven Drug Discovery with OpenBench CEO James Yoder

James Yoder of OpenBench explains a success-driven model for early-stage drug discovery using computational screening and AI-powered hit scoring.

53:37Full transcript below
JY

James Yoder

Founder & CEO at OpenBench

Overview

Early-stage drug discovery is notoriously capital-intensive and failure-prone. Biotechs invest heavily in hit identification—the first step in finding compounds that interact with a disease target—often without certainty of success or a clear path to return. This significant upfront cost and inherent scientific risk can stall innovation and deplete critical R&D budgets.

James Yoder, Founder & CEO of OpenBench, recognized that advanced computational methods could mitigate this early discovery risk. With a background in statistics and data science, Yoder built a business model where OpenBench takes on the scientific and financial risk of hit identification. Clients pay only when OpenBench delivers novel, potent compounds meeting pre-defined criteria, shifting the burden from the biotech.

In this episode, host Ross Katz speaks with Yoder, who details OpenBench’s journey and its “success-driven” approach. He explains how their data systems, built on active learning and proprietary scoring functions, efficiently screen trillions of virtual compounds. The conversation covers how OpenBench assesses target druggability, establishes success criteria, and continually refines its models with project data, illustrating how a data-first strategy redefines drug discovery economics.

Key Takeaways

Shifting drug discovery risk requires rigorous upfront data assessment.

OpenBench’s “success-driven” model moves financial and scientific risk from the biotech to the discovery platform. This is enabled by a free, one-week feasibility study that rigorously assesses target druggability and establishes specific, measurable success criteria. This upfront, data-informed diligence is essential for profiting from a risk-bearing model.

An active learning loop efficiently searches trillions of virtual compounds.

To work with ultra-large virtual libraries—trillions of compounds—OpenBench employs an active learning loop. This process involves sampling, molecular docking, proprietary rescoring, and building proxy models to infer affinity across the vast chemical space. This approach prioritizes promising compounds for synthesis and testing, moving beyond brute-force enumeration.

Stratified search strategies improve hit diversity and relevance.

Rather than uniform sampling, OpenBench stratifies its search within massive virtual libraries by reaction subspace. This prevents oversampling common chemistries, like amide couplings, and helps identify good chemistry complementary to specific binding sites. This method yields more diverse and relevant chemical matter, optimizing the chance of novel hit identification.

Collaboration data builds a global scoring function for continuous improvement.

Data from every OpenBench collaboration—including both active and inactive compounds—feeds back into a global machine-learned scoring function. This continuous data flywheel improves the model’s ability to predict molecular affinity across all target classes. Benchmarks confirm that the scoring function’s performance improves with each project, increasing the likelihood of future success.

Related: CorrDyn serves the biotech and life sciences industry by building data systems that drive real discovery. We help clients apply machine learning to complex scientific problems and conduct data assessments to ensure their technology strategy supports business outcomes. For insights on building efficient data operations, see our post on realizing data value in biotech manufacturing.

Full Transcript

Jason: Hey everyone, it’s Jason, producer of Data in Biotech. Quick one, every episode of this podcast is now on YouTube with animated glossaries that break down the technical terms we discuss. Just search Data in Biotech on YouTube to watch. Welcome to Data in Biotech, a podcast from CorrDyn where we explore how companies leverage data to drive innovation in life sciences. Every two weeks we sit down with an expert from the world of biotechnology to understand how they’re using data science to solve technical challenges, streamline operations, and further innovation in their business. Here we go.

Ross Katz: James Yoder, welcome to the Data in Biotech podcast.

James Yoder: Yeah, thanks for having me, Ross.

Ross Katz: Awesome, well just to kick us off, would you mind introducing us to your background and what led you down the path of drug discovery?

James Yoder: Sure thing. It’s a little bit of a circuitous route. I studied statistics in undergrad at Harvard and then I found myself working as a data scientist and started moonlighting essentially as an ML consultant for a public biotech in North Texas, of all places. While I was doing some research trying to figure out how to build reasonable molecular property prediction models, a high school friend of mine who’s a brilliant AI researcher put me onto using some message-passing neural networks that Justin Gilmer and his colleagues at Google Brain published in 2017. Working with my now co-founder Jim Thompson, we went from there. The models worked well enough that two outsiders could make an impact on small molecule drug discovery and were naive enough to start our own company on the basis of one anecdotal success, and then the OpenBench dream began.

Ross Katz: What I would love to do is understand a little bit of the history of OpenBench because as I understand it you’ve gone through a business model change over the course of your history. So would you mind talking about the founding story of OpenBench and what you learned and where you are now in terms of the model that you’re using?

James Yoder: Because I was a data scientist by background and my co-founder was a software engineer, we thought it was natural to start a software company. We were planning to do molecular property prediction as a service where we’d expose some predictive endpoints through a ChemDraw interface and medicinal chemists, computational chemists would go in there, draw their molecules, get predictions on any range of properties and use those predictions to influence what they made and tested next. We ended up building basically a prototype of that software and selling it to a few biotech companies, licensing it to them on a pay-as-you-use-it basis. We pretty quickly realized it was not a very good business for us at the time. The models were so-so, it was very data hungry because it was a supervised learning system so we needed a fair amount of data to train these models. There was some literature data available, but not enough to make things interesting. Maybe we’ll come back to this, but some of the federated learning approaches for instance that TuneLab and Eli Lilly are doing is very interesting to me because they have the data handy and they have a consortium that they’re building. So the models weren’t great and the business wasn’t great either. It’s a fairly shallow market. There wasn’t a ton of willingness to pay at the time for biotechs buying properties predicted on a screen. The understandable ethos was why don’t we just make compounds and test them? That’s the way we’ve been doing things for a long time, it works pretty well. We found basically that the usage metrics just weren’t that good even for the companies that we signed on. So we assessed that, okay, this isn’t going to be a viable business. We went back to our initial pre-seed investors and said, hey, what do you think? Should we give you your money back? We’ve hit a wall here. And they said, no, you’re smart young men, figure something else out. And this was in the end of 2020, beginning of 2021 where it felt like money was free and they were like oh, whatever, you’ll figure something else out. So we ended up pivoting into our business that we still run today. We hired a brilliant chief scientist, Lewis Martin, who’s still with us, who had an actual science background in biophysics and computational chemistry. And he had developed some methods that we had basically seen his publication, he had put together a preprint on a way to computationally effectively screen ultra-large virtual libraries and pick top hits without having to do brute force characterization of every single molecule in that library. We liked the approach because it was very practical, very pragmatic to say, okay, we’ll lose a little bit of precision here, but we’ll be able to do this a lot cheaper. We picked up Lewis, we figured out okay, we have this interesting concept, how are we going to actually get people to buy it from us as opposed to all the other software players out there? Why don’t we bundle this up into an operating model that could be attractive in its own right? We developed the success-driven molecular discovery model that is our differentiated thesis to this day. This idea that okay, don’t trust that our technology works, we’ll deploy it ourselves, we’ll make compounds at our own cost and scientific risk. We’ll test those compounds in your assays and if and only if they work in those assays are you then on the hook to purchase the qualified hits from us. That’s a brief history of how we arrived, how we conducted this pivot in the spring of 2021. And pretty quickly we found a couple companies that were willing to take us on. We hit the ground running from there and are proud to still offer the success-driven model today.

Ross Katz: I want to unpack the success-driven model a little bit. You mentioned that you’re making compounds at your own cost and scientific risk. Can you walk us through what a relationship looks like with a new company that comes to you with a molecule or a set of molecules and how you work with them to determine what the compounds are that they’re then purchasing back from you after it meets some sort of standards or criteria for them?

James Yoder: Absolutely. In general, what a company comes to us with is a target. They’ve identified that target, perhaps through their own platform or perhaps they find it interesting because of other work that’s been done against the target or work that’s been published in the literature. Folks come to us when they want that novel potent chemical matter that can be the starting point for developing an asset that will eventually make it to the clinic. Oftentimes it’s not a molecule or set of molecules they’re bringing us, but rather just the identity of a target that they think is promising from a therapeutic perspective, from a translational perspective. Our task is to find novel chemical matter that engages that target in an interesting manner. To that end, we take a look at the target and really we’re looking at the structure of the target to see whether or not it’s amenable to being drugged or being liganded is usually the first consideration. Can we find binders to the target? To bind a target usually you need some pocket to bind to, some groove, some cleft, some site that has concavity that has opportunity to make interactions that would drive binding affinity. The way that we assess these targets, there’s nothing really special to it. We’re looking at the same thing that any medicinal chemist would look at or drug hunter would look at. Is there opportunity to make very high hydrogen bonds? Is there opportunity for hydrophobic burial more broadly? Are there key interactions that will drive affinity? We conduct this feasibility diligence and we try to be very pragmatic about it. Somebody sends us a target, they send us whatever structural information they have available, they send us maybe some literature that they think is the key literature around the target. In certain cases they say, we don’t actually have a structure of the target, but you guys can try to generate some predictive structures and work from those, and that’s increasingly common. We do this assessment in one week and we do it for free so that we can abide by our success-driven promise, which is that the first dollar they spend is to purchase interest of chemistry, something that’s developable. That is the first interaction: this feasibility diligence once a target is disclosed under CDA. Usually, I’d say about two-thirds of the time we come back to the prospective collaborator with a thumbs up. And maybe another one in 10 times we have a thumbs sideways where we say, okay, we’ll need to generate some data to validate that this site actually exists. And then sometimes we just have to give a thumbs down because there’s not sufficient opportunity to bind the target. It may just be that they have a protein-protein interface in mind that they want to drug that’s very flat and featureless and not good to approach with a small molecule. So an assessment that we make and we go from there.

Ross Katz: I guess I’m really surprised at how frequently you’re saying yes to people coming to you with proteins that are drugable. I’m interested in how do you handle the activity and then also how do you account for the frequency with which you’re able to say yes with this success-driven model given that a lot of drugs fail is the prior that I have there.

James Yoder: This assessment is just the first step. The next step usually is to consider okay, we have a target, we think it’s druggable or ligable using our technology, what are the assays that are available to confirm that experimentally? There may be a suite of cell-free assays that we start with or biophysical binding assay, looking directly at measuring binding of a small molecule to a protein target, biochemical assay that maybe judges binding by saying okay, we know if we bind to this site that we’ll be competing with some endogenous substrate and we can measure the competition to get an IC50 or a Ki in an inhibitor, if it’s an inhibitor design setting. We next look at these assays. Are the assays reliable oracles for specific engagement of the target? And if they are, then okay, second box is checked, we’re that much closer to striking a success-driven collaboration. The third major box is to define the success criteria. Given the nature of the target, the binding site, the prior art if there is any history on the target, we define success as a set of experimental and non-experimental criteria. The experimental considerations are usually potency first, maybe we look at potency in a primary assay and some orthogonal assay, we want to make sure those corroborate each other. We want to potentially look for selectivity or there’s other aspects now of the target product profile that start filtering back to the hit profile that we’re trying to achieve. There may be known off-targets that drive toxicity. Okay, let’s make sure we avoid those in whatever we’re delivering. There may be certain things around mechanism of action or maybe even a desire to confirm via structural biology. Let’s get a crystal structure to confirm that this is indeed binding at the site of interest. There’s all of these assays that we say okay these are going to work and then we set thresholds essentially in each of these assays to define success. Oftentimes for a well-established target we want sub-micromolar potency in some cell-free model. It’s a baseline threshold for success. If the target’s never been drugged before, we might have a looser criteria for success. Something that’s even double-digit micromolar as long as it’s confirmed in multiple orthogonal systems could be valuable as a starting point, acknowledging that there’s ample medicinal chemistry that would have to be done to get that to a potency range where it could be a good drug. There is a very case-by-case bespoke success criteria development.

Ross Katz: What I hear you talking about is that you’re establishing a set of standards for what would be a commercially viable thing worth purchasing from your partner in advance and you’re trying to establish the scientific criteria, the experiments that you would use in order to say this meets those criteria or this doesn’t meet those criteria because ultimately those criteria are going to be contractually governing in terms of whether your partner is obligated to purchase the compound back from you. And there’s also the, if you lower the criteria that need to be met then you can potentially lower the price that needs to be paid for the outcome. What I do think is data science relevant about what you’re doing is that you have to assess upfront not just how likely is this to succeed but what are the scenarios under which we will need to spend certain amounts of money in order to ensure the success of the compound that you’re developing. Am I thinking about that right or how would you update that?

James Yoder: No I think that’s very accurate. And I think it is an interesting problem from a data scientific and broader scientific perspective. How do we assess the risk of doing this? There’s all sorts of what we help epistemic risk, that’s not a phrase we’ve coined, but a phrase that exists to say okay we’re embarking on a scientific effort, there’s a lot of unknown here. If we fail we’re eating the cost of failure, we’re not a big company that can afford to fail very often. We’re trying to provide a good service, we have a reputation to protect, we can’t afford to fail very often from that perspective. How do we take all those risks? How do we pair those to our presumed out-of-pocket costs? Where the cost is also going to be associated with a risk. If we think it’s a hard target we’re going to be making more compounds, we’re going to be doing more characterization. How do we mitigate those risks using technology? How do we improve our ability to predict these risks given our history running the success-driven model? And then how do we simplify all of those risks and costs into a single dollar amount that we can quote to a collaborator? That has essentially been one of the main projects of the success-driven model is trying to figure out how to do all of that.

Ross Katz: You are taking on all of this risk but that also, I’m imagining from a potential customer’s perspective, one of the downsides of entering this is that you’re exposing your IP to a potential partner without any financial transaction in advance. I’m just interested, who are the types of companies, of biotechs or researchers or members of the biotech ecosystem that are ideal partners for you to work with for whom the success-driven partnership model is a great opportunity for them as well as you?

James Yoder: The success-driven model drives us towards clients who are broadly speaking resource conscious, which these days is pretty much everybody in drug discovery. A lot of ink has been spilled over the fact that drug discovery is expensive, that it’s slow, and people are looking for ways to use technology to try to improve the efficiency of drug discovery. In our niche we’ve taken a bite out of this by offering the success-driven model. The earliest clients were all in biotech. Small molecule therapeutics, biotechs that may be pre-seed or seed stage or they may be public biotechs. But in general there’s a lot of innovation happening where these are biotechs that are servicing new targets that are willing to put money behind those targets but also want to apply those resources in an efficient manner. They say, okay OpenBench, we have this new target, nobody’s drugged it before, you claim that you can target it for us, that you can find some start of chemistry. Have at it. And if you guys fail we’ll still have our budget to try a DEL or a high-throughput screen or other more conventional method we could try. That’s very common, where we’re a first line of screen. It’s also common that people come to us after they’ve exhausted their conventional efforts. They say we’ve already run a high-throughput screen, we’ve already run our internal libraries, we’ve tried to do some work maybe starting from literature starting points but we don’t have anything that’s motivating us to apply additional resources on hit-to-lead and optimization so OpenBench, if you can meet these success criteria we know that our bosses will basically give us the resources to keep working on this project. The board will approve of releasing the next tranche of funds to get this to lead stage. That’s another instance where we’ll come in as the back line. And that’s always the most satisfying projects for us where we can pat ourselves on the back and say okay they weren’t able to do this using conventional methods but our approach was able to solve this problem. For the longest time and still today our primary value prop that we’re selling is don’t pay us unless we succeed. Increasingly we have a track record to say work with us because we’re able to do stuff that nobody else in the world is able to do. We have this track record established through these previous projects and that’s pretty cool and I think that will open up over time a different sort of market where it’s less the cost conscious and more like the stereotypical pharma buyer doesn’t care about saving half a million dollars by accidentally running a screen that doesn’t work, they care about finding that novel chemical matter that could someday become a blockbuster drug. I think the track record will be a platform for us to sell to that sort of buyer in due time. We’re trying to figure out a little bit on the fly how to appeal to that buyer because we’re so used to message discipline around the success-driven model, and some people just don’t care about that. About saving or losing the money or running a screen that doesn’t work. I think the technology is broadly applicable but the earliest market was the smaller biotechs.

Ross Katz: That’s interesting and it also strikes me that because you’re doing that initial screen for free, you’re also giving information to potential customers just by engaging with you. So there’s an exchange of information. They’re exposing the target to you but you’re also exposing to them your evaluation of that target. Are you willing or not willing to undertake a success-driven collaboration that actually produces the novel compound that binds that target and meets all of their criteria. So that in and of itself is interesting. Before we move on from the success-driven model, would you just share what your success rate is thus far? What percentage of collaborations have resulted in success and then, if you have an assessment of your break-even success rate. How successful do you need to be over the long term in order to succeed as a company? I’d be interested in that as well.

James Yoder: The success rate that we’ve achieved to date is just a matter of fact. We’ve now run 16 projects to completion with commercial partners and 12 of those were successful. 75% to put a number on it. In those success cases, in the majority of those cases we’ve sold more than one series. There’s a fee for each series in general is how we construct the agreements. I think that is basically right in line with our target success rate thinking back to this law of large numbers. What we aim for is scoping a project where the likelihood of success is around 80%. Though we try to be open-minded. If we assess a priori that there’s a 50% chance at success, we should be able to quote a deal to that success rate that’s still profitable to us from an expected value point of view. I wish sometimes that we just had standard pricing and we said okay we’ll accept this amount of scientific risk for this set price because it would make our BD a little bit easier. We’d be able to go in that first meeting and be like this is exactly what it costs. But it’s not a one-size-fits-all project because no drug discovery program is the same as the next. Over time we have been willing to take on more risk for higher prices essentially. There’s new sources of risk that we find to be interesting from a market expansion perspective or from an impact perspective. It’s always going to be more challenging to work from a predicted structure of a protein than an experimental structure. Predictive models have gotten a lot better, there’s a new set of epistemic risk that those models carry, there’s different sources of uncertainty of how a binding site is arranged that we’re trying to model. It unlocks the ability to do structure-based screening against a wholly host of targets that haven’t been crystallized before or that are defiant to being crystallized. That’s really cool but it also means that we have more risk. Maybe over time ironically our success rate will get worse but our business will become more viable. I believe that could be true. In general, it’s always painful to fail but we have the ability to sleep at night knowing we are bearing the cost of failure when we do fail and we haven’t charged an arm and a leg to our clients for that failure. Now there’s still time and effort that they put into these projects, don’t get me wrong, it’s always painful to deliver or to come to the conclusion that we’ve failed on a project.

Ross Katz: One thing about your model that’s really interesting to me is that you seem to have come up with a model that gets around the principal-agent problem that you often have when you’re doing partnerships with researchers or consultants, and I say this as a consultant myself. The thing that enables you to do it is that you’ve created this set of computational approaches to working with your customers that are sufficiently advanced that they give you the information you need to make decisions on the fly about which of these projects is worth working on and also some information about what is the path to getting there because you need to also price out how much investment needs to be made in experimentation, in computational cycles, in people working on the project, etc. To the extent that you can, I would just love to unpack, what is the end-to-end data engine that’s setting you up to be able to do this kind of model? Let’s start from the very beginning. You’ve gone through this feasibility study, you have settled on the criteria, you have agreed on a price, you’ve signed a contract. Just from there, from a computational perspective, what happens from end to end?

James Yoder: The first thing that we’re doing is figuring out what are the models of the protein that we want to screen. We’re setting up a horse race internally. Essentially we’re going to at the end of a virtual screen make somewhere like 600 to 1,000 compounds. We have a pretty good sense of what the library is that we’re going to be screening. Typically we’re screening these ultra-large virtual libraries. There’s libraries these days numbering in the trillions of compounds that are commercially available that have not yet been synthesized. Our preferred vendor for this is a company called Enamine, who are the global leaders in this space in terms of pioneering these ultra-large virtual libraries for commercial use. They are phenomenal partners, shout out to our friends at Enamine. They basically say okay, here’s a set of building blocks we have available and a set of reactions that we can run to combine those building blocks. That’s half of the equation for the virtual screen, that virtual library. It’s not even explicitly enumerated, to explicitly enumerate it would be computationally challenging because it would be hundreds of terabytes of compounds in a flat file and it’s just not very ergonomic. It’s not something that you can just copy onto an e3 instance. A lot of what we’re doing now is on the fly. We’re sampling reaction subspaces within this large chemical space, preparing to sample those in a way that we can turn the crank on is part of our data stack. Ingesting these libraries from Enamine and similar companies and preparing them for screening is the first piece. I got a little bit ahead of myself. Behind that is figuring out the structures that we want to screen. If we’re working from a predicted protein model, there may be different snapshots essentially that we’re taking, predicted snapshots that have slightly different arrangements of the binding site. In principle, sometimes binding sites are rigid, but a lot of times they’re breathing. A protein is sitting in solution and it’s moving all the time. These are very dynamic systems. But what we’re doing to screen first at trillion compound scale is just taking static snapshots that we’re going to model ligands into. We’re picking those snapshots. What do we think are the most interesting? We actually run the virtual screen. We basically say okay we’re going to take these large virtual libraries and we’re going to start sampling ligands from those libraries and modeling them in the binding site. In the binding site we are basically generating a bunch of poses that this ligand could take and scoring those poses. This is a concept that we did not pioneer. Basically molecular docking has been around since the 1980s and has been worked on for a long time by people much more intelligent than us and we haven’t really reinvented the wheel here. There’s some new methods that we use sometimes for pose generation, increasingly co-folding methods can be useful, like all-atom protein-ligand co-folding. But we also don’t really do any research in that area ourselves, we’re basically just receiving the state-of-the-art both in more conventional docking and in co-folding to generate poses. And then we are using our own technology to score those poses. In certain cases it’s called re-scoring. Essentially to try to use our machine learned methods to predict the affinity of a ligand to a protein if it occupies that pose. The pose is important because the affinity is being driven by all of these physical interactions between hydrogen bond donors and acceptors and the distance between a donor and acceptor matters a lot. There’s all sorts of interactions like that that are driving affinity, there’s certain hydrophobic burial that’s guided by how the ligand is actually sitting in a binding site. Is it displacing certain waters that may be in that binding site in solution? Things like this are driving affinity and essentially that’s what we’re modeling. We’re guessing could any given ligand occupy a certain pose and what is the affinity that it would have in that pose? We do these predictions at scale. We’re sampling on the fly, sampling canonically, docking, re-scoring. At a certain point we basically feel like okay we’ve built enough of the sample of ligands mapped into affinity that we feel like we have a training set. We can use that training set to build then a proxy model. It’s a lot easier to infer affinity using this model across trillion compound chemical space. It’s a lot more computationally efficient to do that rather than to dock and score every molecule. We use this proxy model as a tool to infer and then search through the chemical space. Really what we’re doing these days is we’re looking through each reaction subspace to try to find the highest scoring ligands. We use this proxy model to prioritize ligands to feed into the scoring oracle. Basically it’s an active learning loop. We actually generate this computational label, this score, we train a model using molecule into label and then we select the next sample based on that inference and we do it all over again. We’ll do that until we basically hit certain stopping criteria where we’ll see a saturation of score after a few iterations. Then we have our highest scoring molecules across these different reaction subspaces, we end up selecting our favorite reaction subspaces and focus on those to start winnowing down, basically going from trillions to hundreds that we’re going to make and test. All of this is basically happening out of curiosity, mostly using Python scripts that are being executed from a shell script. We run our own scoring functions in Python essentially that are trained and we’re using those trained weights. After we encode a protein-ligand system, we’re basically passing those encoding through a model to get that predictive affinity output. All of the active learning loop is also happening basically using supervised approaches in Python. We’ve built this whole auto screen infrastructure is what we call it internally so that we can make sure we’re running stable releases so we reduce any technical risk. We have enough scientific risk; we try to mitigate as much as possible technical risk when we’re undertaking these screens.

Ross Katz: I would love to hear from you, what have you all learned about searching these massive libraries and finding the subset of compounds that are going to be best both for computational evaluation and then also for bringing together in order to train the scoring algorithm that you have further down the pipeline?

James Yoder: In general when we design molecules, we’re almost never making the molecules because we think they’ll be good training data points. We’re making them because we think they’ll be good hits. We think there’s a good chance that they are novel, potent, developable and that they’ll essentially satisfy the success criteria that we spoke about. It so happens that the vast majority of the molecules that we make and test in any given project, I would say over 99% of the molecules we’ve made are molecules that have never been made before and that generally occupy a chemical space that’s not exemplified in the patent literature. Maybe saying the chemical space hasn’t been exemplified is a little bit of an overstatement, but these precise molecules have not been made and tested before in the vast majority of cases. There’s this nice byproduct which is that if these molecules are active or inactive, we can have basically a label on those molecules, a label that was generated by predicting this property using our existing vast iteration of the scoring function. We could take that and feed it back into the scoring function. Just to address the way that you phrased that question, we’re usually not doing search with the objective being that we’re doing some sort of uncertainty quantification that’s giving us that next molecule that’s most effective. Sometimes in the purely computational loop that is happening, where we’re trying to surface molecules that we think are going to be high scoring via inference or that are high perplexity via inference. But when we’re actually making compounds, it’s usually because they’re already high scoring. Now a lot of those are inactive. The vast majority of the compounds we actually make have never been made before and they’re inactive. It’s very useful to be able to feed those back into our training set and to train new scoring functions that learn from that. It’s like scolding a child who’s done something naughty. You have to learn from example. You told us that these 100 molecules would be active and only three of them were actually active, take a look at these 97, take a good hard look at them and think about what you’ve done by surfacing these. I’m not a parent yet so that’s probably bad parenting advice to try to rub your child’s mistakes in their face but I think the analogy stands. We’re always making molecules that are inactive and to be able to integrate those into a training set is valuable. We’ve learned a lot about search partially by necessity. Over the course of OpenBench’s existence these libraries that we screen have gone from billion compound scale which was able to be enumerated in a matter on gigabyte scale into trillion compound scale. There’s a form factor there that we’ve had to respond to and that the industry more broadly has responded to. A lot of people are thinking about ways in which they can search in reactant space as opposed to product space. The number of theoretically accessible small molecules that are in a drug-like chemical space is beyond comprehension, it’s on the order of 10 to the 60 is what some people have said. And those people aren’t just talking out of, they’re not just people have done these calculations, and we’re just capturing a very small piece of that that’s low-hanging fruit from a synthetic accessibility perspective. But even within that low-hanging fruit there’s a possibility to fall into local minima if you don’t stratify these spaces according to different reactions. In our experience there’s a lot of amide coupling reactions that occur and a lot of building blocks that are amenable to amide coupling in the Enamine library. We’re not the only ones that have observed this. There’s tens of billions of amide couplings. Maybe you could end up oversampling those if you just take a random sample of chemical space and dock and score it and then go from there. Doing this stratified sampling and thinking about almost building an individual model for each reaction subspace has been a huge boon for us in terms of enriching hit rates and enriching the diversity of chemical matter that we’re making and testing, which I think is a big part of the point. If you’re screening trillions of compounds you really want to be able to take advantage of the full diversity. It may be that there’s smaller, quote-unquote, reaction subspaces that merely have tens of millions or hundreds of millions of compounds but where there’s really good chemistry that’s complementary to a binding site, it may be missed if you’re not being judicious about how you’re searching and sampling these massive spaces. Just because a simple random search of tens of millions out of trillions, even if you sample a million molecules, you might get only a couple of exemplars, a couple of representatives, and then you’re training your proxy model on very sparse data for that particular reaction subspace and at least our oracle isn’t good to pick up on that with such a small sample size. Hopefully that’s making sense, it’s a somewhat complex thing that we think about a lot so I may be underspecifying the shape of the problem and how we think about it but there’s been a lot of methods development in search. What makes our scoring function bespoke and the great thing about the success-driven model more broadly is that we are bearing this cost and scientific risk and it basically creates a permission structure in which our collaborators, unlike the standard fee-for-service MSA relationship with a service provider that I spoke to earlier, our collaborators say, okay you’re bearing this risk, you’re making these molecules at your own risk, you’re paying to have them tested, we’ll license back to you after we purchase these libraries the rights to use these libraries to improve your technology as well as the data that’s been generated in our assays. That’s really cool because now we’ve run a couple dozen of these and that’s really high quality purified samples tested in primary assays against a wide range of targets, target classes, binding sites, and there’s certain global biophysical principles that we’re able to learn with sufficient data, layering that on top of high quality literature data that we’ve pruned and cleaned up. In particular it’s exciting to know that there’s a data flywheel that’s feeding back into trying to avoid these pitfalls that we’ve fallen into before.

Ross Katz: It leads to the potential hypothesis that as you do more and more of these collaborations in addition to getting better at selecting the potential collaborations that are likely to result in success, you’re also able to, because the scoring function of the libraries is built on the data that you’re creating as you’re going through these collaborations, that you’re also going to increase your own success rate by gathering that data. I’m just interested, are you already seeing that with the data that you’ve gathered so far and what’s your sense of how applicable one version of this for one client is to the next compound that you’re developing for the next client?

James Yoder: It’s a really interesting question. I think in principle we can say with confidence that based on the benchmarks that we’ve developed the data we’re training on from previous collaborations is improving our scoring function. And that improvement of the scoring function in principle, exactly, increases the chance that we succeed on the next target. The scoring function that we have is a global scoring function, so it’s not something that is even fine-tuned on whatever literature data may exist for a given target. There’s a lot of conjecture really around how our training data improves our scoring function from one project to the next. What we see in general is that in 2021 we fashioned a benchmark using a basket of targets. Kinases, GPCRs, other enzymes, other receptors, a suite of targets for which there was ample literature in here are molecules that bind these targets. We basically took all those actives, people don’t really publish their inactives that much so we basically created a property matched set of decoys, inactives, and it’s essentially a classification task, for a given target is this molecule active or inactive? We’re able to generate classification metrics on this task, your classic receiver operating characteristic metrics, what’s the area under the curve for a given scoring function telling us whether it’s able to differentiate actives from inactives. Then we would look at across these targets, for some of them we have high AUCs, .8, .9, .99 for certain targets because our training data is enriched with structures, protein-ligand systems that are close to whatever that target may be. For example if we’re looking at a competitive kinase inhibitor, a type 1 inhibitor, there’s a lot of crystal structures that exist in the literature that show you need to make this hinge binding interaction to drive potency in most cases and there’s sufficient depth and diversity of training data to be good on those targets pretty much out of the box. But there’s a bunch of other targets where we underperformed. What we’ve seen over time is that yes, now today basically our worst case performance across any of the targets is better than what our average performance was in 2021 in large part due to the data that we generated, cleaned up over the course of our collaborations as well as improved our ingestion of data from the literature and improving the method itself, the actual algorithm and how we featurize these systems or encode these systems. That’s been a huge boon for us in terms of the enrichment of hits in the libraries that we make. I have good reason to believe that that oracle will continue to improve over time, our ability to predict affinity, which should be a positive feedback into the way that we bear risk and our ability to drive good value for our collaborators. That’s an exciting observation.

Ross Katz: For sure and the thing that strikes me about that is if I understand correctly the benchmark you created prior to a lot of these collaborations so it’s not as if the benchmark is weighted toward the collaborations that you had, the benchmark is representative of all the potential collaborations that you might have had and so what you’re seeing is that the selective collaborations that you’ve had are improving your performance on a basket of collaborations that might have looked very different than the ones you actually had.

James Yoder: Exactly. The way I’d restate that is the targets we’ve worked on in collaboration do not overlap with the benchmark targets. I think actually just by coincidence one target we’ve now worked on is in the benchmark. But that’s one of many and it’s a good sign that we’re learning something globally about these protein-ligand systems that we’re able to apply to the next target that we haven’t seen before.

Ross Katz: Awesome. Well before we wrap, we’ve spent a lot of time on your model and on the data science approach that you use to be able to enable that model. Would love to hear about some examples of companies that you’ve worked with and how those projects went.

James Yoder: We always have to be sensitive when talking about our collaborations because we want to protect the competitive interest of our collaborators. It’s one small point of frustration for me because I think we do all this cool work but we can’t really show it to people because it’s not patented yet by our collaborators and they’re trying to carry it forward. But there’s a couple of collaborations I think I can speak to because we put out some press. One of those is with a company that was one of our first to whom we almost owe a lot named Tavros Therapeutics. Tavros was a functional genomics platform that spun out of I believe some work at Duke and they’ve since been acquired by a company called Vividion which themselves are a Bayer subsidiary down in San Diego. Tavros though their emphasis was really on the biology. They had this CRISPR platforms surfacing synthetic lethal targets for application in cancer therapeutics. They didn’t have the chemistry expertise to launch these projects internally, they didn’t have their own libraries, small molecules, they didn’t have existing vendor relationships for high-throughput screen or anything like that nor were those screens something that they could prosecute internally. We found them at a good time where they were very cost conscious, they were very young, seed-stage company when we first came across them in 2021. We had this newfangled idea and their founder Owen who I admire a lot and who’s a very scrappy and effective entrepreneur was kind of like, okay. We had good fortune where we had one case study we had developed internally which was on the same target class that they were going after. We said look, we can do this. We had actually some relationship with one of their consulting medicinal chemists prior, which was also good fortune for us. They said okay, we’ll give this a shot. They saw the appeal of the success-driven model, we gave them a pretty sweetheart deal because they were one of our first collaborators, we were still doing a little bit of price discovery. It worked so well against that target that we decided to re-up and undertake a multi-target collaboration after the initial target. In 2021 we launched our first target, we delivered some really nice quality chemical matter in 22 and then 23 we decided to take on multiple targets together coming out of their platform. That was a great story for us because it showed that we were able to serve young companies that were very cost conscious and that we were able to be complementary to companies that were really strong on the biology platform where we have really no expertise. Despite not having expertise in their targets or their CRISPR tech we had our own operating model that could fit really well and be very complementary to them. We know that they were able to take those molecules and improve upon them substantially, which is also encouraging. It’s always a question that people have for us. We don’t maintain any formal audit rights. But it’s always encouraging to know, okay, the goal was to deliver something that was potent, novel, developable, we know experimentally that these are potent, we know they’re novel just because we can reference the literature, but are they really going to be progressible? And that’s a little bit of an unknown until you start chemistry. They were able to make a lot of progress with the potency and selectivity of these molecules and getting proof of concept in more advanced models than just the biochemical systems that we were originally testing them in. Another case on the other end of the spectrum: whereas Tavros was very small, we were working with a company OnKure Therapeutics with whom we were working with a really seasoned set of extremely good drug hunters, medicinal chemists and computational chemists and biochemists and structural biologists who had existing relationships with CROs and a lot of capabilities internally on the chemistry side. They were just going after a really hard site that was essentially a novel site and they saw OpenBench as an opportunity to take a flyer and say okay, you guys think you can do something against this difficult site, prove it. We did, which is awesome. We were ultimately able to release that we found the most potent known chemical matter at that binding site of interest. It was a nice feather in our cap when we were able to share that with the world and these are scientists that we worked with for whom I have an immense amount of respect and who have been around the block. A lot of them came out of the Array BioPharma diaspora. Array has been one of the most productive research houses this century, they were a services company then developed their own pipeline, got bought by Pfizer in Boulder and in Boulder a lot of their scientists went into the wind and a lot of the seeds germinated all over Boulder and we got to work with that crew who had done some really impressive stuff that we didn’t even realize we had read their papers previously as cool work. A different case, working with a more established company where they had a lot of in-depth expertise around a target but we were able to support novel discovery against a novel site against this target that they knew well and was a nice compliment, a collaboration that I also cherish in my memory.

Ross Katz: As we head toward the end of our conversation, I’m just interested in hearing from you, how do you see the success-driven model evolving over the next three to five years?

James Yoder: This is a million-dollar question. Maybe someday a billion-dollar question, I don’t know, there’s a deep enough market for it to be a billion-dollar question. Within hit discovery, we’ve been able to make a lot of impact, but we also acknowledge that we’re just taking this first step on a 10-step progress. There’s a natural extension to deliver chemical matter that’s even more advanced. We’ve never set any in vivo criteria within our success parameters for example in an industry collaboration. Being able to bite off more and more. In 2022 all of our success criteria were specified in cell-free assays, we started moving into cell-based assays, into early ADME as incrementally biting off a little bit more towards what a target product profile might look like. Continuing to onboard those sorts of things I think is exciting for us and maybe we’ll end up being more of a lead discovery house than a hit discovery house. Continuing to get creative with our collaborators on ways to serve really young companies is an area of excitement for us. Right now I think there’s a lot of need for young companies to have resource-efficient discovery services available to them, especially because we’re at a low-water mark in terms of seed and series A funding and funding of research in particular. Trying to find a way to enable new platform companies, emboldening them to launch new projects I think is exciting and something that we can do. There’s a market that we feel like we can create there given the success-driven model. Solving the chicken-egg problem where it’s like okay, we’re not going to be able to raise money until we have chemical matter. We don’t have to pay OpenBench until we have chemical matter. Maybe we’ll be able to find that matter for them that they’ll be able to raise on the back of. Exploring that almost downmarket from where we currently sit is also quite appealing. Upmarket from us as lead discovery and working with larger and more established biotech and pharma companies as well is I think an opportunity. We see increasing opportunity to work with other CROs, too. A lot of CROs have this incumbent fee-for-service business that we can work with them directly. A lot of times the characterization that’s being done in our collaborations is with CROs doing the biochemistry, the biophysics, the structural biology work for confirmation. We depend upon these CROs and working more closely with them is an area of interest. We have a certain moral clarity around sticking to our services mission. Really we want to serve scientists and there’s a lot of AI companies that have gone before us. We think we’re good at hit discovery and that we’re so brass that we’re willing to offer the success-driven model. We don’t want to fall into the same pitfalls that a lot of companies before us have fallen into where they’re like we have an edge on this small part of drug discovery and so now we’re going to develop a therapeutic pipeline. There’s a lot of reasons that venture-backed companies go in that direction because there’s an easy path towards the billion-dollar outcome if you have a safe and effective drug and you’re helping patients then usually there’s a billion-dollar outcome there. Services are a lot less sexy and a lot smaller market generally. We’re okay with that though. That’s the basis on which we founded the company and I think there’s a lot of opportunity for us continuing in the service vein. I think trying to still figure out some of the details along the way, getting even better at risk assessment, getting even better at taking on more custom synthesis. We prosecute these ultra-large libraries, what’s the next step there? Improving our methods to justify synthesis where we’re spending a thousand dollars per compound, which is closer to the order of magnitude you’re spending in a more custom synthetic regime is an area of great interest to us as well. We’re trying to do some methods development there with branching outside of our existing ultra-large virtual screening. More broadly, we really believe in success-driven discovery as a meaningful theory of change for AI in drug discovery. If AI in drug discovery is going to make an impact, what matters is the results, not the technology. That’s very much been our ethos from the beginning, inspired by others who have gone before us who have made that their ethos. There’s a quote from Anthony Nicholls who was a founder of OpenEye, talking about the importance of delivering, when in doubt deliver results, not the tech. I think continuing to plow that furrow will be ultimately fruitful for us.

Ross Katz: Is there a single challenge or roadblock that you see to taking on a lot more collaborations or taking on bigger collaborations and scaling up the work that you’re doing as a services organization?

James Yoder: I don’t think that these are insurmountable but some of the roadblocks today, one is the more ambitious we get the more capital we need. That’s one thing about the at-risk model that we engage: we need working capital. Trying to figure out how to do that in a way that makes sense for ourselves and our long-term goal. There is a little bit of a tension between venture doesn’t usually fund services companies. How we capitalize the business is something that we’re trying to figure out. We’re revenue generating unlike a lot of biotech companies as well so obviously we can bootstrap to a point and that’s a lot of what we’re doing. But at some point it may make sense to take on capital and how we do that is a big question. I think we’re outside of the box for a lot of, definitely for biotech VC and even for tech-bio or tech VC. Trying to figure that out. Right now I see that it’s been a hard market for biotech companies raising money themselves and especially raising money behind research.

Ross Katz: And where can people go to learn more about OpenBench and the collaborations you’ve been a part of?

James Yoder: You can learn more about OpenBench by reaching out to me directly if you so desire on LinkedIn. You can mention that you listened to the Data in Biotech podcast and heard what we’re working on. If you want to check out our website it’s openbench.com or you can email [email protected] to get in touch with me as well.

Ross Katz: For sure. Well James, it’s been great to have you on the podcast. I really appreciate the time and look forward to connecting down the line.

James Yoder: Awesome. Well thank you Ross.

Jason: And that’s it for this episode of Data in Biotech. If you enjoyed the episode please subscribe, rate, or leave a review in your podcast platform of choice. See you next time.

Frequently Asked
Questions

How does OpenBench manage the financial and scientific risk of its "success-driven" model?
OpenBench conducts a free, one-week feasibility assessment for each target to determine its druggability and establish precise success criteria. This upfront diligence, combined with their internal risk assessment models and a 75% historical success rate, allows them to accurately price projects and take on the scientific and financial risk of hit identification.
What kind of data underpins OpenBench's predictive capabilities for drug discovery?
OpenBench trains its global scoring function on high-quality, purified samples tested in primary assays from its own collaborations, alongside pruned and cleaned literature data. This dataset includes both active and inactive compounds, allowing the models to learn global biophysical principles and differentiate promising molecules from non-binders.
How does OpenBench efficiently search the trillions of compounds in virtual chemical libraries?
OpenBench uses an active learning loop within its autoscreen infrastructure. This involves sampling reaction subspaces, performing molecular docking, and rescoring poses with machine-learned methods. They then build proxy models to infer affinity across the entire chemical space, prioritizing high-scoring ligands and stopping when scores saturate.

Need a data partner for life sciences?

CorrDyn helps biotech and pharma companies build the data infrastructure that accelerates research and operations.

Book an intro call