Listen on
Overview
If you ask a general model how many approved drugs target PD-L1, it will tell you three. Grounded in DrugBank’s data, the answer is six, and the model sounds exactly as certain at three as it should at six. Lisa Downey’s argument is that the scarce input in AI-assisted drug discovery has moved. Twenty years ago customers wanted more data. Today every team has data lakes and a capable model, whether Claude, Gemini, ChatGPT, or something built in-house, and what they lack is the layer underneath: facts complete, reproducible, and traceable enough that a team would stake a go/no-go decision on them and defend it to a regulator.
Lisa Downey is CEO of DrugBank, a structured biomedical intelligence platform that began at the University of Alberta, where co-founders Mike Wilson and Craig Knox set out to build Google for bioinformaticians. DrugBank now holds more than 156 million structured data points harmonized across 20-some ontologies, is cited in more than 62,000 papers, and is used by nine of the top 20 global pharma companies. Downey previously built Clarivate’s genomic and rare disease data business and held leadership roles at GlobalData, and has spent almost 20 years in healthcare and life sciences data.
Host Ross Katz and Downey start with where the time goes. For most teams the lost cycle time sits in data prep rather than analysis, and automating on top of a layer nobody harmonized moves the mistakes as fast as the insights. A defensible answer is one a regulator can trace. When a safety signal comes under review, “the model surfaced it” does not settle whether a relationship is causal or correlated, and Downey’s four questions for anyone building or buying a reference data layer start from that test.
Key Takeaways
A model missing half the answer sounds as sure as one that has all of it
The PD-L1 question is Downey’s example. A general model answers three approved drugs. Grounded in DrugBank, the answer is six, three more once off-target interactions are counted. Downey attributes the gap to context rather than to the model, which is reasoning over fragmented, incomplete sources and cannot tell what it is missing. Its confidence at three matches what it should be at six. A grounding layer is built to come back empty when the answer is not there, and DrugBank treats no results as a valid answer rather than a failed query. A general model in the same position fills the gap with something plausible.
Reproducibility is the second test, and Anthropic’s benchmark measured it
If you ask a general model the same question four times, Downey says, you will routinely get four different answers. Grounded in DrugBank, the same question returns the same answer, with every drug, trial, and phase traced back to the record it came from. She cites the benchmark Anthropic ran with its latest model release: a basic biology task, pulling viral sequences from a public NCBI database. The best frontier models, Anthropic’s and others’, returned the right answer as little as 17 percent of the time and gave different answers on repeat runs. Given a deterministic tool, the same models went north of 99 percent. Anthropic’s conclusion, as she relays it, was that the bottleneck was the data infrastructure underneath the model.
Cycle time is lost in data prep, and skipping the prep moves mistakes faster
Speed alone is a trap in Downey’s account, since anyone can move fast and be wrong. For most of the teams DrugBank works with, the lost time goes into building the data fabric: weeks, sometimes months, spent reconciling the same drug, chasing whether a relationship is real, and rebuilding context that should already exist. The risk comes when a team closes that gap by skipping it and automates on top of a layer that was never harmonized or validated. Mistakes then travel as fast as the insights.
In drug development, nobody finds out until it is very expensive. Downey’s goal is fast and defensible, and she places both the cycle-time saving and the quality of outcome in a complete grounding layer under the team.
Depth turns a graph from a lookup into something an agent can reason over
DrugBank uses many open and public sources as starting points or QA checks, and pulling them into a static graph only goes so far. The work Downey describes is structure (hierarchies, ontologies, mappings, directional relationships), automated validation, and the process that keeps all of it maintained. Depth means knowing how strongly two things are linked, by what mechanism, and on what evidence. One customer wanted to know whether failed programs had failed because of the wrong target, or whether a different modality would have worked. Before DrugBank, the team pulled trials from a competitive intelligence subscription, chemistry from a separate database, and pharmacology from another tool, then stitched them together by hand and hoped the entity names reconciled. Their agent was sophisticated and still hit roadblocks. On a connected graph, they could start from any target and follow its pathway, every drug that hits it, the companies behind those drugs, and every trial down to eligibility, outcomes, and adverse events.
Human-over-the-loop curation only works on data connected from the start
DrugBank began with a scientist approving every AI-surfaced finding, which was workable while the scope was drugs and drug targets and set a high bar for quality. Two things ended that model. Science moves faster than a curation team can be staffed against it, and once DrugBank connected drugs to targets, trials, and diseases, checking each output in isolation became impossible, because a connection cannot be approved one row at a time. Experts now drive and oversee systems across large spans of the data instead of signing each slice. The shift depends on coherence underneath: when everything is connected, one expert validating one relationship cascades through dozens of related facts, and a messy dataset would fall apart the moment a human came off each row. The outside check is academic use. At 62,000 citations, errors get caught and reported back, and the internal bar becomes whether a fact would hold up to a reviewer trying to prove DrugBank wrong.
Regulators want receipts, and a causal claim needs more than literature
Downey separates deterministic data, sourced back to a fact with a paper trail, from probabilistic output, which is plausible and predicted but cannot go to a regulator with the same confidence. She cites the first FDA warning letter tied to AI, sent to a manufacturer that had leaned heavily on AI agents to generate its quality documentation. The FDA’s message was that AI does not relieve a company of responsibility for having a qualified human and facts behind every record. Safety is where she sees the distinction decide outcomes. A regulator reviewing a safety signal presses on whether the relationship is causal or correlated, and “the literature mentions them together” does not survive that question. A translational safety team at J&J used DrugBank’s data to separate causal target-phenotype relationships from correlational ones, drawing on mechanism of action and on-target and off-target effects that literature alone, with its fuzzy language, does not pull apart.
Four questions separate data infrastructure from marketing
The first risk Downey sees is scope. Teams build a reference layer in-house so they can do the real work on top, and building the foundation becomes the work while the scientists wait. Whether a team is building or evaluating a vendor, she asks the same four questions. The first is whether the data is AI-ready and harmonized, which for a builder can mean 18 months of crosswalks between Mondo, RxNorm, ChEMBL, and UniProt, and for a buyer means checking whether the vendor did that work or only says it did. The second is whether every fact can be traced back to a source and a date. The third is whether it is maintained or was scraped once and called complete. The fourth is whether it can connect to what you already have, including licensed commercial datasets whose terms may block the use you intend. If it passes all four, in her words, it is data infrastructure, and if not, it is probably marketing.
Ross adds a method for answering them. His approach starts from a benchmark set of the questions a team needs answered and the responses it expects, then prototypes in-house, tries external vendors, and picks the architecture that meets the benchmark on the team’s time frame and budget. Rebuilding DrugBank itself would be a 20-plus year effort, and at that scope his advice is to buy it.
Related: CorrDyn builds the data quality practices that make a reference data layer traceable and maintained for pharmaceutical R&D teams. Also on grounding agents in structured biomedical data: Cody Schiffer on knowledge graphs in biopharma and MCP in the Enterprise: Cost, Security, and Workload Routing.
Full Transcript
Ross Katz: Welcome back to Data in Biotech, I’m Ross Katz. For 20 years, the thing everyone wanted in drug discovery was more data, but that problem is over now. Data isn’t scarce anymore, and neither is reasoning. Every team has a capable model, whether it be Claude, Gemini, or ChatGPT, or something in-house. The scarcity has moved to another part of the equation. According to Lisa Downey, our guest today, it’s the layer underneath the model. The set of facts that is solid enough that you’d stake a decision on them. Lisa is CEO of DrugBank, a connected, machine-readable map of drugs, targets, diseases, and trials that has been curated continuously for 20 years and is cited in more than 62,000 papers. If you ask a general model how many approved drugs hit a given target, it’ll give you the answer with total confidence. Nothing about the way that it answers will tell you what information entered its equation. On this episode, we get into why that gap exists, why speed without a scalable and sustainable foundation is a trap, and what it takes to stand behind an answer that AI gives you when a regulator asks for the source behind it. Lisa, welcome to the show.
Lisa Downey: Thank you. Good to be here.
Ross Katz: Awesome. Well, you’ve spent your whole career on the vendor side of healthcare data with GlobalData, DRG, private equity, a few early-stage startups before landing at DrugBank. I’m interested in what was different about DrugBank that made you want to join the team?
Lisa Downey: Yeah, it’s funny, early on in my career, 20 years ago, data was the only thing customers wanted to get their hands on. More data, more data, more data. There was an abundance of data and lots of vendors out there and all of them being very hand-wavy about why their data was better differentiated, so on and so forth. A bunch of platforms started to emerge, every team had a platform and a login. It quickly got to a point where customers were drowning. Biopharma and MedTech plowing tons of cash into trying to make fragmented data usable. They had fatigue from the platforms, they had a bunch of data lakes, but still couldn’t really make use of the piles of data that they had internally. When I came across DrugBank, it was co-founded by Mike Wilson and Craig Knox, they had built it because it’s what the science needed. They were scientists themselves, started at University of Alberta. I was just frankly blown away. It was a big wow moment for me seeing that they had connected, curated, and made machine-readable this enormous expansive span of biomedical data. I was really excited, that was a bet that I wanted to lead.
Ross Katz: Yeah. Getting into DrugBank, as you mentioned it started in academia and the founders described it as Google for bioinformaticians. If you wouldn’t mind, just set the stage for us, what is DrugBank and then who relies on it today?
Lisa Downey: Yeah, so simply put, DrugBank is the grounding layer for biopharma R&D. The co-founders, as you said, set out initially to build Google for bioinformaticians. What that meant was they took the abundance of drug and drug target data and information from across literature, drug labels, chemical databases, in a dozen or more vocabularies, and connected it, resolved it, made it machine-readable. At that time, 20 years ago when they first set out doing this, it was important. But now in the age of AI, it’s got the urgency to match that importance. That’s how it transitioned away from that academia project effectively into a reference layer underneath pharma’s AI. Structured, harmonized, continuously curated map of not only drugs and drug targets, which is where it started, but now target diseases, trials, commercial context to match that. We’ve got over 156 million structured data points today, all harmonized across 20 some-odd ontologies. We’ve got nine of the top 20 global pharma relying on DrugBank alongside the broader scientific community. We still have academics utilizing our data. We’re cited in more than 62,000 papers now. There’s no shortage of data, no shortage of reasoning layers. What customers come to DrugBank for is that grounding layer that sits underneath.
Ross Katz: Can you unpack a little bit about what it means for DrugBank to be that grounding layer underneath? I understand what you mean by reasoning layer as being ChatGPT, Claude, Gemini, internal in-house LLMs that companies are using, and the agentic harnesses that sit on top of them, but I’m interested in how DrugBank connects as that grounding layer.
Lisa Downey: So grounding is what gives those LLMs the facts. There are two invaluable contributions that I would point to. One being completeness and the other being reproducibility. If your team is working on target identification, pretty high-stakes work, I think we can all agree, complete and reproducible data is really the whole ball game. Let’s say you’re trying to understand how many approved drugs target PD-L1, for example. If you ask a general model, it will tell you three. The real complete answer, when that model is grounded in DrugBank data, the answer is six. There are six approved drugs that target PD-L1, three more when you account for off-target interactions. That disconnect isn’t the model being broken, it’s reasoning over fragmented, incomplete context. The really big concerning piece of that is that it doesn’t know what it’s missing. It believes that it knows. It’s sounding just as certain and as confident and as sure at three as it should be at six. The second one is reproducibility. Ask a general model the same question four times and routinely you will get four different answers. Grounded in DrugBank, you’ll get the same answer every single time. Every drug against the targets, the trials, the phase each one is at, every figure traced back to the record that it came from. When the answer isn’t there, a grounding layer will tell you so. We built it that way so that no return results is a valid answer rather than a failed query. A general model in that case is going to fill the gap with something plausible. That’s a huge liability. One interesting proof point here is Anthropic’s benchmarks with their latest model release, a basic biology task, which was pulling viral sequences from a public NCBI database. The best frontier models, theirs and others because they didn’t just run the benchmark against their model, returned the right answer as little as 17 percent of the time and gave different answers on repeat runs. When those same models were given a deterministic tool, the data accuracy went north of 99 percent. Their own conclusion was the bottleneck was never the model. It’s the data infrastructure that sits underneath it. For R&D, that is the infrastructure we’ve spent 20 years building.
Ross Katz: Right. If we think about what your baseline ChatGPT or Claude or Gemini would do if you asked one of these questions, maybe it would do a series of web searches against the public internet, against academic literature, maybe if you’re using Claude for Science, for example, you’re plugged into an existing database, but there’s a limit to the amount of time it’s going to spend searching through the long tail of academic literature and extracting the right information. A grounding layer like DrugBank, having already, to a high degree of quality, preprocessed all of that information and put it in a format that makes it easy for the model to understand and to ask the right question of, allows it to get to the right answer very quickly and then also allows it to get to the right answer reproducibly over time. Am I thinking about that right?
Lisa Downey: You are. Yes, exactly.
Ross Katz: Yeah, awesome. I want to get into how DrugBank is constructed, but before I do, you mentioned target identification as one example. Would you mind talking about the specific decisions that your customers make using DrugBank?
Lisa Downey: Yeah, and I’ll back up a little bit and just start with why it matters to us. DrugBank’s North Star, our big hairy audacious goal internally, is to get a thousand new therapies to patients. I’ll talk about the money and the time that we save our customers, but that’s not the ultimate end goal. That’s really best thought of as the unlock. Every dollar, every year that a team burns on the wrong call is a therapy that’s going to reach patients much later or possibly not at all. The decisions that we contribute most directly to would be go/no-go decisions. Which target to advance, for example, which to kill before it consumes years and hundreds of millions of dollars on what would be a predictable safety liability. Where white space exists, so you can understand where to place a bet versus where the field is already crowded. For a platform biotech, which validated target actually fits their modality, because every high-conviction target a competitor picks up is a program that they have lost. Or for a business development team, we work very closely with them on which asset in a crowded class is the diamond worth in-licensing and which one only looks good because you’re missing half the picture. What all these things share is that being wrong can be catastrophic. Being slow is nearly as bad. It’s well documented that getting a drug to market can cost in excess of a billion dollars, it takes a lot of time, only a small fraction of candidates ever make it. The job for our data is to let the team make that call on a full landscape rather than just a slice that happened to be in the tool that they had open and their teams are regularly working with, and to make it faster with the evidence that they need to be able to defend it. Kill the wrong things sooner, back the right thing with conviction.
Ross Katz: Yeah. As we were talking previously about DrugBank, you mentioned that one of the things that DrugBank does is reduce the cycle time of decisions that need to get made. I’m interested in where that lost cycle time typically comes from and how DrugBank enables that cycle time to get reduced.
Lisa Downey: Yeah, so I will say, I believe that speed alone is a bit of a trap. Anyone can move fast and be wrong. The goal is to be fast and defensible. The good news is for most teams that we’ve been working with historically, the cycle time’s not lost where they think. It’s rarely the analysis. Usually it’s in the data prep, the building out of the data fabric, and trying to stitch together a good foundation before anybody is able to even ask the questions. We see it constantly. Teams are spending weeks, sometimes months, reconciling the same drug across five different languages or trying to chase whether a relationship is real. Rebuilding context that should already exist. A decade or two ago, as I said off the top, the bottleneck was volume of data. Not anymore. It’s definitely more about collapsing that distance between data and decision. Where that pressure to go faster creates risk is when a team tries to close that gap by skipping it. We see that too. Automating on top of a layer that was never harmonized or validated. That’s where we start to see that speed on that shaky foundation just means the mistakes are going to move as fast as the insights. In this industry, you don’t find out until it’s very, very expensive. The real unlock isn’t just strictly pushing a team to go faster, it’s putting a complete defensible grounding layer underneath them. That’s where you get both cycle time and quality of outcome.
Ross Katz: Yeah, and I can imagine the contrast here being the homegrown knowledge graphs, where people say, well, we’ve got LLMs now and we’ve got access to all of the same literature that you have access to. I’m interested in understanding a little bit more about what differentiates DrugBank from both your competitors in the space and then the other homegrown knowledge graph type systems that would allow people to create the semblance of a grounding layer that might give them the appearance of moving fast, but might make it so that they’re making decisions that are somewhat less defensible.
Lisa Downey: Yeah, and I’ll start by saying that there are a lot of fantastic open and public sources and many of those we use ourselves as starting points or QA checks in some cases. But pulling those sources together into a static graph only gets you so far. The real value is in building out that truly curated data infrastructure with, and this is a big key part of it, the process to continuously maintain and update and refresh it. Once AI agents are reasoning over these knowledge graphs, that becomes even more important. What you need is for your knowledge graph to be two things, which might sound obvious, but the first one is usable. The second is trustworthy enough for the agents to reason over. I’ll get into a little bit of the layers underneath, which most people probably would not think of or certainly appreciate until you are in the throes of trying to do it yourself. First is the engineering and the structure. That’s the hierarchies, the ontologies, the mappings, the directional relationships. Then there’s the validation and review. Automated QA, deep linking, cross-checking. Depth is what turns the graph from a lookup into something an agent can reason over. It’s not just that two things are linked. It’s how strongly, by what mechanism, on what evidence. It’s the difference between an answer you can act on and one that you’re going to have to go and re-check. I’ll illustrate with a real example to make this a bit more tangible. We had a customer who came to us and they wanted to look at programs that had failed and ask, did it fail because it was the wrong target? Or would a different modality work on it? We love that kind of question because you can’t brute-force it. You have to take the trial outcome, connect it to the drug, to its modality, to its target, and to the biology around it. Before they had come to us, they were pulling trials from a competitive intelligence subscription, they were using a separate chemistry database and pharmacology tool, they were having their team manually stitch it together hoping that all the names and entities reconciled. They couldn’t just point their agent at it. They had developed a very sophisticated agent, but they kept coming up against roadblocks. When they came to us and they saw our connected graph, and this gets into why the connection is so important and the depth of it is so important, they could start from any target, see the pathway it sits in, every drug that hits it, the companies behind them, every trial across every disease down to the eligibility, the outcomes, the adverse events. That was it for them. We had collapsed it all into one layer that their agent could properly reason on and not just do lookups.
Ross Katz: Right. You can jump from the target to the modality and the mechanism to the evidence supporting it and the source of that evidence, all in one go. You’re able to make the connections that these complex types of questions, like did it fail because of the target or the modality, raise, while also opening it up so that you can see the evidentiary support for the reasoning and the response that you’re getting back from the AI agent, and you can weigh it yourself versus having to take the AI agent’s response as being true on its face. Am I thinking about that right?
Lisa Downey: Exactly. The benefit is too, it is this complicated web, but you don’t have to start from the target either. You can start from the chemical structure, you can start from if you have the right kind of connection and you’re not going to get bungled up by different ontologies that don’t speak to each other, don’t properly harmonize.
Ross Katz: Yeah, so you mentioned making it usable and trustworthy through both the engineering structure and the automated QA. Would you mind talking about both of those elements, the engineering structure and how it’s created, and then the automated QA and how humans are involved?
Lisa Downey: Probably the best way to talk through this would be our methodology in how we started versus where we are right now. Where we started was human-in-the-loop. That was the right model for where DrugBank started. Our foundation was drug and drug target data. Relatively speaking, a fairly narrow and well-bounded scope. When the domain is that contained, you can have a scientist approve every AI surface finding and it’s completely workable and it sets a super high bar for quality. That worked initially. As we grew, the two things that caused us to really rethink that model were the pace at which science moves. You really can’t keep up to the science and staff against it in a way that is scalable and repeatable over time. The second piece of it was, and this was the bigger one for us, as we broadened our scope outside of strictly that drug and drug target focus and we went into more of a cross-context data type of approach, and again connecting the drugs to the targets to the trials to the diseases. Checking each output in isolation doesn’t just get slow, it gets complex to the point of being impossible. The value is truly in those connections and you can’t approve a connection one row at a time. That’s what led us to do what we’ve continuously done over the course of time, which was disrupt ourselves. The only reason we could is because of that initial foundation. The original drug-drug target data was never a pile of data, it was deeply curated and connected from the start, from the beginning. The graph itself does part of the validating for you, and when everything is connected, one expert validating one relationship cascades through dozens of related facts. You can only move to that kind of human-over-the-loop model if everything underneath it is coherent enough to trust that cascade. A messy data set would fall apart the moment you took a human off each row.
Ross Katz: So when you talk about human-in-the-loop versus human-over-the-loop, can you talk a little bit about the differences between those two things and how that manifests in the processes and the technology that you’re using internally?
Lisa Downey: The way that we’ve transitioned it is from using our proprietary method of automation from the early days of pulling the drug and drug target data and information, we had our curation team set up so that they were looking at individual slices of the data. There was human signature on every data point. That was manageable. The team when I was first looking at DrugBank was quite lean and nimble compared to many others out there. It was really about pulling in the right sources and being able to look at them effectively and having the right scientists overseeing, sitting within in-the-loop, so it wasn’t strictly machine just running away with it. The human-over-the-loop does require what I would describe more as experts driving systems as opposed to doing the validation within the loop. Overseeing and orchestrating and making sure that nothing is broken. There is still human signature on it, but, as I was describing, it’s propagating across many different views of the data instead of sitting in a narrow slice. More concretely, we don’t have our curators just looking at trials data, for example, or just looking at specific ontologies, to overseeing large spans and pieces of the data to make sense of it and to make sure that there is true harmonization taking place there.
Ross Katz: How do you assess the quality of the system that you’re creating and the extent to which the human-over-the-loop system is maintaining the high standard of quality that you’ve set out for it?
Lisa Downey: So we have methods internally, but frankly I think I’d point most to the academic community still using our data to the extent that they are. We’re working deeply with pharma and biotech. We still have a significant and meaningful footprint within the broader academic and scientific community. We’re now at 62,000 citations, if anything that’s accelerated. For us, that’s really the kind of feedback loop that most commercial data sets certainly do not have. It means that every error is going to get caught and comes back to us and gets fixed. It compounds, and that really leads to a lot of the discipline internally. When you know your data is going to show up in peer-reviewed work, good enough isn’t good enough. Our bar then becomes, will this hold up to a reviewer who’s trying to prove us wrong? That’s a standard that our customer inherits when they build on us. I think that’s the main thing that I would point to. Yes, we have our own QA and we do the manual checks to make sure that what we’re doing is valid and it’s making sense and we have systems in place, but that is a really important overlay and stress test to our data.
Ross Katz: Yeah, that makes sense. I’ve heard you draw this line between deterministic and probabilistic data, where the stakes of the decisions that your customers are making using the data that you’re providing to them require data that is deterministic rather than probabilistic. Can you lay out the distinction that you’re making there and explain why you think it matters?
Lisa Downey: Yeah, so I think it matters in many settings, but I’ll say within this one, within a regulated industry, whatever gets submitted to regulators, it’s not just about the final confident-sounding paragraph and deliverable. You need to have the paper trail to back up said deliverable and what’s getting submitted. Deterministic data, data with receipts sourced back to a fact, is something that you can submit to regulators confidently. Probabilistic can’t. It’s plausible, it’s probable, it’s predicted based on the massive amount of data that it’s overseeing. Regulators are already starting to act on this. There have been subsequent ones, but I recall the first FDA warning letter that was issued earlier this year that was tied to AI. It was to a manufacturer that had leaned heavily on AI agents to generate its quality documentation. The blunt message that the FDA issued was, look, AI doesn’t relieve you of responsibility. You still have to have a qualified human and facts and data to stand behind every single record. Where deterministic data would change the outcome, take safety, for example. We work a lot with safety and pharmacovigilance teams. When you bring a safety signal to a regulator, the question they’ll press on is whether the relationship is causal or just correlated. Saying the model surfaced it isn’t obviously a satisfactory answer, or the literature mentions them together, that’s not going to survive that question. We had a translational safety team at J&J Innovative Medicine that we work really closely with and they used our data to specifically draw that line, separating causal target phenotype relationships from correlational ones. That’s exactly what a model that learned the world by reading scientific literature could not do and could not give you, and it’s exactly what decides if a safety signal is going to hold up. When a regulator challenges the decision, it can’t just simply be what was generated by a model, even if it has a lot of data contributing to it. You have to be able to say, here’s the source, here’s the evidence, here’s exactly where it came from.
Ross Katz: Yeah, since you mentioned the J&J study, if you’re willing, I would love to have you unpack what it was about DrugBank’s data that enabled them to draw the distinction between correlation and causation, and why the DrugBank data set and grounding layer is so effective for that use case.
Lisa Downey: It comes down to the knowledge graph again and the inputs to it. Many other reasoning layers out there in addition to general LLMs use literature alone as the primary input. In literature comes a lot of fuzzy language that doesn’t really disseminate and pull apart causality from correlation. The knowledge graph that we have built does, because it has all of the other inputs and data points, so you’re not just strictly looking at the correlation, you can see causality by pulling in the mechanism of action and you can see the target effects, the off-target effects. Being able to do those additional connecting of dots around other data points and sources rather than the flat literature alone is what enables our customers to be able to get deeper and to be able to distinguish that fuzzy language into what’s true, what could be backed up, what’s evidence.
Ross Katz: Yeah, so the evaluation of causation versus correlation is ultimately a judgment call that needs to weigh the full body of evidence that has been surfaced in the academic literature. If you’re only looking at an AI interpretation of the conclusion of a given paper, and you’re asking it to do all of that evidence weighing on the fly as it’s researching a lot of other things, it’s impossible to know the quality of reasoning that was applied to the particular question of correlation versus causation, and certainly your ability to track it down and reproduce it yourself is limited as well. Am I thinking about that right? Is that accurate from your perspective?
Lisa Downey: Yeah, and that’s where I would go back to the reproducibility angle and going back to asking the same question four times. If you’ve got the DrugBank grounded layer, it’s going to give you the exact same answer with the evidence to back it up, the paper trail, the facts, versus if you’re asking an LLM without a grounding layer, that’s where you’re going to get the probabilistic answer, which may be a different answer every time you ask that same question. It really comes down to what the model is synthesizing. Is it structured data that it’s reasoning over or is it a whole pile of text and it’s forgetting segments and fragments of the sentence when it was doing that review?
Ross Katz: Yeah, and also the structured thought and process and review that went into distilling down the relevant and most important bits of information and structuring them in a way that makes them searchable, so that remembering or forgetting particular words, or particular aspects of the research, or attending to the correct aspect of a given paper, does not matter for the purposes of a given question that’s being asked. At least that’s my understanding. Would you add to that at all?
Lisa Downey: Yeah.
Ross Katz: So I’m interested in how teams are using DrugBank today. How do they expose DrugBank to either large language models that they’re building or fine-tuning in-house or to the reasoning layers that we’ve talked about previously?
Lisa Downey: Yeah, so that’s one of the things that we have been very deliberate about, and that is being extremely interoperable for customers. We have MCP connectors, you can plug us into Claude, into Copilot, into ChatGPT. Some of our customers have their own LLMs and that’s how they want to work with the data. We have other customers who simply want us to deliver so that they can host it on-prem, and they’ve got access to CSV flat files, and they get a catalog of SQL queries so that they know how to maximize utility and value of DrugBank. The way that it’s deployed really depends on the digital maturity of the team and the size of the team and how broadly they want it to be embedded across the organization. It’s fairly flexible, interoperable by design. We’ve started to create a catalog of MCPs, including DrugBank-aware MCPs, because some of the feedback we got from long-standing customers pre-MCP era, where they were getting the flat files, they’re saying to us, we don’t know what we don’t know. We need to understand better how we maximize the value. We don’t know the edges, we’re making some assumptions. We’ve done a lot of work to make sure that we can provide the maximum utility and value across organizations.
Ross Katz: Yeah, so MCP is, I mean, we’re on Data in Biotech so I assume most people know what MCPs are, but MCP is Model Context Protocol server, and so it’s a way of surfacing tools, including tools for accessing data like DrugBank, to large language models, especially the foundational models that most people are using today, which are very familiar with gathering. I’m interested, from the perspective of your customers, in how they think about when to use the MCP connection to DrugBank versus when they choose to bring the structured data into their own ecosystem. Is it that when they have the structured data in their own ecosystem, they’re trying to build out the knowledge graph to include elements of their own internal research, or what are some of the use cases that you’re seeing for why people choose one or the other?
Lisa Downey: I think in addition to the digital maturity angle, it comes down to security as well. One of the technologies that we have is called UltraLink and it’s a dynamic mapping tool. It is something that, when quote unquote let loose inside our customer’s four walls and they have our data on-prem, they can deploy internally to then have DrugBank be the infrastructure that they build on top of. If they want to make connections between other data that they have, that’s what they will use and leverage to do that. It really comes down to preference. Many of them have security considerations and want to be certain about what’s going to be surfaced, what’s not going to be surfaced. I will say, importantly, we’re very blinded when it comes to what our customers are doing. Obviously none of their data gets surfaced on our side, there’s no concern there. We don’t even check their queries, nothing like that. For them it really comes down to how they’re going to maximize utility. If we have a smaller team, so we work with what I would deem de facto AI R&D teams, in those cases, typically they have agents that they’ve constructed, and so they will have their agent query our MCP, but it will fetch from our connector, it will use their internal data. Then we have other customers, like, we work with Flagship Pioneering Intelligence, they use DrugBank to ground their data exchange. It’s bolted in there and then they’ve got data that they’ve built on top. There are mixed flavors, I would say.
Ross Katz: For biotech companies that are deciding whether to build their own reference data layer or bring in a partner, what are the sorts of questions that you would have them ask themselves before committing a team to a project like that?
Lisa Downey: Yeah, so I love this question. We get it often. I think the first thing I’d say is be careful how you scope it. Because teams think oftentimes that they’re building the reference layer in-house so they can do the real work on top. What we often see happen is building out that foundation becomes the work, and they never get a chance to have their scientists do the building on top of it. That work continues to get pushed out. Whether you’re building in-house or you’re looking at a vendor, I pretty consistently would ask the same four questions. The first one is, is the data actually AI-ready and harmonized? If you’re building, that means, are your people going to be spending 18 months building those data crosswalks between Mondo and RxNorm and ChEMBL, UniProt, etc.? If you’re buying, you want to understand if the vendor has actually done that work or if they just say that they have. The second is, can every fact be traced back to a source and to a date? If you can’t trace it, you can’t defend it, and that’s true no matter who assembled it. The third question I’d ask is, is it maintained? Or was it scraped once and called complete, and does it have any consideration for being able to add more data as the science moves? The fourth question is really, can you connect it to what you already have? In some cases, customers have other commercial data sets that they’ve licensed, but the license restricts them from being able to do what it is that they’re setting out to do in building this reference layer. Those are things that you want to check. I would say, if you ask those four questions to your team internally with your capabilities to be able to build it, or anyone who’s trying to sell to you, those are going to be mission-critical, all four of them. If it passes, that’s data infrastructure, if not, it’s probably marketing. I think those are probably the most complete set of questions that I would ask.
Ross Katz: Yeah, I love those questions. I would just also add, as the head of a team that builds data infrastructure and AI integrations as well, that fundamentally it’s about what is the use case, and you mentioned the scope question. Obviously if you’re trying to rebuild DrugBank from scratch, that’s a 20-plus year effort with more resources than you’re probably willing to invest in building it, you should really just buy it. But if it’s a narrowly scoped effort and you know what questions you’re trying to answer. I mean, this is really the thing that we see, that people have an idea of generally what they’re trying to build, but specifically the questions they’re trying to answer and the quality responses they’re trying to get, starting with a benchmark data set of this is what we’re driving at, and then you have the ability to do a prototype in-house, try external vendors, discover what is the solution architecture that gets you to the responses that you expect on the time frame and within the budget that fits your business. I think that if you start in a very concrete way like that, in what I would call a scientific way about doing this kind of discovery, then you’re going to end up in a much better place than just trusting, this is what we think we can do for how much and this is what these vendors say they can do for how much. What do you think about that?
Lisa Downey: Well said.
Ross Katz: Awesome. Comparing DrugBank to the publicly available options and the other vendors in this space, where would you say that DrugBank is the best fit for customer problems and where would you point people toward other vendors or other solutions?
Lisa Downey: Yeah, and this is an answer informed by what our customers repeat back to us, because we often ask them. We had very recently one of our customers say, look, I don’t need an all-singing all-dancing thing from every single one of my vendors. Any vendor who tells you they could do all the things, you should probably walk away from. We’re the grounding layer. Drugs, targets, diseases, trials and the context that connects them, structured for an agent to reason over. We don’t do structure prediction, for example, we don’t do market access or forecasting, and you mentioned your own business. Nobody should do this alone. You need to have help. You need to have the folks that have been there, done that before. This is their core competency. If you need help designing the stack or running the transformation of bringing in a reference layer, for example, that’s exactly where strong consultants and AI partners can help move those efforts along and keep them on track, according to what those outcomes need to be and what you’ve stated. Again, what we provide is the grounding layer that they and their teams can build on, getting to defensible options, and the judgment staying that of our customer. I brought up previously a little bit about the security and protection of data, and I do think that that’s a really worthwhile question for our customers to be asking of vendors, especially today, which is what happens to your data? There’s real distrust in the market right now, and increasingly so, for where your data goes once a vendor has it. For us, it matters that ours runs where your data already lives. Your compounds, your results, query patterns, etc. They never leave your walls. We don’t want your data, we want your trust. If you’re wanting to have a different type of working relationship, DrugBank probably isn’t it.
Ross Katz: Yeah, that makes a lot of sense. As we head toward the close here, looking ahead, you’ve called AI the ultimate amplifier of both good work and of mistakes. As more of the industry is wiring AI into discovery and manufacturing, where do you expect that amplification to help the most and where would it worry you?
Lisa Downey: Yeah, it’s funny, we see this, we live this every day. It really is the ultimate amplifier. It’s going to amplify the good and the bad with equal efficiency. The way I’d answer this, I think that it starts with where the industry is on the curve. For a few years, early days, pharma was treating, and many other industries, but pharma certainly was treating AI as a tooling problem. Trying to get models and co-pilots into folks’ hands and expecting productivity to follow. Then it became an efficiency story, sometimes a cost story, eliminating headcount, things like that. Where we see the more advanced organizations really starting to excel is realizing that the true impact is only going to be realized with an operating model change. They’re breaking down data silos, we’re seeing scientific teams and platform teams starting to sit together where maybe they hadn’t historically, re-architecting so that AI can actually operate. That re-architecting of the data, the workflows, the teams, it’s really hard specialized work. It’s a discipline, as you know all too well, of its own, and why partners do this kind of transformation and do it well, and those partners matter to a great extent right now. Our role through that is to be the stable layer underneath while all of that is in motion. Getting back to your question about where this amplification is going to help most. It’s where AI removes the friction between scientific talent and usable data. If you can keep your brilliant scientists doing the brilliant science, that’s where you’re going to see an amplification of the people, of the talent and not just the compute. Where it worries me is that same force pointed at an unstable base. The development of all these tools and agents and additional reasoning layers, if they’re pointed at an unstable base, that’s where you scale poor data practices, weak governance, siloed orgs, flawed assumptions. On a shaky foundation, it just makes mistakes travel further faster than the insights. The key will be to reorganize thoughtfully around your data, your workflows, your decisions, and give AI something solid to stand on. Those are going to be the teams that are ultimately going to net the biggest returns and yield the highest results in this kind of transformation. Look, I realize these transformations are really, really disruptive and can be super difficult and challenging for these teams. I think if you’re going to choose one thing to try to keep stable, it has to be the ground truth underneath. That’s really what it comes down to. In a reorg, that’s what’s going to keep everyone oriented while everything else is moving around it.
Ross Katz: Yeah, I think in this era of everything speeding up, of the attraction of being able to ask a question in natural language to an agent and get an answer that seems plausible, it’s easy to get trapped by the illusion of moving fast and just taking a shortcut. What you’re highlighting is just the importance of building on that stable foundation and doing the work, or working with people who have done the work up front, that allows you to build from there. Lisa, this has been fantastic. Thank you for joining the podcast. Really appreciate it.
Lisa Downey: Thank you, Ross. It was a lot of fun. Thanks for having me.
Ross Katz: That’s my conversation with Lisa Downey. The most dangerous failure in agentic workflows is this unsourced confident answer that we typically get. The reason why it’s dangerous is because it’s impossible to know what has been omitted from the answer that you’ve been given, and without the source of the information, it’s impossible to verify the confident answer that you’ve received. You have the opportunity for both type one and type two errors. Her example was you ask a general model how many approved drugs hit PD-L1, and it says three, but the real answer is six. That is in many ways worse than a made-up answer. Every drug that it gives you might be a real one, nothing on the page might be wrong per se, but there might be three additional drugs that are left out. Also, without the ability to verify the source or to trust the source from which it came, then the answer is effectively worthless to you. Running with an answer isn’t going to get you anywhere. The problem here is that AI has made it more and more difficult to smell a bad answer. We’re used to looking at an answer and saying, oh, that sounds right, or, I think that holds up because all of my follow-up questions get me what I would expect. But these models have been trained to be fluent, and they sound just as fluent when they’re wrong as when they’re right. Those instincts that we had for evaluating things on their surface are no longer as valuable as they once were. I think the counter-intuitive thing here is that the better the models get, the more the curated data like DrugBank increases in value. There’s this idea that if models know everything, then it’s effectively a database and you can just ask the model and it’ll give you the right answer, or it’s going to be able to boil the ocean of the internet on the fly, but I think that’s computationally unrealistic. The value then becomes what context, what foundational grounding layer can you give to the model that allows it to get to the right answer much faster and with sources that you can verify as you’re asking the question. For me, as someone who leads an organization that builds agentic AI systems, I think the question is where does your organization’s data, or the things that you build custom as part of that grounding layer, add distinct value, and how do you plan to combine that with the other systems and sources out there that have done the work for you to drive that value in the future. I just want to thank Lisa again for joining the podcast. I’m Ross Katz and this has been Data in Biotech. I’ll see you next time.
Jason: And that’s it for this episode of Data in Biotech. If you enjoyed the episode, please subscribe, rate, or leave a review in your podcast player of choice. See you next time.







