Listen on
Overview
Biopharma companies have traditionally grappled with significant knowledge work overhead, often dedicating 50-80% of analysis time to merely collecting, integrating, and structuring data before any insights can even begin. This manual, linear approach limits the scope of inquiry, drives up costs, and can lead to missed opportunities in a rapidly evolving market. The emergence of large language models (LLMs) offers a fundamental shift, transforming unstructured scientific literature and internal documents into an interrogable data modality, making it feasible to automate much of this foundational data work.
In this episode, host Ross Katz speaks with Alex Telford, Co-Founder and CEO of Convoke, a company building an AI operating system for biopharma. With his background in life sciences consulting and biochemistry, Alex saw firsthand the inefficiencies plaguing drug development and market analysis. He’s joined by Majid Ahmed, who brings extensive experience in large-scale AI and data infrastructure from the automotive autonomy sector. Together, they discuss how Convoke is applying LLMs to automate knowledge work, enabling biopharma to move from static, linear processes to continuous, iterative decision-making.
They explore the technical challenges of building a rigorous semantic layer over millions of diverse documents, managing compounding errors, and distilling massive analytical outputs for human consumption. The conversation highlights how this approach allows companies to canvas entire decision spaces, asking previously cost-prohibitive questions, and ultimately driving a more data-driven, ‘Moneyball-ification’ of the industry.
Key Takeaways
LLMs Redefine Biopharma Inquiry, Beyond Simple Efficiency
LLMs dramatically reduce the cost of data collection and structuring, enabling analyses that were previously too expensive to consider or impossible at scale. This capability moves companies from limited scope reviews to complete explorations of all potential targets or therapeutic approaches. The real gain is not merely time savings, but the ability to identify novel opportunities and avoid missing significant upside in a competitive market.
Pre-processing and Semantic Layers Are Critical for Production LLM Systems
While LLMs are powerful for on-demand queries, deploying them at scale for complex domains like biopharma demands extensive preprocessing. Building a well-tested semantic layer and knowledge graph over millions of documents allows for thorough and accurate retrieval, overcoming the limitations of general deep research agents. This upfront data engineering investment mitigates compounding error rates and bad data quality, which can otherwise derail large-scale deployments.
Surpassing Information Generation: The Next Hurdle is Information Consumption
As LLM systems generate analysis at unprecedented scale, the new bottleneck becomes the human capacity to consume and act on this volume of output. Companies must build new systems and interfaces to distill thousands of reports or analyses into actionable insights for decision-makers. This shift requires re-thinking workflows beyond simple report generation, focusing on continuous decision-making and automated distillation.
Trusting AI: Source Control and Iterative Feedback are Non-Negotiable
Gaining trust for AI-driven insights in biopharma demands clear provenance and control over information sources, as regulatory and internal standards often dictate acceptable data. Systems must provide full traceability from output back to specific document snapshots, alongside mechanisms for iterative human feedback and critique. This approach allows users to guide the model and build confidence in its results before broad dissemination.
Related: CorrDyn helps companies in biotech and life sciences develop strong AI strategy and deploy scalable data engineering solutions. For more on maximizing data insights, see our blog on realizing value in biotech manufacturing.
Full Transcript
Jason: Hey everyone, this is Jason, producer of Data in Biotech. I’m excited to announce that this podcast is now an AAPS media partner. Join us on-site at PharmaSci 360, November 9th through to November 12th in vibrant San Antonio, Texas. We’ll be interviewing leading scientists about the data trends that span the pharmaceutical development and manufacturing spectrum. Data in Biotech thanks AAPS for this opportunity to share emerging science with the global pharmaceutical community.
Majid Ahmed: Large language models let you work with text like scientific literature in a way that was previously not possible.
Alex Telford: Think of it a bit like a building. When you’re building a building you have the scaffolding around it, and you need the scaffolding to help the model get to the right answer. But we shouldn’t be precious about the scaffolding we’re building.
Jason: Welcome to Data in Biotech, a podcast from CorrDyn where we explore how companies leverage data to drive innovation in life sciences. Every two weeks, we sit down with an expert from the world of biotechnology to understand how they’re using data science to solve technical challenges, streamline operations, and further innovation in their business. Here we go.
Ross Katz: Alex Telford and Majid Ahmed, welcome to the Data in Biotech podcast.
Alex Telford: Thanks for having us.
Majid Ahmed: Happy to be here.
Ross Katz: Awesome. Well, just to get us started, would you mind introducing us to each of you and giving us an overview of the work that you’re doing?
Alex Telford: Sure. I’ll go first. Alex Telford, one of the founders here at Convoke. My background’s in life sciences consulting. Before founding Convoke, I studied biochemistry at university, and then I worked at a boutique life sciences firm called Charles River Associates for almost seven years, helping companies with strategy, business development, competitive intelligence, med affairs, the whole gamut of work you can do from taking your drug out of discovery into the clinic and onto the market. Got pretty excited about language models and the potential they had to rethink knowledge work in the life sciences. That’s when I teamed up with Majid and my other co-founder Vikas, who’s not here today.
Majid Ahmed: My name is Majid. My background’s in technology and startups. I was previously at a company building AI and data infrastructure for the automotive industry, particularly in autonomy. Spent six years there. Before that I was at Google, then Google Cloud. Then was excited about applying some of the things I saw particularly in building large-scale data infrastructure to biology and the biopharma industry. When I met up with Alex, got excited about working on Convoke together.
Ross Katz: Awesome. Well, will you introduce us to Convoke and the work that you do?
Alex Telford: Sure. We describe ourselves as an AI operating system, but to make it more concrete, we’re more like a workspace. You can log into the platform, upload documents, internal documents. This could be PDFs, PowerPoints, scientific papers. Really agnostic to format. Then we combine that with documents we’ve sourced from the public web and APIs we connect to like clinicaltrials.gov, PubMed databases and so on. We bring that all together into one unified environment. Then we allow you to essentially do work on that. Use language models to write documents, to analyze competitors, to create landscape reports, a mix of document drafting capabilities and also structured data sets that we’ve compiled. We have a set of our proprietary databases of clinical trial outcomes, of competitor drug statuses, indications. A really broad set of capabilities, and we’re trying to build something akin to the Bloomberg for the life sciences industry.
Majid Ahmed: Our broader ambition is to automate all of the knowledge work that goes into bringing a new drug to market. We’ve raised our seed round, we announced it recently where we raised nine million from Kleiner Perkins and Dimension. We’re working with a broad set of customers including top 30 pharma to small biotechs to scientists that are just doing the initial thinking about commercializing the research they’re working on.
Ross Katz: Can you give us the story of how you decided on this particular approach to using language models in the biotech space, or how did you determine what use cases you were going to go after and how you were going to attack those?
Alex Telford: When I was in consulting, I felt like my teams were spending an inordinate amount of time just pulling information from one source or another, integrating it together before you could even start thinking about getting to insights or informing a decision. It seemed like there had to be a better way than going to clinicaltrials.gov to get the latest trials, and going to the competitors’ websites, and then pulling some patents from a different source, and integrating all these things you might find on Evaluate or Clarivate. We would spend maybe 50 to 80 percent of the time of doing this analysis on just collecting data and munching it together and doing the manual entity resolution between indication names and drug names. We did this for every customer we had. Every biotech company or pharma company goes through the same development process, regulatory process to get their drug approved. There’s lots of similarities in the types of questions that companies are asking and the analysis they need to do to inform decisions. I just felt like we have this huge amount of manual work that goes into informing any decision that felt extremely inefficient, plus this very common set of analyses or work products that need to be created for each customer. There was never any way to build software around this until language models started to get more capable and you could actually build these flexible tools to both collect data and structure data, but also to generate outputs directly. You could build this whole end-to-end process going from data to deliverable. That’s what got me excited about moving out from consulting into the tech space.
Ross Katz: Majid, coming from an outsider’s perspective, how did you look at this problem in a unique way, or what did you feel like you could bring to the table in terms of attacking this set of use cases differently than other organizations were attacking it?
Majid Ahmed: As I was learning about the industry and about biotech more broadly, a lot of that was reading Alex’s blog, reading other resources out there. The two things that were really interesting to me: one, in autonomy development we saw very clearly the players that became the most advanced built this integrated data feedback loop where they were able to take the data they were generating from their simulation testing, in silico testing, from their labs both vehicle-in-the-loop, hardware-in-the-loop labs, and from their commercial fleet, actual customer vehicles that are out on the road. Take all of that data and feed it back automatically into a next iteration of the models and the software that’s then rolled out to vehicles automatically. It was very clear this same thing was happening in biology and in biopharma development broadly where you want to take your in silico testing data, your wet lab data, your animal model data, clinical testing real-world data, and use that in an automated fashion to move from a linear development process into an iterative development process for your therapies. One of the things I was thinking about as I was trying to think about where could I apply myself in the industry is where can you support customers on that journey? What are pragmatic ways where you can add value today that support customers moving towards that development model? The other thing along those lines that was really interesting to me was it was very clear large language models let you work with text, like scientific literature, in a way that was previously not possible. You are now transforming text into a data modality that can be analyzed, can be interrogated in a way that you could previously work with tabular data. You’re unlocking, and in life sciences where a lot of the underlying information is text, is papers, press releases, other text data, you’re unlocking this capability that was previously not possible. Those two things were very exciting to me and kind of led into these broad strokes, and then obviously talking with Alex and his experience at CRA and his experience through consulting led us to narrow in on a few use cases where we can add the most value today, which is med affairs, competitive intelligence, business development.
Ross Katz: That’s really interesting, and I’m trying to visualize for those use cases, the market intelligence, the competitive intelligence, the document generation, the types of questions that customers are asking where they need to gather this external data and then layer it on top of their internal data. That all makes sense, but then looking at moving from the linear feedback loop to that iterative feedback loop that you’re talking about, Majid, how does Convoke facilitate making those processes more iterative?
Majid Ahmed: I can give you a concrete example, but we think this will be a process that plays out over many years. You have to be very pragmatic about taking on small tasks that add value to our customers today. One concrete example is improving study design. Being able to do trial design in a way where you’re looking at all the historical data that has been previously published within a specific disease area or an indication, and using that to make quantified decisions on how can I design my trial to give me readouts that give me all of the sufficient information I need at the lowest cost possible while giving me the highest PtS, probability of technical regulatory success. Our customers are able to do that in a bit more of a quantified fashion rather than doing it based on experts that are going off of instinct and what they’ve seen previously.
Alex Telford: One thing that’s interesting about the models is you can now collect data and structure data much more cheaply than you could before. It’s still relatively expensive to go and get a language model to say process everything you could find on PubMed, but it’s actually feasible for a company with our budget to do that. Whereas if you paid a consulting firm to do that, it would be hundreds of millions of dollars to create that data set. Because you can create these new data sets, you can ask questions more cheaply that would never be worth asking before. I could ask things like what’s the correlation between ALT elevations and nausea in every phase three psoriasis trial? That’s maybe not a very interesting question, you wouldn’t want to pay a lot of money to run that analysis. But because the data is already structured, you can now run that analysis in a few clicks. Because you can ask questions that would be too expensive to ask before, you can hopefully find new insights.
Ross Katz: Interesting. I want to open the hood and get into the technical details of how Convoke processes a question like that, but I’m curious, in addition to what you were saying about maximizing the effectiveness of trials or minimizing the cost or improving the practical executability of a given protocol, are there particular use cases, questions that you envision customers will ask that you’re already seeing over and over or that you expect to see over and over?
Majid Ahmed: We’re seeing a few questions. One is, as Alex mentioned, understanding for a particular set of studies, maybe for a disease area, what are all the clinical outcomes, biomarker, safety issues that have been reported, and are there correlations in that data across different timeframes? That can be an interesting analysis for many downstream questions and decisions that you would need to make. We’re also seeing repeated questions around exploring targets. As you’re thinking about indication expansion or target exploration, exploring that in a structured way of can we do large-scale analysis for each target of what is the strength of the clinical data, what is the strength of the pre-clinical data, what is the genetic association evidence for this target and a particular disease area? Those are some things we’re seeing repeated pretty regularly.
Alex Telford: One question we worked on recently, a concrete example, is are there opportunities for us to run shorter trials in certain indications? If you can run a trial that can show efficacy at an eight-week time point as opposed to a 12-week time point, that can be a significant cost saving or acceleration to the whole program. Finding those opportunities to shorten the time into a primary readout was one potential use case because you can analyze the different values of a given endpoint over time across many indications at once, and then you can automatically detect, okay, well here for this particular indication, let’s say atopic dermatitis or something, we can see that the eight-week endpoint is very highly correlated with the 12-week efficacy for this endpoint. Actually, if you use this as your primary, you could potentially run an eight-week trial.
Ross Katz: Interesting. I’m certain that a lot of your customers or potential customers are experimenting with LLMs on their own as well, building their own RAG-based deep research agents from across the different foundation model companies and also building their own homegrown solutions to ask questions of their data. I’m interested in how you differentiate the work that you’re doing at Convoke from the big players who are trying to solve all the problems and from the internal data teams that are trying to pull themselves up by their bootstraps.
Majid Ahmed: One of the big technical investments that we’ve made up front that we think is leading to differentiation in the outputs we’re generating is we built this large-scale semantic layer over the industry. For every document that we’re processing, every document that we’ve sourced that we put into the system, we are using language models to transform it into a structured format. For example, for a study we’re interested in what are all the different arms of the study, what are the subgroups within those arms where different results were reported, what are the efficacy measures that were reported, what are the safety measures, what are the biomarkers and so on. For each of those entities that were mentioned, we then want to tag it within our broader semantic map. That way, when a user comes in and kicks off a task, for example, I want to see a landscape for non-small cell lung cancer, the agent can use the semantic map to quickly find the thousands of documents that are relevant for answering that question and do a comprehensive review of all those documents. That’s fairly different from how a deep research agent would approach that problem. A deep research type agent would kick off a few searches, usually within the web, pull in a sample of the right public documents and then use that to create its output. Oftentimes the output can be high quality, but it really lacks comprehensiveness.
Ross Katz: If I’m understanding correctly, you’re doing all of this backend work to create a knowledge graph or a semantic layer, a series of indexes that allow you to understand the relationships between the different documents that you’re searching and between the entities that are discussed in those documents. At query time, that enables you to pull in the most relevant information that is grounded in the domain that your customers are used to. When they use a word like protocol, it’s not in the general term, it’s in the specific term that you would find inside of a clinical trial. As a result of bringing in that richer information or more relevant information, you’re able to give better answers or generate better documents that are relevant to your customers’ needs. Am I thinking about that right, or how would you add on top of what I just said?
Alex Telford: I think that’s the right way to think about it. In principle, it’s a difference between pre-processing and doing everything on demand. In principle, you could have a deep research agent do everything on demand every time, but it’s going to be very expensive and it’s going to take a long time and it may make mistakes. By pre-processing everything and building a robust data pipeline and evaluation pipeline, we feel like we can give the model a head start when it gets to these questions that are specific to our domain.
Ross Katz: Interesting. You’ve got this process that’s using LLMs to pre-process and organize to the extent possible all of the data that exists out there in the healthcare biotech ecosystem that might be relevant to a question that one of your users might ask. I’m wondering, what did you learn about the process of building a knowledge graph or building a backend for a system like this based on trying to do this at scale? Because you’re talking about, I’m assuming, millions of documents, terabytes of data that need to be accurate to a degree that baseline search and indexing wouldn’t get you.
Majid Ahmed: One thing is that it’s hard. That’s why we did end up going down the venture route and raising money and trying to build a world-class engineering team here in the Bay Area to go after this problem, and we think a lot of the folks trying to do this in-house will see the same thing over time. There’s a lot of technical challenges that we’re working on. One is dealing with the scale of the documents as you mentioned. Building robust data pipelines that are trending the line very carefully with the rate limits from the model providers, working across the model providers, dealing with the long-tail accuracy problems as you’re working with that scale of data. For example, building pretty comprehensive evaluations of all of the model outputs we’re generating and really quantifying what is the error rate that we’re seeing on a regular basis and building systems where we can react quickly to any degradation in the error rate and make sure our final output is high quality.
Alex Telford: I was fairly naive when we started. When I started working on the problem I just put 100 clinical trials, uploaded them to the API and just told it to structure it. It does a pretty good job and you think, oh this is going to be easy, the models are so great out of the box at structuring this data. Then when you start scaling it up to millions of documents you realize that even though the error rates may be small for an individual model call, they compound in ways that could completely destroy your database if you’re not careful. One thing I also didn’t appreciate is just how much bad data is out there. Even something like clinicaltrials.gov which is nominally structured or at least partially structured, there’s a lot of poor quality data in that system. If you’re trying to build a clean version of that information you actually have to do a lot of processing that is not just LLM-based to clean everything up.
Ross Katz: That’s interesting. It also sort of makes me think that given the scale of documents that you’re working with and the specificity of the questions that are going to be asked of those documents and how much those questions depend on the role of the user and their perspective on the world, it strikes me that there’s infinite ways that you could structure a knowledge graph to solve those problems. Figuring out the level at which you construct it and the level at which you expose it to the agent at query time is a decision that’s very much related to the user who’s asking the question. I’m interested in how you think about the landscape of tasks that a user might ask of the system and how that relates to the backend architecture and how you handle the questions that are asked at query time.
Alex Telford: Obviously that’s a trade-off you make when you’re doing any pre-processing instead of doing everything at query time. You have to enforce some structure, and the structure may or may not be what the user exactly wants. But the trade-off you get is it’s much faster. I like to think that given I did a lot of this work for a number of years, I have a good idea of what the most common types of questions are. We built the schemas that we’re extracting around very common use cases. I want to see a table of drugs against this target and I want to see in the table the generic name or the sponsor and the modality, the indications, the latest development, all these things. In the clinical outcomes, I want to see the primary endpoint and I want to see the time point and the measure and the population. A lot of things that I think are just generic analyses that, even though they may be consumed by different functional areas within a company, you’re really using the same underlying data. But then there is stuff that is completely one-off or unique. There we have to make a distinction between, okay, is the schema we’re serving up sufficient or do we need to also get the model to go back into the document and retrieve something on demand?
Majid Ahmed: One of the design decisions that we made is we transform each document into a structured predefined format of the information that is relevant for us. For example, the reported clinical outcomes as Alex mentioned. The decision we made is we don’t want to create this knowledge graph that is a perfect encapsulation of all of the biological pathways, the biological reality of the world. We’ve kept it very simple intentionally, where a simple understanding of the relationships between different entities, and then tags to the source documents that mention those relationships. You’re able to then, when you’re doing the task live, the model can go and search through each of those documents and pull in the context into the prompt or into the agent’s resources. What we’ve found is the models in the agents have pretty strong reasoning ability to be able to come up with the biological understanding that they need to do the task.
Alex Telford: It’s an interesting debate we have quite a bit internally: how much structure do you want to put on the world? The models can discover some structure that is maybe not visible to us if you just train on everything. But then again, there’s certain information that’s evergreen. You’re always going to want to know what’s the target of a drug, probably, I imagine. It seems like that’s an evergreen piece of information. We have to try and pick out what are the bits of information that are always going to be important to decision makers and then probably leave a lot of the rest to just the models to pull on demand or train into the weights. It’s a pretty common area for debate because if you look at the history of machine learning, what tends to happen is the structure gets dismantled over time and a lot of it just gets pulled into the model’s weights.
Ross Katz: That’s interesting, and it strikes me that you’re architecting the system to improve retrieval, and then you’re relying on the models to handle the generation from there. If you can improve the context that gets into the LLM and make sure that precision and recall are both as high as possible, then you’re in a good position to get the answer to the question that you’re looking for.
Alex Telford: Think of it a bit like a building. When you’re building a building you have the scaffolding around it, and you need the scaffolding to help the model get to the right answer. But we shouldn’t be precious about the scaffolding we’re building because as the models get better, we’ll end up probably dismantling quite a lot of it. We need the scaffolding now to improve the models, but it probably won’t be there forever.
Ross Katz: Majid, you mentioned evaluations earlier. I’m imagining you have evaluations that you’re doing of each phase of the process that you’re working with, but then you also have the customer-facing evaluations and the feedback that you’re getting from the customer in order to put that back into the system and the workflows that you’re doing. I’m interested in how you think about each of those steps, and if you’re willing to give some insight into the types of things that are in your rubric and how you use them, I’d be interested in that as well.
Majid Ahmed: You’ve framed the problem in a nice way. The two ways we think about evaluations is an end-to-end evaluation that should match very closely to how the customer is thinking about the output, and if it’s scoring highly there, the customer should also score the output highly. Then, can we break that end-to-end task into each of these sub-tasks and can we build evaluations for each of the sub-tasks? That’s just useful internally for if things go wrong we can quickly hone in on where it’s going wrong and where we think we can drive the most improvement. One of the things we’re aiming to do is be very quantified in that process, really always just look at the numbers of where we are. We also think it’s useful for all of us to spend time doing the manual data labeling to create those evaluation sets that are then measured pretty regularly. All the engineers on the team will spend a bit of time doing the manual data labeling, trying to learn what we need to learn to do that.
Alex Telford: The other thing I would say is that, in addition to the quantified evaluations, these more like vibes-based evals are also quite important. Something as complex as writing a regulatory document or a long report informing a decision whether or not to go after a particular target is pretty hard to break down into good or not good. The customers’ opinion on the quality is also very important, even though it’s hard to operationalize that in the same way as a quantitative eval that we have.
Ross Katz: Is there an aspect of your front-end interface that’s gathering that feedback for you? If so, what are the components of that?
Majid Ahmed: Exactly. It’s fairly simple in the product. You can highlight a piece of text, throw in your feedback. For other non-text outputs we have feedback buttons as well. That all feeds into an automated system where we’re collecting that, we’re seeing how that output was generated and we’re trying to see if there was a hallucination, for example, what was the cause of that, and incorporating that into our evals as well.
Alex Telford: This is more aspirational right now, but we’ve heard from a number of customers that there’s some desire to create, I suppose, personas inside the system. If I want to write a document and have a document critique me as if it’s the German payer that I’m going to be pitching in a few months, then how can I get feedback into the system so the persona that the model inhabits is some useful stakeholder? We’re still thinking through how to actually build that into the product, but it’s a pretty interesting idea for how these systems would evolve to have many different personas and characters that you can work with almost like an army of colleagues.
Ross Katz: That’s really interesting and it strikes me that there’s a form of feedback that I hadn’t considered yet, which is feedback from your system to your customer or to your end user about the output that was generated and ways that they might not even have considered that it is insufficient to enable them to utilize Convoke to make it even better for, for example, the German payer. Am I thinking about that right, or is that where that goes?
Alex Telford: The way some customers are using it is, let’s say I’m writing a regulatory document or I’m writing maybe a piece of material that’s going to go out externally, some communications to a KOL. Normally that would have to go through a process of something like medical-legal review or another review process internally. What can be quite helpful is because those processes can be really lengthy, you have to coordinate with many different internal stakeholders, you can actually do a lot of the initial iteration just one person working with the model. Get the model to critique my document for the robustness of all the claims I’m making, give feedback, and then I can make a really robust document that by the time it goes for internal review, you’re actually already in a very robust document and it doesn’t have to circulate around the organization endlessly and you get quick sign-off.
Ross Katz: Interesting. I want to ask one more question on evals and then I’ll move on to some of the other technical stuff I’m interested in hearing from you all. Given that the outputs that you’re producing are the sorts of outputs that need to be evaluated by domain experts, and even the domain experts, to the point you just made, Alex, need a model to tell them where their blind spots are with regard to the other stakeholders inside their own organization. How do you think about the human-aided evaluation process that happens inside your four walls where you’re constructing your evaluation set for end-to-end testing and also trying to reflect the domain expertise that it’s hard to expect a software engineer to be able to take on?
Majid Ahmed: We do have a team of experts, including Alex and others that are formerly life science consultants, PhDs that are taking on the evals where you really do need that expertise, so the end-to-end evals, manual review of the final data sets that are generated to find the things where, okay, the sub-tasks all performed well, but actually led to this overall failure for some other reason outside of the sub-tasks. That’s something we’re continuing to invest in. We’re actually spending a lot of our time on building really good tooling for them, so you’re able to have the experts review the data pretty quickly, add in annotations pretty quickly, and that’s all fed into the system for automated improvements.
Alex Telford: I’ll also add that it’s important for us as a company to be users of our own product. That’s always important, but I think it’s especially important with language model-based products because we don’t actually know what the capabilities of language models are. We’re continually discovering them. We need to test the boundaries of a system and see where the failure cases are and where the success cases are, and hopefully we find the failure cases before our users do so we can build the scaffolding around that, and then when they go and try out that use case it’ll be successful. It’s a pretty interesting dynamic where, unlike traditional software, we don’t know everything the models are good at yet, we don’t know everything they’re bad at, and it’s changing all the time.
Ross Katz: It’s so interesting that, as with traditional product companies who do traditional software, obviously you have usability testers who are going through your website all the time and testing the front end, but now you have these intelligence and text-based outputs that require usability testers that have the requisite intelligence background, domain expertise to really help you improve a system like the one you’ve designed. As you’re going through the development of this platform, you’ve got a variety of technical hurdles. We’ve talked a lot about accuracy, evaluation, avoiding hallucinations, but there’s also latency, throughput, cost. I’m curious, were any of these challenges particularly difficult, and if so, how’d you go about solving them?
Majid Ahmed: One continuing challenge is working at scale. Being able to ask a new question, process millions of documents, create the right outputs to answer that question quickly. A few components of that is the job management to run these jobs at scale, having the right infrastructure components to work with that, making sure you’re not building any bottlenecks in the system. Now that we’re relying on these frontier model providers, being able to work with their systems well, understanding when they have capacity, and building that into our code. Those are things we’ve invested a lot in and we’ll continue to work on. Anything else you’ve seen that’s been challenging, Alex?
Alex Telford: One thing that’s been interesting to see from the customer perspective is if you give customers new tools that allow them to do things at a greater scale than they ever could before, they’re going to push it to its limits. Previously when I was at a consulting firm and we were doing opportunity assessments or target assessments for a company, we would negotiate for the scope and we’d go, okay, maybe we can look at 10 targets and maybe 15 targets if you pay us 300 grand or something. It would be that order of magnitude of targets we could go and review the literature and the competitive landscape and all that stuff. But now you don’t really have a limit. We had a project recently with a customer where when we kicked it off, we were thinking okay, can we look at maybe 50 targets or 100 targets? Eventually they realized that no, they could actually look at every target and go run that at scale. But now the problem is that they’ve generated so much analysis they have no way to read it all. How do you manage that? It’s a new way of working where you can actually with one click of a button get a system to go and write a pretty comprehensive report on every single target in the human proteome and why it may or may not be a good fit for your platform. These are all fairly high quality reports. But then they’re like, how do you use that information? The company cannot churn through that information yet. You need to build systems upon systems to deal with the volume of output you can produce.
Ross Katz: That’s so interesting, that when you’re operating at that scale of information, even the vibes-based evaluation is insufficient because the human being can’t even pick up a vibe on hundreds or thousands of pages of an analysis like that. You’re unlocking the ability to ask questions that they never would have thought to ask, but also enabling them to ask questions that even if you give them the answer, they don’t even have the capacity to consume.
Alex Telford: You need to build new systems, new interfaces that help them consume that information and distill it down because no one’s going to read a million reports, even if you go and assign all these agents to go and look at every target in extensive detail. You as a decision maker need a way to distill that information.
Ross Katz: Given that that’s the case, how do you think about quantifying the value or selling the value of what you’re able to do to existing and potential clients, or is it really just, try us and you’ll see how much more efficient you could be? That sort of thing?
Majid Ahmed: There’s some clear cases where we’re saving time, we’re reducing spend that they would have put into a consulting company, where the value is clear. But we do think in the cases where you’re looking at every possible target, that’s harder to quantify. You’re creating new opportunities that previously didn’t exist. In those cases we want to create that value for customers and we want them to work with us, and we’re trying to build a long-term relationship where they can play around with the product, start seeing that value, and over time we both capture a piece of it.
Alex Telford: I asked one of our customers who’s a platform biotech how they see the value of the product we’re building, and it wasn’t about time saving, even though that’s where you can get an advantage. It was more like, if you have a really broad platform that can go after many different targets, target many different tissues of the body, your decision space is very large and you can’t look at it all as a human. The value she gets from a platform like ours is, I can canvas the entire decision space and I feel comfortable that I’m not missing the mountain of gold over the hill because I sent something over there to go and look for it. The actual value is you’re finding the opportunity for massive upside, and it’s less about I’m saving some time, although that is nice.
Ross Katz: That’s really interesting, and it’s interesting to think about it as able to find the global optimum inside of this unstructured data space that you’re operating in rather than settling for whatever local optimum you would have had to find if you were limited in the amount of space that you could explore in that landscape.
Alex Telford: It’s the same thing for writing a regulatory document. I can generate 50 variants of a document and I can find the globally optimal version of that document, and I can get models to critique it until it’s perfect, which is something that you can’t spend that much intelligence or just human labor on some of these tasks where now you can.
Ross Katz: Interesting. My understanding is also that your customers can bring their own data to the platform and layer it on top of all the external data that you’ve gathered for accomplishing these goals. I would love to hear from you, how do you structure the system to make sure that the security and sanctity of your customers’ data is there but also that you’re able to generate all of the insights that are needed for the customer on top of that data, and that the insights from that data aren’t leaking into the evals or into the system later on?
Majid Ahmed: That’s a good question. We’ve built the system from the ground up where it can be taken and deployed to the customer infrastructure. Some customers choose to do that, and they’re using it in a completely isolated environment where they’re feeding in their own data, and they own and manage that infrastructure. That being said though, a lot of customers do use our shared multi-tenant subscription offering as well, and we’re very careful in having the right safeguards in place or isolation in place for all the data that we take from them, and having good policies internally for how we access that data when we’re looking at it for debugging. We don’t take it lightly that we’re working with our customers’ most sensitive IP, and we’re trying to build a company culture that really takes that to heart.
Alex Telford: As someone not from a software background, I was pretty fortunate to be working with Majid and Vikas who have done similar kind of enterprise deployments in the automotive industry for a long time. I think we take it very seriously.
Ross Katz: Obviously when you’re working in the biotech space where IP is such an important consideration, I’m certain it’s a question you get all the time. Another question I had was, obviously you’re never done integrating data into a system like this, it’s just going to continue to grow and evolve. I’m interested in how you think about versioning the sources or managing the state of the knowledge graph or making sure that the outputs that are being produced are reproducible so that if there’s a bug in the output you can trace it back. These are all the kinds of problems that come about when you’re dealing with a non-deterministic element in your tech stack. Just interested in how you think it through.
Majid Ahmed: It’s a good question. For all of the sources that we capture, we take snapshots at the time we process them. We know the state they’re in when they were processed, and then when we process that data, we create outputs that are also snapshotted and versioned, and that’s integrated into this broader semantic map that’s live. Under the hood that’s implemented as a Postgres database, that’s pretty simple. But then what we do is we’ll take regular snapshots of that database as well, and we transform it into Parquet files that let you do faster sharded columnar analytics over those outputs. It also lets you version the output. Any downstream task actually references a specific Parquet snapshot, and you can follow those. You have full traceability to the source documents and the snapshots of the source documents for any output that was generated.
Ross Katz: That’s really interesting, and the emergence of tools that allow you to query Parquet as snapshots in that way is, I’m certain, an enabling factor in that. That’s great. What have been the biggest challenges in convincing biotech teams to trust the AI-driven insights that are coming out of the Convoke platform?
Alex Telford: I think the main thing is you’re moving from a world where you know the person who is doing the work and you can watch them doing the work and you can ask them about it in a conversation on a Zoom call afterwards, to a world where you have all these background agents, for lack of a better word, doing a lot of the work to gather the information, structure it, and write the reports. Language models are not like humans. They’re spun up for a task and then they’re effectively spun down again, they don’t maintain context or state between prompts. You can’t really ask them why they did something in a certain way, and if you could, they wouldn’t give you an answer that was reasonable anyway. The main thing is just, how can I trust this system? How can I trust this system is doing something in a way where I can believe the outputs, I believe the right process was followed? Building trust is an ongoing process and I don’t think we’ve completely nailed it yet, but the way we’re approaching it is by building systems that give visibility into what the models are doing and controllability over what the models are doing. If you have a multi-step process where the language model is going to do something more agentic, you can do a series of tool calls and search PubMed or write and iterate on it. Getting visibility into what process it did and where it got the data from, traceability tools to sources, has been important. Just as the users use the models more, they do get a better sense of where these systems are weak and where they’re powerful, and then they can integrate it more into their daily workflow. Change management is difficult when you’re starting from using an entirely new technology that people have not been used to working with yet. It’s just going to take time for everyone to integrate language models into their work.
Ross Katz: Do you feel like you need to put a label on the outputs that you’re providing to people, communicate the uncertainty or the limitations of the approach or the models so that they can better consume the information, or do you feel like it’s common knowledge that LLM-based outputs require being double-checked by humans?
Alex Telford: We try to be upfront with where we think the system is strong and where it’s still under development. But actually, a part of why we think a company like ours makes sense is you can build those guardrails into the product experience. We can expose functions to the user where maybe there’s a button that they press, and that button behind the scenes is a prompt to a language model that does some process, but they’re not inputting prompts. A lot of the variability in quality you get is in variability in what context you feed to the model and how you prompt the model. If we can control those aspects, you can get a more reliable, consistent experience versus if you’re just interacting through a text interface. Just exposing the users to ways of working, interfaces that constrain how you use a model in the use cases, can get around a lot of those issues.
Majid Ahmed: I’ll add to that as well. One of the features we have in our document generation product is the ability to then iterate with the model on improving the document. You’re reading the document within the product, you’re leaving comments, and then the model can go and work on those comments. That’s something that some customers really like compared to using a deep research where it just creates the output and it’s one and done, you can’t do much with it. Our users are iterating with the product before then disseminating it widely within their organization.
Ross Katz: Interesting. They can see the intermediate outputs and then respond to the intermediate outputs to guide the agents to go back in the right direction?
Majid Ahmed: Exactly. Once the final output’s created, you can then give feedback on the final output and say, actually I didn’t like the way that you wrote this, can you make this into a table instead or add a few more sources here?
Ross Katz: Do you give them the ability to guide the agents toward different sources like, ‘I read an article in STAT two weeks ago that addressed this exact thing. Can you go check-’ maybe this is a bad example, but I know that people who work in this landscape are constantly consuming information, so I imagine they might have ideas about sources as well.
Alex Telford: This turns out to be a very important part of the product. When we first started going out with the document generation feature, customers were interested, but they were only willing to use it when we gave them source control. Source control is such a fundamental part of how work gets done. If we’re working with a med affairs team, they have a validated library of sources, and they only want to use those sources, they only want to use ones that have got internal sign-off. Or if you’re creating a regulatory document, you can only use a paper or a piece of material or experimental data that is cleared for release. If the model goes and collects something from the web, that’s not useful. It can only use specific sources. The sources are constrained to even sections as well. You need pretty robust source control before you could even start doing work in a product like ours. It’s something we had to build pretty early.
Ross Katz: That makes a lot of sense, and when the provenance of the sources is something that the user understands in advance, then it makes it much easier for them to trust the outputs that you’re giving them versus when it could have pulled from literally any page from the internet. That seems like a critical aspect of what you’re building.
Alex Telford: Even though we spend a lot of time gathering information from the internet, that’s just such a small piece of the information that’s in the world of an individual user. So much more of it is internal knowledge, and that can be documents, but also internal context that’s not necessarily encoded in a format that’s easy to access for the model. We need to find ways to get all that information efficiently into the system, what’s in the world, what are the documents that exist in the organization and data that exists in the organization, and then what is the data that exists in the user’s head that hasn’t been encoded or written down yet? That’s what we call tacit knowledge. I know that Greg in regulatory knows how to write a document in the format that I like it. How do I communicate that shared knowledge between me and Greg now between me and the language model? We need new systems for that.
Ross Katz: That’s really interesting and it brings up a couple extra questions for me. Obviously the knowledge that the organizations you’re working with is not just in unstructured document format, it also exists in spreadsheets, it also exists in databases, it exists in FASTA files, it exists in a variety of different places. Are you all integrating that more structured information in the variety of formats that it’s there, or is that something that you think will be valuable, or for now are you selecting use cases where the unstructured information is the best path for Convoke to drive value?
Majid Ahmed: We have integrated structured data as well including Excels. We’ve actually been pleasantly surprised at how well the models perform with tabular data as an input interspersed with unstructured text data. We have not done much on raw experimental lab sensor data. That’s something we might add in the future, but right now focus on documents whether they’re Excel or text documents or PowerPoints.
Ross Katz: Interesting. One of the insights that was shared on one of our past podcast episodes was the idea of MCPs being the future interface to especially those structured data sets that people might be searching through. I’m curious, how do you all think of the role, if there is a role, for MCPs and the way that you plan to integrate information into Convoke in the future as that standard is emerging?
Majid Ahmed: We think it’ll be important. The thing that will be important is building multi-agent systems where you can have each agent only have a subset of the tools available to it and perform really well on analysis over some subset of the data. Then, whether it’s MCP or another communication format between the agents working well and working reliably in a way that can be tested. That’s where we’re building towards moving as well, being able to build these multi-agent systems with an agent having a different expertise.
Alex Telford: You generally do better by not fighting the models. MCP is a format for providing information on APIs that the models are being explicitly trained on. It’s often good to try and adopt standards the models are trained on. Models like Python, they like TypeScript, they don’t like OCaml or whatever, some of these more obscure languages. Generally if you want to use language models you get more benefit out of the stuff that’s more represented in their training data. That’ll be self-reinforcing for MCPs most likely.
Ross Katz: That makes a lot of sense. Majid, I take seriously what you’re saying in terms of scoping the agents appropriately that they’re really good at what you’re telling them to do but also they can coordinate effectively so that they can accomplish the goal that they’re designed to accomplish better. As we head toward the end here, I would love to hear from you, is there anything on the horizon for the future of LLMs that you think you’re particularly well-positioned to take advantage of, or that you think might be a risk to the way that you’re developing Convoke today?
Alex Telford: I’m pretty excited about language models becoming truly multimodal. We have some multimodality now, you can feed an image into the model, you can get the model to generate images, but actually if you try to get them to really parse a graph or really understand a video, they don’t really understand it. They make a lot of elementary errors in reading out information from a graph. As we move towards these systems that are truly able to reason in visual space to a similar degree that a human can or reason about a video clip, that opens up a whole new world of information that at the moment is not as represented in the models’ weights as we might like, and we have to do so much work to accurately parse a graph right now because the models believably can’t read them that effectively. So much of scientific data is in graph format, it’s very visual science, biology.
Majid Ahmed: I’ll add to that as well. It’s really fun seeing this emergence of these new systems that convert this raw intelligence from these language models into useful outputs. Whether it’s letting these models become agents and being able to create and follow plans, or building the right systems that can put context into the model, this is all obviously very new, and it’s fun being at the forefront of that field and seeing these systems become more productized over time. I think we’re going to be in a good position as the best systems become available, whether it’s from academia, open source, or commercial systems, being able to implement them specifically for our customers and specifically for the domain.
Ross Katz: Awesome. As the future unfolds and Convoke continues to do what you’re doing, creating this iterative feedback loop rather than the linear feedback line that has existed previously, is there anything in biotech that you expect to fundamentally change from interacting with tools like Convoke?
Alex Telford: I think we’ll probably move from making these decisions on a scheduled cadence. Like I’m going to have my portfolio review meeting every six months or every year, or we’re going to do a big target landscape refresh every six months. Because you can actually do a lot of the input work continuously and cheaply and easily and autonomously. I think you’ll move to much more continuous decision making about things like how you construct a portfolio and whether or not to advance an asset. Obviously, once a trial is running you can’t just, so some things are set in stone and have to run their course, but there’s a lot of things that can be moved to more on-demand, on-time decision making. How do we respond to competitor data? A new target, some information about a new target comes out, either from a clinical trial or in the literature. Can we immediately jump on that and really generate a new TPP and an investigative plan and submit everything and just get it rolling in hours instead of months.
Majid Ahmed: To add to that, study design will be particularly impacted by a move to a more data-driven world. Every aspect of your study design will be backed up by pretty rigorous data that shows, hey, this is the best decision for us for our platform, for reading out the right data at the end of this.
Alex Telford: It’s like a financialization or money-ballification of the industry.
Ross Katz: Interesting. What occurred to me as I heard you talking was that the arrival of new information being the trigger for refreshes of questions that we’ve asked before or questions we might like to ask, and the pervasiveness of intelligence means that you don’t have to have a human think of a question in order to have that question asked. If you know in advance that when the result of this experiment becomes available, you’re going to want to ask a series of questions, that can all just happen for you and then you can review it on your timeline versus having to know that the new data exists and then ask those questions.
Alex Telford: You can think of processes or workflows that are done emergently by an organization as code now. These things can just happen, and organizations can be effectively computational systems.
Ross Katz: Very interesting. As we head toward the end, is there anything we missed, any final thoughts you want to share?
Majid Ahmed: I’ll add that I’m sure there’s many software engineers listening to your podcast as I was, curious about a move to biology. Feel free to reach out, happy to share more about what we’re building and what we’re up to. We’re looking to grow our team pretty significantly over the next few months.
Ross Katz: Awesome. Where can listeners go if they want to learn more about you and your work at Convoke?
Alex Telford: The website is convoke.bio. You can find us on LinkedIn, Twitter as well.
Majid Ahmed: Thanks for the time.
Ross Katz: Fantastic. Well, Alex, Majid, it’s been a pleasure talking with you, and look forward to connecting down the line.
Alex Telford: Thanks, thanks for having us.
Jason: And that’s it for this episode of Data in Biotech. If you enjoyed the episode, please subscribe, rate, or leave a review in your podcast platform of choice. See you next time.






