Listen on
Overview
Designing a clinical trial today can take years and cost billions, largely due to reliance on unstructured document-based processes. This traditional approach frequently leads to high patient burden, site complexity, and expensive regulatory amendments that delay critical new medications. Data leaders in biotech face immense pressure to accelerate drug development while controlling costs and maintaining quality.
In this episode, host Ross Katz speaks with Patrick Leung, Chief Technology Officer at Faro Health, who brings a fresh, data-first perspective from finance and big tech to the highly regulated life sciences industry. Patrick explains how Faro Health is systematically addressing these inefficiencies by transforming clinical trial design from a document-centric task into a data-driven process.
The conversation covers Faro Health’s dual-model AI system for generating clinical-grade documentation, its unique approach to quantifying patient and site burden, and how AI can provide predictive insights to optimize trial protocols. Patrick also discusses the practical challenges of AI adoption in biotech, managing data privacy concerns, and his vision for how AI can lower the barrier to developing and deploying new medicines.
Key Takeaways
Clinical trial design is a data modeling challenge, not merely a document authoring task.
Traditional clinical trial protocols, often drafted in Microsoft Word, introduce errors through information duplication and prevent quantitative analysis. Faro Health’s approach models trials using a structured data schema, enabling real-time metrics on patient burden and cost. This shift from unstructured documents to structured data is essential for efficient, patient-centric trial design and execution.
Generating clinical-grade documentation with AI demands a layered approach where an evaluation model scrutinizes the output of a generation model.
Naively asking a large language model (LLM) to generate a protocol often leads to errors and made-up information, which is unacceptable in clinical trials. Faro Health employs a dual-model system: one model generates content based on clinical expertise and structured data, while a second ‘critic’ model—informed by extensive checklists—refines the output based on both machine and human feedback. This iterative process ensures clinical quality and regulatory compliance.
Transforming unstructured historical trial data into structured, comparable insights is crucial for optimizing future trial designs.
Faro Health processes thousands of publicly available clinical trial PDFs, extracting and structuring data on activities, schedules, and subtle nuances. This rich, queryable database allows sponsors to identify similar trials with lower patient burden or better efficiency. It provides data-backed guidance to optimize new protocols, helping avoid common pitfalls and costly amendments.
Rigorous AI governance, including isolated model instances and transparent safety protocols, is paramount for securing sensitive pharmaceutical IP.
Biotech companies view internal documentation, especially related to new molecules, as crown jewel IP. Patrick stresses that private data must never be shared with public models. Solutions involve running open-source models locally or using guarded cloud instances from vendors like OpenAI, alongside clear AI governance policies that ensure data safety and prevent issues like script injection attacks.
Related: CorrDyn helps biotech and life sciences companies implement AI strategies that drive real business outcomes. Our data engineering expertise builds the foundations for these systems, often leading to significant data cost optimization.
Full Transcript
Jason: Welcome to Data in Biotech, a podcast from CorrDyn where we explore how companies leverage data to drive innovation in life sciences. Every two weeks, we sit down with an expert from the world of biotechnology to understand how they’re using data science to solve technical challenges, streamline operations, and further innovation in their business. Today, we’re joined by Patrick Leung, Chief Technology Officer at Faro Health. Patrick shares his journey from finance and big tech to life sciences and explains how AI is transforming clinical trial design and reducing inefficiencies. He and Ross discuss the inefficiencies in traditional clinical trials and how Faro Health leverages AI to optimize protocol design, reduce patient burden, and streamline operations. Patrick also shares his insights with Ross into the challenges of AI adoption in biotech, the evolving role of regulatory agencies, and the future of AI-driven drug development. Here we go.
Ross Katz: Patrick Leung, welcome to the Data in Biotech podcast.
Patrick Leung: Thank you. So glad to be here.
Ross Katz: Well, just to kick us off, can you give us an introduction to your background and what brought you here today?
Patrick Leung: Sure. I’m currently the Chief Technology Officer of Faro Health and we are an AI-based SaaS platform for clinical trial development. I have a long background in data science and AI, having worked at Google and Two Sigma Investments and various other companies in the past.
Ross Katz: How did your career lead you to Faro Health, considering you’ve been in finance previously and other things like that?
Patrick Leung: Yeah, I think I realized at some point that I really want to work in companies that have a real mission. Previously I started my own company doing ecological restoration, and when it was time for a change, I really wanted to find a company that helped to I guess make the human condition better. Life sciences is so fascinating to me. It’s a very unique industry. Being relatively new to the field, it really took me a while to realize, it takes up to a dozen years and maybe a couple billion dollars to get a drug to market. That just seems like a lot. And there’s real human lives in many cases at stake here, or at least the alleviation of suffering in the sense that there’s so many new incredible medications and molecules coming up and it takes them a long time to get them to patients’ hands. I felt that was a really worthy mission that Faro was endeavoring to pursue. I came in last year and have brought a lot of my knowledge and experience into this organization.
Ross Katz: How do you feel like your background has brought a unique lens to the kind of work that Faro Health is doing, and where are the areas where you feel like your background positions you well to accomplish certain goals at the company?
Patrick Leung: Yeah, I think spending over 10 years at Google helped me a lot to realize what a really high-performing software engineering company operates like. And I think that not having a long background in life sciences and not being, I guess, indoctrinated into a lot of the ways of doing things helped to bring a fresh perspective to this. It really shocked me that something as complex and as high-stakes as a clinical trial was essentially being mastered in Microsoft Word in many cases or some kind of document authoring system. And to me, this is a complex project plan that really deserves to have its own dedicated, highly structured, and advanced system. This is really what Faro has set out to do, and I think that my bringing in the application of AI and data science and other areas like private investing and conversational AI helped to bring a fresh perspective on how to solve these problems, going back to first principles. I think that’s really borne out in the last year to 18 months during which we’ve been really bringing AI into this company and to this field.
Ross Katz: I want to explore some of the AI aspects of it a little bit later on in the conversation, but before I do, you mentioned that a lot of clinical trial designs are drafted in Microsoft Word — that’s the document type from which the clinical trial designs are run. Can you give us your perspective on how not having a more structured approach leads to inefficiencies in the way that clinical trials are designed or conducted?
Patrick Leung: Yeah, we’ve been really delving into the protocol document and how best to generate that using AI. We can get into that later. But essentially what I realized was that these documents contain a ton of quantitative information that’s been summarized in various places. And this information gets repeated in many other documents further down the line, the ICF and other critical documents during the course of the lifetime of the clinical trial. In many ways, storing such information in an unstructured format like a Microsoft Word document is probably the worst thing you could be doing as far as avoiding errors arising from duplication of information. And if you have to go back and change something, making sure it gets changed in all the other downstream documents — it’s just really inefficient. I know there’s different measures you can use like reusing components or using templates, but really what you need is a structured data store, ideally backed with analytics, to drive this process.
Ross Katz: What does this look like from the perspective of customers of Faro? I noticed that you recently announced a partnership with Recursion to optimize and accelerate clinical protocol design. From the perspective of a new pharmaceutical company that’s coming to the platform, what does it ask of them, and why do they prefer the workflow that you’re offering them to the workflow that they’ve traditionally had to do?
Patrick Leung: Yeah, I think the main thing we’re asking of our customers is to get used to a different way of doing things. Instead of writing a document, we ask that you use a really elegantly designed user interface, a web-based SaaS solution, to start modeling these trials. Our goal is to really shorten the amount of time it takes from when you start modeling the trial to when you can actually gain insights into the design of that trial. We provide metrics on patient burden, on cost, and complexity, to give you an idea of how this trial is going to be for the patient when they actually, further downstream — it might be a year or two years before patients start getting enrolled, coming into those sites. What is that experience going to be like? Is it going to be too burdensome? Are we likely to run into trouble enrolling patients because of the amount of time they have to spend outside the office, going to the hospital, or the amount of blood that’s being drawn? So we try to get you that information as quickly as possible so that you can essentially design, so that sponsors can design a trial that is amenable to patients, patients-friendly and also site-friendly. That’s really important to ensure the trial’s going to be successful and completes on time and is efficient.
Ross Katz: Can you unpack that a little bit? How do you quantify the patient burden or the complexity from the site perspective, and how do people who are using Faro leverage that information to improve the design?
Patrick Leung: On the one hand we’ve built up a really comprehensive what we call a biomedical concept library — a library of all the different activities that can occur during the course of a trial, thousands and thousands of these activities organized into this standardized library. And this contains cost information that we’ve accrued from various data sources. On the other hand, we have this tool that allows you to rapidly put together a schedule of activities and a schema for that trial. We use the library to calculate the insights based on the SOA, the schedule of activities, that you model using our system. In doing so, we’re able to compare with benchmarks and give you some idea of how this trial stacks up — is this an unreasonably burdensome trial, or is it maybe incomplete in some way? We give you guidance on how to address that.
Ross Katz: Right. And I imagine that the success of the trial is sort of based on the patient burden and the site complexity as well, because if it’s burdensome on the patients and complex for the sites, then you’re probably going to have more situations where the protocol isn’t being followed — compliance with the trial protocol just generally is not occurring. Am I thinking about that right?
Patrick Leung: Yeah, exactly. And to be honest, designing a trial that’s really patient-friendly and site-friendly is a lot of work. Traditionally that involves trying to figure out what trials have been completed either in my own company or externally in the public domain, published to clinicaltrials.gov. Which of these trials are somewhat similar to mine that I can use as guidance? How did they fare? Which ones had issues enrolling patients? Which ones had amendments with the FDA? Each amendment costs millions of dollars in terms of both material cost of writing the amendment and filing it, as well as delayed time to market, because an amendment costs you six months plus of time. Designing a patient-friendly and site-friendly trial that will pass muster with the FDA and won’t require amendments is a lot of work. It requires a lot of research. We are basically in the process of incorporating AI to help automate that. Not only do you have a really elegant tool that can provide you with insights, we can also provide guidance and say, ‘Hey, it looks like trials similar to this historically have had amendments. You might want to rethink this.’ This is really an ongoing process that’s super exciting — leveraging the latest advances in large language models to increasingly make the process of designing a complex clinical trial a lot easier and a lot less error-prone.
Ross Katz: I’d love to dig into that. You’ve shared two use cases so far of how Faro is using AI — the document generation use case and assisting with design and identifying warning signs about the trial that’s being designed and how that might lead to amendments or additional cost or time later on. Can you give us the overview of how Faro is using AI in those use cases?
Patrick Leung: Yeah, these are two pretty different areas and require different approaches. I’ll do them one by one. On document generation, we’ve been thinking about this and prototyping and evolving this for well over a year now. Naively you might think, these large language models like GPT are so capable, surely we can just ask it to generate a protocol. And you can’t really do that, for a number of reasons. First of all, the LLMs aren’t as smart as we think they are. For anyone listening who doesn’t really understand how LLMs work, the slightly tongue-in-cheek mental model I’d like to encourage you to think about is super well-informed, super sophisticated auto-complete. What that means is that the LLM is game for anything. It will start writing things and producing text no matter what, but it’s not always true. What we found was that no matter how much prompting, no matter how much guidance we gave it, the LLM would invariably miss things or make things up. That’s really dangerous when it comes to clinical documentation, both in terms of wasting time and submitting things potentially to the FDA that require a lot of amendment and feedback, as well as just your own clinical process and ensuring that your clinical trials are really high quality. So we’ve devised this dual model system where we have one model that generates — that incorporates a lot of clinical expertise into the prompts that we use and incorporates data both from the Faro study designer as well as the sponsor’s own documentation. And then we have another model that essentially checks the output. It also embodies a lot of clinical expertise that we’ve downloaded from the brains of our clinical expert team and put into a bunch of checklists and an evaluation framework, and we score the output coming out of the generator model. We use that scoring function or checklist to refine the query that generates the protocol section and produce a second cut based on both machine and human feedback to make it better. This is what we’ve found is needed to really generate clinical quality documentation for things like the trial protocol. That’s what we’re doing on the document generation side. It’s an ongoing process. Some of these protocol sections, as many listeners will know, are extremely complex and require both a lot of quantitative information as well as a ton of verbiage that describes the process, and so we’ve been systematically going through the clinical trial protocol subsection by subsection to generate those individually.
Ross Katz: Can I ask some questions about that before you move on to the other application of AI? You’ve got the generator model and then you’ve got the critic or the evaluator model. Are you using third party models for both of these? Are you fine-tuning third party models or are you utilizing open source models or fine-tuning open source models for these purposes, and how do you think about that?
Patrick Leung: We’re pretty model agnostic right now. We built the first version of this technology using GPT just because it was most familiar, available in Azure and so on. But there’s been a plethora of alternatives coming out. I tend to favor open source. I’m super intrigued by DeepSeek, there’s a lot I could say about that. But suffice to say that we are rather agnostic when it comes to the LLM we’re using, and in fact there probably will be a number of different LLMs we use that are more suited to individual operations that we need for this system. It really comes down to which models perform best and are the most cost effective while maintaining the quality required for each use case. What we’re finding is that certain types of content — for instance, if it’s super highly structured and requires a lot of quantitative information — certain models are better at that, and then if it really requires a certain tone or template to be followed, you can probably get by with a lesser model for that. OpenAI has started to bifurcate its models based on cost versus performance, and I think there’s going to be a lot of that kind of bifurcation as different use cases become obvious and prevalent. DeepSeek is super interesting because it’s really changed the cost curve on inference and also training. They’ve made some really incredible advances in training the model and are able to offer similar services to the more established players at an order of magnitude or more lower cost. I think that’s going to have a really disruptive effect that’s going to help provide these LLM services at a lower cost to everybody.
Ross Katz: You mentioned that you’re incorporating the sponsors’ documentation into what you’re generating. Have there been any privacy concerns about incorporating their documentation into the generation of the clinical documents given that you’re using OpenAI or potentially using a third party API?
Patrick Leung: Yeah, life sciences is highly regulated — there’s a lot of sensitivity around internal documentation, particularly those related to new molecules, which is the core IP crown jewels of a sponsor. Every single time we engage with a sponsor and we start talking about AI, everybody asks us this question: ‘Is my data going to end up being made available to other people via these models?’ And the answer is no. There’s no way we’re ever going to pass any private information to a publicly available model. The good news is that we have open source models that we can run locally in our own cloud instance. And vendors like OpenAI do provide a guarded instance where you can deploy a model like GPT-4O onto your own Azure machine instance or cluster and are assured that none of that data ever makes it back into the general purpose OpenAI system that is available to the public. That’s actually essential — it’s a sine qua non when it comes to offering AI to life sciences companies. You have to have that story in order. Moreover, I’d go further than that. What we’re increasingly seeing is that sponsors are concerned about AI governance. What is the safety protocol around this even if you do have this kind of isolated instance? How can you be sure that the output is going to be safe? That there’s not going to be some kind of guardrails that are overcome, maybe script injection attacks and things like that? We’re starting to see these concerns coming in as the LLMs have become more widespread — people are starting to discover there’s these really bad edge cases that we need to be careful about. We are getting ahead of the curve there and we have a whole AI governance initiative where we’re essentially going to be thought leaders in this. We want to publish our AI governance policy so others can copy it and subscribe to the same standards that we do as far as keeping this information safe.
Ross Katz: Anytime you’re working with LLMs, you have to think about benchmarks and evaluation criteria so that any changes you’re making to the prompts or any updates you’re making to the system for generating, you’re able to see where it’s regressing and where it’s improving on the things that you’re trying to do. How do you think about building and maintaining the evaluations for a complex system like this that’s generating a very complex document with regulatory implications?
Patrick Leung: I think it’s really key to have a human in the loop as well, at least during these early stages. There might come a time where we’ve generated so many protocols and evaluated so many protocols that in 99 percent of the cases it’s going to be completely foolproof because we’ve seen it all, but we’re not at that situation yet. We just haven’t tested or generated hundreds of thousands of protocols simply because that would take a lot of time and money. As you can imagine, if we’re doing a great number of LLM calls for every single subsection we’re generating and there’s potentially a hundred plus subsections in a document, that really adds up. Definitely having a human review the output is critical. There’s always going to be edge cases, situations where despite our checklist and despite everything, the output does require some correction that can only come from having an expert actually look at things. But our goal is to really reduce the percentage of cases where that’s necessary and to provide tools that make it easy for someone to give some overall direction, some correction that then gets factored into the output.
Ross Katz: I’m interested in what you said at the end there about how you incorporate the human feedback — which is in many cases qualitative — into the process of making your document generation process more effective at getting it right the first time or the first couple times.
Patrick Leung: That actually plays into the strengths of these LLMs. Because you don’t have to reduce or transform your qualitative feedback into quantitative rules, which you would typically have to do if you’re fine-tuning a database or some more traditional content management system. Instead, you can actually provide overall direction like, ‘Hey, you’re being a little bit wordy. Reduce the number of words.’ And the LLM can take that kind of super high-level qualitative feedback and turn it into a refined result. However, the output of such tuning needs to be evaluated to make sure that nothing gets missed — that in obeying the human feedback of making things shorter, you don’t miss things that are important now. We need to continue with that checklist-based evaluation.
Ross Katz: For sure. Yeah, that makes sense. Can you talk a little bit about how AI is used in the clinical protocol design process?
Patrick Leung: Yeah, this is a newer area for us. There are definitely things that we’ll announce during the course of the rest of this year. But basically the idea is that we have all these clinical trials out there in the world. If you go to clinicaltrials.gov, thousands upon thousands of trials sitting there in PDFs. We’ve developed a data pipeline and a knowledge modeling system that takes those PDFs and essentially decomposes them into chunks of text and images and tables and really uses the LLM to interpret those and to vectorize the text — so we have a RAG kind of retrieval augmented generation system that takes all those different chunks of text and creates a knowledge base system to be able to access them. But we also look at the tables and interpret them, and link the information in those tables — for instance, each activity in the SOA, we extract that out of the table from the PDF and we link it to the relevant sections of the text. That gives us a richer view on what the structured information in that PDF is. One way of thinking about this is it’s almost the opposite of document generation — instead of taking structured data and generating the document, we’re taking the document and extracting the structured data. Why would we want to do that? This provides us with a really rich database of trials that we can perform really advanced queries on. We can say things like, ‘Okay, given this trial you’re working on in Faro, give me five examples of trials out there in the world that are similar to this one, maybe it’s an oncology solid tumor trial, but that have a lower patient burden.’ Let’s say I’m analyzing a trial that I’ve built before or that’s nearing completion and I have this feeling like, ‘I think this is going to have a high patient burden.’ This seems like a big bloated trial. What do I do? Maybe you’ve copied and pasted it from some other trial and added a few things and the scientists wanted a bunch of extra measurements and you’re getting this feeling like, ‘This is probably getting a little too burdensome.’ Our insights might tell you, ‘Yeah, the patient’s going to be spending eight hours a day at the site, or there’s an unreasonable amount of blood being drawn here. What do I do?’ We can basically perform a query against this structured clinical trial repository that contains thousands upon thousands of trials decomposed into structured data and make these kind of very researchy, high-level requests: ‘Just give me some clues here. Show me some trials that have a lower patient burden and why. Or just give me some guidance — I want to reduce my patient burden. What do I do?’ This is going to save a lot of time and result in much better trials in the future because we’re hearing time and time again from our sponsor clients that they have this pipeline of trials, and because of the difficulty of actually creating an optimized trial, there’s a lot of copy and paste or best practices that end up becoming bloated over the years. We can really help cut into that by providing these very incisive insights into how to improve the metrics of this trial.
Ross Katz: I have several questions on the back of this, but the thing that comes to mind first is that what you said up front about the data model and the schema of how you understand a protocol document is the critical middle piece between the protocols that you’re creating and the protocols that you’re decomposing, so that you can compare these together. Am I thinking about that right?
Patrick Leung: It’s so interesting because that was the initial thesis behind Faro before we were even thinking about AI much. There were just a few years, quite a some number of years of just really in a very detailed way modeling trials in a very structured way — all the nuances of the different activities that get performed during the course of a clinical trial, all the different chemistry panels and other things that go into that trial, the recurrence scheduling, the nuances of the population schema. We had that already. When LLMs came along, it was really just a natural extension of that structured database — now we want to generate documentation as well as input tens of thousands of trials into the system so it can be analyzed in a really structured way. It puts us in a pretty strong position that we have this data model that we’ve built up over many, many years and have built extensive reporting and analytics on, that we can now leverage into this LLM world.
Ross Katz: Right. So that question you used as an example earlier — ‘Show me examples of clinical trials that are similar to this one that required less patient burden’ — that presupposes that you already have the trials parsed in the same way as the trial that you’re entering, and in a context where you’re able to calculate all of the activities and as a result the patient burden of every trial that’s come into your database, so that you can order them appropriately and look at the semantic closeness but then also sort by patient burden.
Patrick Leung: Exactly. Because it’s one thing to ask the LLM a question and get an answer, but it’s another to walk through the reasoning. ‘Why did you say that? Why are you recommending we look at these trials?’ By having all this nuanced structured information at our fingertips or at the LLM’s fingertips, it can explain — for every single answer it comes up with it can say, ‘Well, I looked at this trial and there are these procedures that are more efficient than the ones you are using, or it doesn’t have these unnecessary procedures that you have in your schedule and so you might want to consider removing them.’ By having all of those details, we can really provide detailed explanation of why. So it becomes less of a black box, less of a ‘Here’s your answer that got generated by the deep learning system’ and more of an actual walking through the thinking and the reasoning behind that.
Ross Katz: What would you say have been the hardest parts of parsing the existing clinical trial documents and getting them into your schema so that you’re comparing apples to apples with the other trials that you’re creating?
Patrick Leung: We studied our own clinical team really closely because they were doing this manually, which is wonderful. We have a team of experts who have this very, very intricate process that sometimes takes a week or more per trial and we got to automate that. We got to really interview them and without question, it really is the kind of detective process of building up the mental model of looking at the SOA, looking at the footnotes, looking at all the different sections that talk about the schedule, doing this with various other parts of the trial as well — sort of doing all that sleuthing to figure out what is the true structure of this trial. Because if you just look at the table, which is relatively easy to do with some PDF parser, that doesn’t give you the full picture because there’s all these gotchas. There are footnotes that say, ‘And by the way, there’s this other thing that you gotta do every two days.’ Or, ‘At the end of the trial, you gotta keep doing this for a number of weeks.’ And it’s those gotchas that can dramatically affect the cost of the trial and the burden that you wouldn’t necessarily pick up on if all you looked at was that straight raw table. The trickiest part for us, and it is an ongoing process, is really threading the needle between the SOA and all the footnotes and all the sections that refer to the SOA and all those nuances that really affect the structure of the trial.
Ross Katz: It strikes me that a lot of the work that’s being done in LLMs right now is trying to reflect the reasoning process or the chains of thought that will lead you to the correct conclusion when you ask a question or you try to get particular structured data out of an unstructured document. But the thought processes that you’re describing are thought processes that have not been written down and published in any meaningful way, and so you have this unique window into what the steps are. Am I thinking about that right?
Patrick Leung: Yeah, exactly. Definitely chain of thought reasoning informs our approach and our approach is very domain-specific — clinical chain of thought, chain of reasoning. And this is new stuff and so we might even publish on this. We’ll see. It really is super interesting going through this process of really breaking down what our clinicians have been doing and automating it using LLMs. It’s a new thing and it feels really exciting to be working on something that’s at the state of the art in this area.
Ross Katz: You mentioned that you’ve got announcements coming in the future and you’re in the middle of this process of using AI to help support clinical trial design. Can you talk a little bit about what are the capabilities that you’re most excited about or that you think people will be surprised by once the problem is finally solved at some point in the future?
Patrick Leung: I’m definitely excited about the document generation just because it’s been such a slog to really get all the sections done. I feel very proud that we’re generating output that our clinicians are saying, ‘This is really good. This would definitely pass muster with the FDA. This is very high-quality clinical output.’ It took a while to get there, so I’m excited about that. But I think the really groundbreaking thing that’s going to shift this whole industry is going to be really intelligently designing these clinical trials. With the benefit of the knowledge of all of the recent trials that normally would take a lot of time to read through and understand and use to inform the new trials that are under development, just being able to do that automatically using LLMs and have this kind of super informed and intelligent research assistant tell you, ‘I’m going to help you design the best possible trial so that you can get this drug to market faster and cheaper and with a better experience for your patients.’ So you can improve your chances of enrolling patients and enrolling sites. That’s really exciting to me because I think that will benefit patients, it’ll benefit sponsors, and also benefit the sites. And that’s what really brought me into this job in the first place — that premise that we could do that.
Ross Katz: That makes a lot of sense. We’ve talked a lot about the platform and about the features that are available and about to be available, but I’m interested in who are the people in biotech and pharma that are actually interacting directly with the system. What are the personas of the users that are getting into Faro, designing the clinical trials, and getting the documents back out?
Patrick Leung: There’s definitely the clinical scientist — the people who are basically tasked with designing the trial, putting it together. We have a molecule, we have a thesis that it can be used to treat a certain condition, let’s put this together. What are the objectives and endpoints? What is the schema design? What are the different activities we need to put into this? These are really the main initial users of the system. I do think that over time we’ll get more clinical operations people who do more of the downstream execution of the trial using the system as well, to interface with sites and to think about population enrollment and all these other downstream challenges, because our system is so well-placed to provide insights into those areas. We also think that possibly in the future people in charge of budgeting might use our system as well because we have a really comprehensive view on what’s required to actually support this trial — all the costs and logistics. There are a number of different user personas that we can consider, starting with the clinical scientist. These are usually the first people that touch our system.
Ross Katz: That makes a lot of sense. In terms of developing new features, whether they be AI features or classic software features for the platform, how do you think about prioritizing the use cases that you’re going to try to meet with the tools that are available to you given that there are so many opportunities?
Patrick Leung: It’s challenging because there are so many opportunities. There are so many ideas we have for making our system more capable and delivering value. And there are other features that are not AI-driven as well that are important, like evolving more enterprise-type features to provide fine-grained access control and other things. We have this whole roadmap containing features of many, many different kinds. We have this ROI-based process where we have some estimate of the impact, the relative value of each feature, against some estimate of the cost, and we figure out what are the highest impact features we can be working on. It’s important for a software engineering organization to have this view of what is the stack-ranked list of features in order of value to the company and to our customers that we need to work on. That’s the way we’ve been doing it. And there’s always changes — there are always things that come up where we realize there’s another opportunity, or if we prioritize this other feature that’s new or further down the list, we can actually move the needle on building this business and making our customers successful.
Ross Katz: Right. You’re doing the impact versus effort matrix but there’s always the element of uncertainty in there, and as the organization grows and learns about what the impact actually is or what the effort actually is, it’s constantly evolving. Obviously getting software and AI tools into life sciences is a challenging sales cycle. I’m imagining there’s a teaching effort that’s required in order to get people to wrap their arms around the way things could be. How have clinical trial teams met what Faro is offering, and have you received any pushback on adopting the tools that you’re bringing to bear?
Patrick Leung: The end users themselves are by and large pretty excited. ‘Why would I not want insights really fast on my trial designs? Why would I not want a really organized way for me to model my trials and gain insights into how to make them better?’ I think it’s more just that enterprise sales is complex and there are many different parties, many different stakeholders that need to be a part of that whole process. With the smaller companies, sometimes they can move really fast because they see the need, they really want it, and it’s just easier because they’re smaller and have less process. But with the larger companies there’s often quite a lengthy procurement process, budgeting, all this kind of thing that needs to happen, and so we’re learning as time goes on about ways to really move more swiftly through that process — and of course a lot of it’s relationships. Enterprise sales is different from the mid-market for sure.
Ross Katz: For sure. And what about the regulatory agencies? How are they meeting the type of opportunities that you’re bringing to bear?
Patrick Leung: Of course there’s been a big change there recently, but back at D-Pharma in September, the FDA was actually at the conference speaking and they seemed very open to AI and to new technologies. I do think they had this goal of being able to actually accept digitalized trials for submission the same way that they do in Europe. We’re more than ready for that. That is in many ways the shift that Faro was founded to make — to move from unstructured to really structured representations of trials to make the whole process more efficient. I think the FDA wants to do this, or at least as of September they did. We’ll just have to see how that plays out. In the meantime, document generation remains really important and so we have this dual structured and unstructured system. Everybody wants this to go faster, everybody wants this to go in a more efficient way. I think everyone’s interests are aligned in this — including the patients, certainly the sponsor, and the FDA as well. In principle, we’re all aligned on wanting to do this more efficiently and faster. It’s just that the devil’s in the details and of course safety is a primary concern as well, so there are counterbalancing factors influencing the evolution of this process.
Ross Katz: For sure. Fast forwarding 5 years, 10 years — when so much of the clinical trial design and submission process is automated and there’s optimization of the clinical trial design happening prior to submission that’s saving on cost — what does the biotech landscape look like? How is it changed based on what Faro is doing?
Patrick Leung: If we imagine a world where the process of getting a drug to market is way more systematized and optimized around producing these clinical trials better — to me the dream is that we lower the bar for developing new drugs in the sense that there are all these companies out there that are innovating in various ways, discovering incredible ways to create new medicines, but in many cases they end up on the shelf because just the prospect of going through a 2 billion dollar process to get to market just isn’t worth it. What I hope is going to happen is that a lot more of these really innovative medicines actually get at least tried. Who knows which of them are going to end up being effective and safe, but hopefully there’s a higher volume coming through where every year there are just more and more innovative medicines hitting the market, helping patients, helping to save lives, relieve suffering. This is ultimately what we want to see happen. Hopefully there’s a dramatic increase in terms of the range and the innovation that comes through with new medications.
Ross Katz: I think it’s a beautiful vision and I hope that it comes to pass. I think you’re right that it’s in everyone’s interest that more things are tried, at lower cost to the system as a whole. You’ve come into the organization, an organization that was already building the structure for understanding clinical trial design with this mandate to bring AI to bear. Obviously there’s a lot of competition for talent in AI right now, so can you share some advice on how you think about building strong AI teams especially at the intersection of biotech and software?
Patrick Leung: This is the key — attracting talent. It is a bit of a hires market out there simply because there’s been a lot of layoffs in the technology space and it’s been a tough couple years. But the top talent that really knows how to wield LLMs and launch AI-based features are in demand. My view on this is that you need to have the mission. You need to have some purpose that can distinguish you. I think that life sciences has that in general — the purpose of this industry is to improve human lives. Great people want to work with great people. The interview process is critical for that. You’re not just evaluating someone, you’re also convincing them that this is a place they’d want to work.
Ross Katz: I know that Faro was already a software as a service company when you got there, but were there any cultural barriers that you needed to overcome in order to reach the full adoption of AI for the platform across the organization?
Patrick Leung: Data science is a little different from classical software engineering — it is much more experimental. It’s called science for a reason. You’re basically running experiments as fast as you can to try to figure out what’s actually going to stick and what’s going to work. There’s always a bit of a cultural shift when it comes to this, that might run counter to doing things in a certain way with tons of user testing and really rigorous productionization. That’s all needed once you know that what you’re doing is going to work, but until you reach that point where you have a system that actually works, it’s just move as fast as you can and try a bunch of things and see what works. Introducing that — shifting some of the thinking towards ‘This is not a deterministic process, we need to explore the solution space and figure out what’s going to work and then we go into the productionization and rigorous testing’ — has been really fun to be able to build that capacity within the organization.
Ross Katz: As we head toward the end of our conversation, are there any upcoming milestones or features that you’re looking forward to releasing this year?
Patrick Leung: The document generation is super exciting — we’re already working with customers on doing this in earnest and we’re really looking forward to launching that. I’m super excited about AI-assisted protocol design and all the benefits that could come from that, and certainly there’ll be some announcements and demos. We have D-Pharma coming up in the fall and we’ll be talking about a lot of this stuff then, and there’ll be interim social media posts and other things to announce our various advances in these areas until then.
Ross Katz: For listeners interested to learn more about Faro Health, where would you suggest they start?
Patrick Leung: You can go to our website pharoshealth.com, there’s tons of information there.
Ross Katz: And for people who want to connect with you online?
Patrick Leung: You can go to my LinkedIn. Look up my name, which I guess is going to be listed on this podcast — maybe we can put a link to my LinkedIn in there as well. That’s probably the best way.
Ross Katz: Awesome. Well Patrick, it’s been a pleasure having you on the podcast today, really appreciate the time and look forward to connecting down the line.
Patrick Leung: Thank you so much Ross. I really enjoyed this. This has been a lot of fun.
Jason: And that’s it for this episode of Data in Biotech. If you enjoyed the episode, please subscribe, rate, or leave a review in your podcast platform of choice. See you next time.






