Skip to content
Dr. Jonathan Usuka — Decoding the Dark Proteome with Dr. Jonathan Usuka
Data in BiotechEpisode 41

Decoding the Dark Proteome with Dr. Jonathan Usuka

Dr. Jonathan Usuka discusses the role of proteomics and metabolomics in drug discovery and the opportunity for deeper biological insights.

46:45Full transcript below
DJ

Dr. Jonathan Usuka

CEO at Sapient

Overview

Biotech companies face a critical challenge: traditional genomics and real-world data often fall short in revealing the dynamic, real-time molecular changes key to effective therapies. Genomics offers a static blueprint, while transactional real-world data rarely captures the true biological responses or patient adherence patterns. This gap limits drug target discovery, clinical trial efficacy, and personalized patient care.

In this episode, host Ross Katz speaks with Dr. Jonathan Usuka, CEO of Sapient, who explains how his company addresses this by pioneering deep molecular characterization. Sapient performs longitudinal profiling of 10,000+ proteins and metabolites per patient sample, uncovering the vast “dark proteome” and previously uncharacterized molecular interactions. Dr. Usuka, with a background spanning Roche, Genentech, and McKinsey, offers a unique perspective on bridging advanced science with practical pharmaceutical R&D and patient outcomes.

The conversation explores how this deep molecular data informs new drug target identification, refines clinical trial design, and reveals hidden insights like drug adherence and environmental exposures. Dr. Usuka details how the FDA’s increasing acceptance of molecular mechanisms as endpoints further validates this approach, paving the way for more precise and personalized medicine.

Key Takeaways

Deep Molecular Profiling Reveals Untapped Drug Targets

The vast majority of human proteins—over 99% of the hundreds of thousands present—remain unexplored as drug targets, despite only 900 proteins currently forming the basis of FDA-approved therapies. Deep proteomics and metabolomics reveal these “dark proteome” proteins and uncharacterized metabolites, offering new avenues for therapeutic discovery beyond what genomics alone provides.

Longitudinal Data Unmasks Real-World Patient Behavior

Traditional medical records and claims data fail to capture critical patient behaviors like true drug adherence or environmental exposures. Serial blood sampling and deep molecular profiling expose these hidden factors. This offers a richer understanding of why patients respond differently to therapies, even when demographic and clinical profiles appear identical.

FDA Acceptance of Molecular Mechanisms Accelerates Drug Development

The FDA increasingly accepts molecular mechanisms—such as directly affecting a disease-causing protein—as endpoints for drug approval, even before long-term clinical outcomes are fully known. This shift prioritizes direct molecular measurement. It aligns with deep proteomics and metabolomics to accelerate therapeutic development for conditions like Duchenne muscular dystrophy and Alzheimer’s.

Beyond Raw Data: Contextualizing Molecular Insights Drives Value

Generating vast amounts of molecular data is only the first step. Sapient’s approach focuses on data science methodologies to interpret and contextualize raw proteomics and metabolomics data with a broad human biology database. This provides biological insights, de-risks potential biomarkers, and avoids the struggle many researchers face in interpreting complex molecular profiles on their own.

Related: CorrDyn provides data engineering to build reliable pipelines for complex datasets, helping biotech manufacturers realize data value. Our data assessment services clarify technical strategies, especially for organizations considering advanced machine learning applications in life sciences.

Full Transcript

Jason: Welcome to Data in Biotech, a podcast from CorrDyn where we explore how companies leverage data to drive innovation in life science. Every two weeks, we sit down with an expert from the world of biotechnology to understand how they use data science to solve technical challenges, streamline operations, and further innovation in their business. Today, we sit down with Dr. Jonathan Usuka, CEO of Sapient, to dive into the frontier of proteomics and metabolomics in precision medicine. We explore how Sapient is creating a new paradigm in life sciences by collecting and analyzing molecular data at a depth and scale never seen before. We discuss how longitudinal patient sampling and deep molecular characterization are transforming everything from target discovery to protocol design and even how we think about patient adherence and real-world data. Whether you’re curious about the dark proteome, dynamic biomarkers, or the evolving regulatory landscape, this episode is packed with insights that will get you thinking about the future of biotech in a whole new way. Here we go.

Ross Katz: Dr. Jonathan Usuka, welcome to the Data in Biotech podcast.

Dr. Jonathan Usuka: Ross, it’s really great to be here. I’m a big fan of the show. I’ve had a chance to listen to several of your pods and thanks very much for having me on.

Ross Katz: Awesome. Well, delighted to have you. Just to kick us off, would you mind giving us an introduction to your background and what brought you here?

Dr. Jonathan Usuka: Sure. Very briefly, my background is in bioinformatics and genomics. While at Stanford, I started a company focused on genomics that got acquired by Roche. That’s how I got into pharmaceutical R and D. I was with Roche and Genentech and then Celgene and BMS for a number of years until I went to McKinsey, a global management consulting firm. There, I focused quite a bit on the use of real-world data to advance pharmaceutical pipelines and find new ways to serve patients. About seven, eight years at McKinsey working on a variety of different projects, left to lead a client for an acquisition by Tempus. There, I was the chief strategy and informatics officer. And then over the last year, I joined Sapient, which I think has a really interesting technology. Here, I’m CEO of a company that’s really driving, the true pioneer, I think, in proteomics and metabolomics.

Ross Katz: Awesome. Well, yeah, we’ve had a chance to have Dr. Mojane from Sapient on previously, so some of our audience might be familiar with Sapient from the previous episode, but for those that aren’t, would you mind just giving everyone an introduction to Sapient and the kind of work that you do?

Dr. Jonathan Usuka: The main thing that Sapient does, we focus on doing one thing and doing it very well. And that is to do a deep molecular characterization of disease. Deep disease understanding from a patient sample. This disease understanding at the molecular level is used by pharmaceutical companies to discover new therapies, to advance those therapies into the clinic at a lower cost and at a faster pace, and then ultimately once those therapies are approved, it helps to get the right dose to the right patient with exactly the kind of condition where that patient would benefit from the therapy. So Sapient is all about the deep molecular characterization of disease.

Ross Katz: Yeah, that makes sense. And so this deep molecular characterization of disease, how are you delivering those insights to pharmaceutical companies to help with therapeutic discovery, advancing into the clinic, and getting the right dose to the right patient?

Dr. Jonathan Usuka: Yeah, so specifically what we do is we’re able to measure a lot of the molecules that are floating around in blood or are in a solid tissue sample. There’s an overwhelming number of these molecules. If we think about there being 20,000 human genes, there are millions of proteins — and especially when you get into the different isoforms of proteins, multiple millions — plus the metabolites that result, you’re talking about many millions of uncharacterized, unknown molecules. Much like where genomics was maybe 15, 20 years ago where the data production, the data generation was faster than our ability to interpret, I think we’re at that place right now for proteomics and for metabolomics. We serve pharma companies, biotech researchers, and academic researchers of disease by helping them to understand what is going on in their samples, to see and interpret all of these different molecules and then apply that to either finding new drug targets, to find new biomarkers that show that those drugs are actually working well in a therapeutic way that won’t invoke a safety issue with the patient, and then ultimately to stratify patients so that those patients that will benefit from therapy are on therapy and those patients that would have an adverse event or a safety issue would not be on therapy. We serve really across from discovery all the way through clinical.

Ross Katz: Yeah, and as I understand it, you’re able to create the proteomics and the metabolomics data sets from advancements in mass spectrometry and then also you’re providing, in addition to the data that comes out of that, advanced insights from your team that’s analyzing that data. Am I remembering that correctly?

Dr. Jonathan Usuka: Yeah, that’s exactly right, Ross. What we do is this molecular characterization — we’re able to detect a broad number of metabolites and proteins from a single sample, numbers like 10,000 per sample. And then we put that together often with real-world data, patient data, phenotypic data to get a much better understanding of the disease state of the patient, underlying demographic influences, different exposures the patient might have to either foods, other drugs, or things as simple as weight or diet and exercise. So you can really break out patient response across a continuum, build cohorts to compare different responses and get a much better idea of what is going to work for a patient.

Ross Katz: Yeah, I would love it if you just compare and contrast the data sets that you’re able to provide with the other data sets that pharmaceutical and biotech research companies might acquire to come to some of the same conclusions. Real-world data, genomics profiling.

Dr. Jonathan Usuka: Yeah, great question. First, a quick definition of real-world data. There are two types of data where you’re really measuring how a therapy benefits a patient. There’s the clinical trial information, which is excellent, lockdown, highly regulated, highly rigorous, disciplined execution of a trial where you control for a lot of variables. Key variables that you would control for are often other conditions — so-called comorbidities, other afflictions a patient has, other diseases, and other drugs that the patient is taking. So you have a very pure data set, a very focused data set, often on a very small number of patients. It’s not uncommon for an oncology drug or a rare disease drug to get approved with under 100 patients in the phase three trial. The real world is obviously very different. The patient treatment context is with a bunch of other therapies, a bunch of other conditions, there’s incomplete record taking with the medical records, a lot of it becomes much more about tracking transactions — who should pay for the therapy or the medical visit or the procedure — and is much less about actually measuring patient outcomes. Real-world data is great because it gives you a much higher N, a much higher sample size when something is in the real world versus a lockdown clinical trial. It also gives you the chance to have much more patient benefit because patients aren’t treated in clinical trials, they’re treated in the real world and you have to deal with these comorbidities and these other potential drug-drug interactions. It’s really critical to think in terms of how your therapy is going to perform in the real world and then get as much information as you can to position that therapy, not just market position to make money, but really position it so that it optimizes for the patient response.

Ross Katz: My understanding is that you mentioned real-world data is tracking transactions, but one of the weaknesses of real-world data is that it’s like a snapshot in time. You’re always looking at a record of all the things that have happened, but there isn’t that sense of a sequencing of things. How do the data sets that Sapient provides compare to that?

Dr. Jonathan Usuka: There are a couple things. Known problems with real-world data sets — as I said, they’re to some extent transactional. And by transactional, I don’t mean that in a bad way. I just mean it’s actually highly accurate data when you’re talking about who’s going to pay for a therapy. Or in the case of electronic medical records, it’s transactional, associated usually with the visit in a doctor’s office and what is discussed during that visit. These records, especially medical records and certainly the claims data, don’t capture a lot of things about patient response. You will not see genomics reports, you will not see much in terms of pathology, you won’t see much in the way of radiological images. So often, for example at McKinsey, we would deduce what happens with the patient — what was the genomics readout from a genetic test from the therapeutic course that was pursued. You would see lung cancer, you would deduce that it was ALK positive because you’d see that patient’s therapeutic course be a tyrosine kinase inhibitor, which are given to ALK positive patients. You can do similar things with breast cancer. It’s really looking at things in a very backwards way, looking at the medical record to deduce what the condition of the patient is at a deeper level than you would see from the diagnostic code. What we’ve done is go much deeper, not so much on the medical record side. The EMRs are what they are. They have good information, but they have these limitations. What we’ve done is invest much more deeply in the longitudinal patient sampling and get the molecular characterization at a number of different time points for the patients. It begins to give you opportunities to see disease state and patient response that you fundamentally can’t see through other data sets.

Ross Katz: Yeah, so you’re longitudinally following the same group of individuals and you’re also getting really deep phenotypic information about those patients. What kinds of insights are you able to uncover with that approach?

Dr. Jonathan Usuka: There are insights directly related to our use cases — the discovery of new drug targets or the identification of new dynamic biomarkers that help to recruit patients or choose which patients should be recruited in a clinical trial, or all the way through really predictive response so you know what might happen in a trial when you design a protocol in certain ways. Those are our flat-out use cases or value proposition for pharma research. We see a lot more though in these samples as well as the samples that are provided to us as part of an engagement. Very often we can see things that would make very good molecular signatures of a diagnostic. We see things that are predictive of long-term patient care, not just the acute immediate response to therapy. We’ve done some really interesting work on conditions that are not necessarily indications — done some work in molecular aging, for example, and understanding the biological age of a patient by looking at metabolites. At first I thought this was a bit silly. I was telling my team we can deduce the patient’s age by one field in the medical record, which is date of birth. But then as we dug into it, we really saw different drivers of what appears to be an older or a younger patient based on different lifestyle factors, different medical conditions. I think we’re really getting at drivers of age more than just a proxy for date of birth in the medical record. It’s an exciting time for proteomics and metabolomics because you’re pioneering these deep views of molecular characterization, whereas before you might measure one or two molecules that you had a hypothesis about. Now you can be hypothesis agnostic. You can really do brute force, look at everything across a large number of patients and see what’s different.

Ross Katz: Yeah, it’s really interesting. I just want to make sure that I understand why the proteomics and metabolomics data sets allow you to draw these sorts of insights. Basically you’re looking at the same patient over multiple time points. The new drug targets come out of the idea that a lot of drugs are targeting proteins, so you have the proteomic profile so you can see where the potential binders are. Predictive response is the idea that you see what happens at a certain time point when they receive the treatment and so you can watch what happens over time and distinguish from a group of patients that received the treatment from people who didn’t. And then the very good molecular signatures — it’s that you just have this broader array of data about what’s happening biologically inside of these different patients and the complexity of the interactions that you can track is much deeper. As a data person, that’s just how I think through the issue, but would you mind correcting or verifying my mental model of how the data set allows you to do these sorts of things?

Dr. Jonathan Usuka: No, Ross, you have a lot of the precursors to what’s actually going on. If we talk about the central dogma of biology — all that is is this idea that DNA gets transcribed into RNA and then translated into a protein. And the protein is what does essentially all of the catalysis, causes the biological reactions to happen in cells. Little bit of RNA catalysis, but very little. It’s mainly proteins that are doing the work of operating a cell. That is of course all guided by what is in the genome, but the genome DNA is fairly static. I think Illumina’s work 10, 15 years ago taking their methods and instruments for measuring DNA directly — reading DNA — and then applying that technique called RNA seq so that they could understand gene expression through measuring RNA, the number of copies of this RNA, was really groundbreaking stuff, but it was just beginning to scratch the surface of the dynamism in a cell. To move from that static DNA view to much more of a dynamic, what is happening now — how is the cell or a group of cells or an organ responding to an agent, either a disease agent or a therapy. When you’re measuring proteins, you’re measuring directly what is doing the catalysis in the cell, what is changing the biochemical profile. You’re also in a world that is much richer than 20,000 genes. As I mentioned, 20,000 genes give rise to — with different spellings — up to like 100,000 different proteins, and then with the ways proteins can be modified after translation, it’s called post-translational modifications, that gets up into the millions of different proteins. We don’t even know right now how many proteins of different forms there are, even though they all come from what we know very well, these 20,000 human genes. Digging into this protein space, which is where drug targets are usually focused — we have, as I mentioned, 20,000 human genes, hundreds of thousands of proteins, and yet Ross, only about 900 of those proteins have been used as a drug target, at least for an FDA approved therapy. There is a huge opportunity going forward — we’re talking about the vast majority, 99 percent of proteins, functions not really known and the drugability, the result of those proteins being drugged, is wide open in terms of therapeutic value for patients. It’s absolutely critical that we understand what those proteins are doing at a deeper level, and then the reactions that they catalyze — those are the metabolites that we’re measuring. You can see not only the proteins directly, are they present, are they knocked down, are they amplified in certain conditions or in response to certain therapies, you can also see the results of those protein catalytic reactions, which are these metabolites, these different analytes, pick those up and get a much more detailed view of what’s happening in the cell.

Ross Katz: Yeah, that’s fantastic. So can you walk us through how Sapient collects and analyzes patient data across all these dimensions?

Dr. Jonathan Usuka: Yeah, so we do a couple things. We have collaborations with academic medical centers here in Southern California, some large centers. We run our own recruitment efforts. We’re particularly trying to build a data set that is not huge. We were talking about these real-world data sets that are based on drug transactions or payers or scripts tracking — those have hundreds of millions of patients in them. Or large electronic medical records, something like an Explorist data set from IBM — those have millions of patients in them. We’re very much focused on a subset of those, 20 or so thousand patients, but we want to characterize those very deeply across 20,000 metabolites and proteins. The choice of those patients becomes absolutely critical. We want to have a diversity in patients, both in terms of demographics, disease types, age, gender, different exposures so that when we design cohorts, we can really account for some of the confounding variables that would limit how robust our conclusions would be.

Ross Katz: Yeah, that makes a lot of sense and it strikes me that when you’re trying to select patients for inclusion into this very deep profiling across multiple time points, you have to consider a variety of factors, but also the level of statistical power or the sample size of patients that you want to have across combinations of all of these different dimensions. Can you give us a little bit of insight into how you account for all these factors in making decisions about who to incorporate?

Dr. Jonathan Usuka: Well, there are two things. There’s who we incorporate and then how we use that patient data in developing cohorts. In terms of who we incorporate, it’s really important to us that we have a cross section of important diseases. Everything from cardiometabolic, where you often require large patient numbers — we like those diseases not because of the large patient numbers required but because often the key outcomes are lab tests that are in the fields of an electronic medical record. I was talking about how radiology reports, images, certain genomics reports are not in the electronic medical record — it’s hard or can be very hard to get to. A lot of the things in cardiometabolic disease are right there. You’re measuring blood pressure, you’re measuring weight. We see a lot of opportunity in areas like GLP-1 in terms of understanding cardiology risks and cardiology treatments, not because we have the large patient size but because the outcomes are readily available. Separately, we’ve been investing in building up solid tumor and liquid tumor data sets, which is directly understanding the patient response to therapy through the EMR, seeing progression-free survival and then directly sequencing or doing the protein profile of those patients. We want to make sure we have a good mix of disease. Even with rare disease — rare disease is obviously rare, it’s hard to find patients that fit the criteria. The phenotypes are often complex. You put together a number of different diagnostic codes because many of these patients have been misdiagnosed. But even when you have a small number of those patients, they are never deeply characterized the way Sapient characterizes our patients. You never have a metabolic profile of 10,000 metabolites or 5, 10,000 proteins. We’re looking for a good mix of indications and a good mix of demographics for the choice of our patients.

Ross Katz: What are some of the biggest challenges that you face in working with this kind of data, gathering the depth of insights that you’re trying to gather across this population which, obviously compared to real-world data it’s relatively small, but still 20,000 plus patients over time is no small feat.

Dr. Jonathan Usuka: Yeah, and to be clear, the 20,000 patients, as you point out, is very tractable now. What is our challenge and where we’re really pushing the envelope is the deep data cube we’ve built around each patient. Each patient has been deeply characterized across metabolites, across proteins, increasingly around their genomics as well. The challenge is building the correct cohort so you compare for disease and then really understanding if your conclusion is robust. I’m a big believer and practitioner of finding robust conclusions from incomplete data. What I do is change variables slightly, build slightly different cohorts and then see if your conclusions hold up. If they do, then you have a robust conclusion that it’s worth building a drug program around, or changing the trajectory of a clinical trial, or building it into the protocol of the next trial. If it doesn’t stand up to that kind of robustness test, then you’ve got to look more deeply at how you actually constructed the experiment.

Ross Katz: Yeah, that makes a lot of sense. Have you found anything particularly surprising on any of your projects extracting insights from this data?

Dr. Jonathan Usuka: Wow, yeah, a couple things. First, it’s exciting to be in a time where so much is unknown. You feel like because of the advances in the Human Genome Project and the deep characterization we have now and understanding how we’ve basically cataloged the human genome, you feel like protein space should really be very well understood because they match — there’s this one-to-many match, but at least there is a match. There’s so much unknown about proteins. And they play so many different roles in so many different diseases as well as healthy states. We’re really interested in the dark proteome, the number of proteins that haven’t been detected and have unknown function. A big part of what we do is identify dynamic biomarkers for pharmaceutical companies. They will have a set of human samples, often from a clinical trial — they’ll have normal healthies and patients that have been treated with the therapy and are responding or are not responding. We dig in there and we find a small number of metabolites that really are highly predictive of patient response. Inevitably, the pharmaceutical company is saying something like great, we’re going to be changing a $100 million clinical trial and redoing a protocol, what are those molecules? And the thing is, unlike genes, metabolites are largely uncharacterized. You don’t even know what these chemicals are, what they do. We’ve done some major advances in two things associated with metabolites. First is our metabolite ID capabilities through AI, where we’re able to identify chemical structures through some pretty interesting methodologies that haven’t been characterized before. And then secondly, we’re able to take that subset of metabolites identified in the clinical trial, look in our human biology database and understand where those metabolites have occurred in other diseases, other patient types, other ethnicities. Do they invoke an immune response? Are they associated with particularly dramatic outcomes? We’re able to de-risk what is otherwise a pretty unknown, risky thing to change a protocol around or to modify a clinical trial around.

Ross Katz: So there’s the dark proteome, the uncharacterized proteins, and there’s also the largely uncharacterized metabolites, but because you’re deeply profiling all of these other patients, it’s not just that you’re seeing something you’ve never seen before from a small set of patients in a clinical trial — you’re able to contextualize what you’re seeing in the broader diversity of biology that’s out there, both related and unrelated to the disease that you’re studying. Am I thinking about that right?

Dr. Jonathan Usuka: Contextualize is exactly the right word, Ross, and we really find that we’re helping the patients and the pharmaceutical companies the most when we’re delivering biological insights, not just raw data. It’s the state of where the industry is now — we can produce a lot more data than we can interpret. Very much like genomics from two decades ago. Sapient is all about applying data science methodologies to understand and interpret that data and provide the biological insights and the context from other patients, not just the raw data that pharma or other researchers really struggle with interpreting.

Ross Katz: Since you mentioned it — your ability to identify molecular structures through computational approaches, and the advancements in data science approaches that are out there — could you share some insight into the types of computational approaches that you’re applying and where you see the most opportunity in the kinds of insights that they can drive?

Dr. Jonathan Usuka: Sure. First, if we go with the information science aspect of it before we even get into comparison algorithms — what information is known? We sketched it in this conversation. Genes, fairly well characterized. Canonical proteins, for the most part, well characterized. All these protein isoforms are all over the place. And in terms of the metabolites, as you correctly summarized, very little known outside of the standard blood analytes that get measured for any doctor’s visit. Then there’s one other thing as you get into more the real-world or medical or phenotypic data — ICD-10 codes, the diagnostic codes, are pretty good, but they’re designed for payers, designed around transactions. Trying to deduce things around medical condition from the diagnosis code only takes you so far. We’ve invested quite a bit in a few things. Obviously bring in publicly available data about genes and proteins, but everybody does that. Another thing we’ve done is really invest in curation. We’ve been working with a local company in San Diego, Rancho BioSciences, to curate our data set, really understand and come up with an ontology that works for all of the different treatments — and there are a lot of different treatments the 20,000 patients can take — and that works for all of the different diagnoses and building complex phenotypes and diagnoses so it becomes much more tractable. Then, once you have that data foundation and we’ve joined a number of different tables from these diverse data sets, you’re able to apply machine learning and do some fairly advanced algorithm development. We use, of course, large language models. We’re very much about empirically detecting signals, not necessarily having hypotheses that impute what a good signal should be or what is likely to be the biomarker. We’re very agnostic and empirical when it comes to our large data sets and detecting what a good signal would be. And then there’s a key aspect to harmonizing and tuning the data sets so that they are readily served up for machine learning. That really goes to understanding biases in the data sets and some of the big advantages we have in terms of how we did the data generation.

Ross Katz: Can you provide some insight into how the data that you’re collecting influences not just the discovery and the clinical trial side, but also patient outcomes and care and what’s happening in the clinic?

Dr. Jonathan Usuka: It’s interesting. The FDA has evolved so much in this area. We were talking earlier about the acceptance of real-world data. Another thing that I don’t think gets as much attention is how much the FDA has moved in terms of accepting outcomes that are not directly traditional patient outcomes. One of the big ones is the fact that they are accepting a molecular mechanism — if you demonstrate that you are affecting the protein that is believed to be disease-causing, you don’t have to wait and do follow-up studies over years, you don’t need that outcome information for the drug to be approved. The FDA is much more comfortable with real-world data generation after the fact to show efficacy, and they’re focusing much more on the molecular basis of the disease. A couple examples: Duchenne muscular dystrophy — the FDA is approving therapies that knock down dystrophin for certain types, certain spellings of the gene of dystrophin, which is believed to be effective in treating patients. The FDA has approved these drugs that directly interact with dystrophin and then over time is measuring the disease progression in these patients over years. That’s a very different mindset from the FDA of 10 years ago where you had to have the patient outcome. Another huge area is Alzheimer’s. These therapies largely are getting approved on the tau hypothesis — a protein tau and its role in Alzheimer’s. These drugs are getting approved based on the fact that they interact and knock down tau or augment it in certain ways. Then we’re tracking these patients over time to see progression of Alzheimer’s, and this is over 10, 20 years that you’d be watching this. It’s a very different FDA, which enables a very different approach to how you think about outcomes. If you have the disease hypothesis that gives you a mechanism for modifying the disease state, as we just talked about in Duchenne muscular dystrophy or in Alzheimer’s, then the protein measurement is the key aspect for unlocking the approval. It’s just a really exciting time to be working with proteins with an agency that really starts to understand the molecular basis of disease.

Ross Katz: So it sounds like the FDA has taken a more welcome view of real-world data and the insights that can be derived from it, and the logic of what you just described plays very naturally into the longitudinal data sets that Sapient is gathering. Have you gotten specific signals from the FDA about the applicability of Sapient’s data sets to these kinds of problems and the value that they can provide?

Dr. Jonathan Usuka: Yeah, so we are a services company supporting pharmaceutical development. We will be in a regulatory package that they will submit to the FDA. Often it’s translational — initiating entry into humans, initiating an early phase clinical trial — or it’s part of the regulatory package for approval. Another big thing that we are not focused on now, but we’ve seen some amazing sources of value that we’d look to partner with other organizations on. We see really good evidence of diagnostic markers, good markers of disease progression, treatment, patient response to treatment, and then stratifying disease state. We’ve done some really good work with the Bill and Melinda Gates Foundation, which is public. Our work with the Bay Area Lyme Foundation, which is focused on long Lyme disease versus acute and really understanding markers of that, which can change treatment and also how you think about paying for treatment. There are a lot of diagnostic opportunities in our data that we’re just not able to focus on now because we’re really sprinting against the pharma use cases.

Ross Katz: Right. It sounds like there’s a lot of opportunity for increased personalization of medicine using the data that you’re providing. Am I thinking about that right?

Dr. Jonathan Usuka: That’s exactly right. I’m leading a small organization, we’re under 50 people. We have this amazing data set. You can’t do everything at once. We’re really focused on those pharma use cases, focusing on understanding disease — new targets — and understanding the patient response to treatment, which is that translational and clinical use case. We will get to those diagnostic and other use cases eventually. Just don’t want to take our eye off the ball right now with what we’re working on.

Ross Katz: For sure. That makes a lot of sense. One more question and then we’ll go to some more forward-looking things as we come toward the end of our conversation. You’ve got this comprehensive approach to data collection that’s very longitudinal in nature — you’ve followed up with a patient in the past and you have this plan to follow up with the same patient in the future. I’m interested in how this approach to deep patient profiling influences the type of relationships that you create with pharmaceutical companies and the value that that provides.

Dr. Jonathan Usuka: Ross, we haven’t talked a lot about it, but when you do longitudinal sampling — and this is sampling meaning you’re getting blood draws and plasma over time — it really changes what you can see in the data and really makes some very tricky, persistent challenges in pharmaceutical interpretation much more doable. A couple examples: drug adherence. Adherence is the idea of a patient in the EMR who is taking a therapy. A patient, based on the scripts data — meaning the data of the transactions with the payer, insurance data — is picking up that therapy from their drugstore. They may or may not actually be taking the therapy. And this is true for chronic conditions, true for acute conditions, shockingly true for really serious conditions like oncology. People just are non-adherent for different reasons at different times. A big part of my work at McKinsey was trying to understand why patients go non-adherent. And they even do this in clinical trials. Even though you are a participant in a clinical trial, it doesn’t mean that you are taking that experimental therapeutic at the time or in the dose or with the frequency that you’re supposed to. Huge challenge in terms of interpreting the results in clinical trials, huge challenge for pharmaceutical companies — not really because they want to sell more, though yes, they do, but it’s also about understanding why a patient would go non-adherent or switch to a competitor’s brand. It’s often because that patient is experiencing an adverse event or there’s a drug-drug interaction that makes it untolerable for the patient. All these things you can’t see just from the medical records, you can’t see it from a questionnaire with a patient because they’ll say yes, you see them in the medical record. If you look at serial blood sampling, longitudinal sampling from the patient, you can see these adherence and non-adherence patterns. And not only that, you can see potential reasons for that non-adherence — things like other conditions, comorbidities, other drug interactions, or just other stresses in the patient’s life, because you can pick a lot of those things up by the analytes, the metabolic profile. It opens up this entire window into understanding the motivations of a patient that you just fundamentally don’t have from the EMR or from other real-world data sets. Stuff like that is just so exciting to be able to see.

Ross Katz: Right. Because of the foundational quality of the data sets that you’re providing, you’re able to open the range of questions that can be asked and hypotheses that can be generated and then tested in analytically rigorous ways beyond just querying electronic medical records.

Dr. Jonathan Usuka: Yeah, that’s right. And these are questions we have always had, but we’ve always had such relatively poor data because the data is correct as answered but doesn’t reflect the actual situation. The therapy’s actually picked up from the drugstore, it’s just not swallowed. The other things you can see are much more related to the exposome. The exposome is what kind of exposures does the patient have outside of things you would see in the medical record or in any way associated with a therapeutic. You tell your doctor you’re a non-smoker, you hardly ever drink, you don’t do illegal drugs, you exercise every day. All of that’s noted in the medical record. Now, is it true? I’m sure in your case it is, Ross. My case it certainly is. But maybe other people aren’t so truthful with their doctor. That is a fact in the EMR, it goes as a fact for any analysis, but if you look at the metabolic profile of the patient at a certain time, you can see exactly whether they’re smoking, whether they’re drinking, whether they’re doing illegal drugs, whether they’re doing as much exercise as they say — 100 other exposures, environmental toxins. All of these things can be picked up with a kind of fidelity and accuracy that weren’t really talked about up until recently. You can do corrections in cohorts, understand treatment outcome differences, and the whole game in pharmaceuticals is to understand two patients that seem exactly the same yet have different outcomes. Why? That’s the key question. If you can get into questions around adherence — did they actually take the medications — and exposures — they seem to have the same lifestyle, they’re the same age, they live near each other, they’re the same gender, they have the same other conditions, and yet one is exercising and the other is not and is a smoker — you’re going to get those different outcomes, which you would never see just from the medical record or from a patient survey.

Ross Katz: As we head toward the end of our conversation, just looking at the rapid pace of innovation in biotech, how do you see the role of data evolving over the next three to five years?

Dr. Jonathan Usuka: We talked a lot about the kinds of questions that people have always had, but there was never the data to answer. Over the next five years, I think we are going to have a much deeper biological canon — a much deeper set of information about proteins and metabolites and their role in disease causality and response to treatment. And I think we’re also going to get much more comfortable with this idea of truthfulness and accuracy in the outcomes data. Things that you can directly detect, like the exposures I was just describing. That has important ramifications for payers, insurance, risk profiles of patients, but it also has tremendous opportunity to develop therapeutics. It’s just a really exciting time to be working with this kind of data.

Ross Katz: For sure. And as you look toward the future for Sapient, how do you think that Sapient could change the way that patients experience healthcare in the future?

Dr. Jonathan Usuka: The big one is going to be that patients are going to be getting a unique therapy in the right dose for their particular condition. And before they get into the therapeutic course, we’re really going to understand problems with drug-drug interaction. We’re not going to discover that stuff in the real world 10 years later that they shouldn’t take it with one or the other. That stuff right now, we do by understanding the biochemistry and the mechanism of action of a single drug. I think in the future we’re going to have a much more combinatorial approach to this and a much more empirical approach to it. Patients are going to be on the right therapy from the beginning and they’re going to experience much fewer adverse events because they’re on that right therapy.

Ross Katz: Awesome. And for listeners intrigued by Sapient, where should they go if they want to learn more?

Dr. Jonathan Usuka: We’re a San Diego-based company, online we’re sapient.bio. We’ve been around for about four years and have increasingly started publishing our results and collaborating with other researchers.

Ross Katz: Awesome. And for people who want to connect with you, where should they find you?

Dr. Jonathan Usuka: Probably the best, most stable is always LinkedIn. I’m Jonathan Usuka, U S U K A. Would love to connect with some of your listeners.

Ross Katz: Fantastic. Well, Dr. Usuka, it’s been a pleasure to have you on today. Really appreciate the time and look forward to connecting down the line.

Dr. Jonathan Usuka: Ross, great conversation, great questions, and really enjoyed talking to you.

Jason: And that’s it for this episode of Data in Biotech. If you enjoyed the episode, please subscribe, rate, or leave a review in your podcast platform of choice. See you next time.

Frequently Asked
Questions

How does investing in deep molecular profiling impact competitive advantage and ROI for a pharmaceutical company?
Deep molecular profiling uncovers novel drug targets and dynamic biomarkers, leading to a higher probability of clinical trial success and reducing development costs. By identifying non-responders earlier and understanding true patient adherence, companies can bring effective therapies to market faster and with greater precision.
What are the infrastructure and data management considerations for handling such large-scale proteomics and metabolomics datasets?
Managing millions of uncharacterized molecules per patient across thousands of patients requires robust data curation, harmonization, and advanced data engineering. Building correct, statistically powerful cohorts demands careful consideration of confounding variables and a flexible data foundation for machine learning applications.
How can these deep molecular insights be integrated into existing clinical workflows to improve patient care?
The data can define molecular signatures for diagnostics, predict long-term patient care outcomes, and stratify patients for personalized treatment. While not Sapient's direct focus now, these insights hold promise for identifying optimal drug combinations and minimizing adverse events in clinical settings.

Need a data partner for life sciences?

CorrDyn helps biotech and pharma companies build the data infrastructure that accelerates research and operations.

Book an intro call