Skip to content
Data in BiotechEpisode 29

The Evolution of Genomic Analysis with Sapient

Mo Jain, Founder and CEO of Sapient, on biomarker discovery beyond the genome and accelerating precision medicine for biopharma sponsors.

43:15Full transcript below
MJ

Mo Jain

Founder and CEO at Sapient

Overview


The human genome was supposed to rewrite medicine. Two decades after the first sequence was completed, genomic data explains only 14 to 16 percent of population attributable risk across human disease. The other 84 to 86 percent is encoded in what people eat, breathe, drink, and carry in their microbiome: signals that live in blood chemistry, not DNA. For biopharma sponsors building drug development programs, that gap is the difference between a Phase 3 that reads out and one that doesn’t.

In this episode, host Ross Katz talks with Mo Jain, Founder and CEO of Sapient. Mo is an MD/PhD with clinical training in internal medicine, cardiology, and preventive cardiology, a molecular physiology PhD, and postdoctoral work in mass spectrometry and large data handling at the Broad Institute at MIT. He spent 15 years in academia as a professor in Boston and at UC San Diego before spinning Sapient out of the University of California in 2019. The conversation covers the case for dynamic markers over static genomics, why proteomics, metabolomics, and lipidomics each carry different biological signal, and how Sapient runs a fully robotic pipeline from freezer to mass spectrometer with more than 500 QC parameters monitored in real time. Mo also walks through Sapient’s relational database of plasma, tumor, and longitudinal outcomes samples, and how biopharma teams use it for target identification, drug screening, preclinical safety, and patient stratification.

Key Takeaways

The genome explains 14 to 16 percent of human disease; dynamic markers carry the rest

If you sequenced every person on the planet and matched it to their medical record, genomic data would account for 14 to 16 percent of population attributable risk. The remaining signal lives in small and large molecule chemistry in blood: metabolites, lipids, and proteins that reflect diet, environment, microbiome, and cell-to-cell communication. Mass spectrometry measures these dynamic markers, moving clinical interrogation from 20 molecules per blood draw to more than 20,000.

Data quality beats data quantity, with 500 QC parameters monitored in real time

Good discovery does not come from the best mathematicians working on large data; it comes from the best mathematicians working on clean data. Sapient monitors more than 500 parameters in real time on every mass spectrometry run, covering sample integrity, chromatography, instrumentation, and extraction. The full pipeline is robotic end-to-end: freezers, pipetting, capping, protein and metabolite extraction, mass spec acquisition, and data processing. That is how QC stays consistent across hundreds of thousands of samples.

A relational multi-omics database amplifies a small Phase 1 cohort

A Phase 1 trial with 100 samples is enough to start a hypothesis, not to make a definitive call on drug response across several hundred thousand patients. Sapient homogenizes proteomics, metabolomics, lipidomics, genomics, microbiome, clinical outcomes, diet, and lifestyle data into one searchable relational database, covering hundreds of thousands of plasma samples with longitudinal follow-up and thousands of tumor samples across diverse cancers. Biopharma clients query that reference data to validate targets, check tissue expression profiles for safety, and contextualize trial findings.

Full Transcript

Jason: Hi everyone, this is Jason, producer of Data in Biotech. Before we get started, I wanted to let you know about our latest white paper. It’s a comprehensive guide to implementing machine learning models in biotech manufacturing. It’s a complete overview of all the potential problems of ML adoption and, more importantly, how to solve them. To download it, simply visit connect.corrdyn.com/biotech-ml. We’ve also dropped the link in the show notes of this episode. Okay, let’s get into it. Welcome to Data in Biotech, a podcast from CorrDyn where we explore how companies leverage data to drive innovation in life sciences. Every two weeks, we sit down with an expert from the world of biotechnology to understand how they’re using data science to solve technical challenges, streamline operations, and further innovation in their business. Joining us today is Mo Jain, CEO of Sapient, biomarker discovery organization that enables biopharma sponsors to go beyond the genome to accelerate precision drug development. In this episode, we discuss how Sapient leverages high-throughput mass spectrometry to generate biological data, the role of AI and machine learning in enhancing data quality, managing complex datasets, and classifying disease states with high precision, and why data-driven approaches are vital for innovation. Here we go.

Ross Katz: Mo Jain, welcome to the Data in Biotech podcast.

Mo Jain: Thank you so much, Ross. A pleasure to be here today.

Ross Katz: Well, to kick us off, can you give us an introduction to your background and what brought you here today?

Mo Jain: Sure. I’ll apologize in advance because it’s a background that is somewhat long and windy as I continue to try to figure out what I want to do when I grow up. Started at a young age, was very interested in science and medicine. I had been enrolled in one of these accelerated medical programs, pre-accepted to medical school through undergrad and then transitioned into MD PhD, did my formal medical education, did a PhD in molecular physiology, went on to clinical training that I absolutely loved in internal medicine and cardiology and preventive cardiology, did my full postdoctoral work at the Broad Institute at MIT in mass spectrometry and large data handling, had some incredible experiences there and was an academic for the last 15 years or so as a professor, first in Boston and then here at the University of California San Diego, loved every minute of academia, the collegiality, the community, the discovery, working with residents and fellows and in the hospital, and then evolved through that to launch Sapient approximately four years ago where I’ve been working on the industry side for the last several years and really enjoying every minute of it.

Ross Katz: Well, with that as the jumping-off point, can you set the scene and introduce us to Sapient?

Mo Jain: Absolutely. Sapient was spun out of the University of California approximately four years ago and was based upon this idea that to realize the potential of what genomics was meant to, with the idea that finally we have means now of interrogating on a very deep level human biology, measuring thousands of factors and being able to leverage that information in order to accelerate the development of drugs and therapeutics as well as diagnostics for patients. That was the underlying concept. Where Sapient was quite a bit different is in the type of data that we were playing with. With the idea that genomics was meant to really transform medicine, in many ways it’s transformed our understanding of what it means to be human and the evolution of humans, but not necessarily has transformed medicine. What we’ve done is doubled down on the alternative side to this, which is what we call dynamic markers, mass spectrometry-based measurements of alternative areas that serve to complement and allow us to go beyond genomics for discovery.

Ross Katz: Interesting. Can you just give us a little background on where you’ve gone since the genomics origins and your founding in 2019?

Mo Jain: Sure. I’ll go one step behind that. Going back to the dawn of genomics, I had the pleasure of being in Boston in 2003 when the human genome was just being completed and parallel sequencing was coming on board. I remember having a discussion with a number of scientists in Boston at the time around this idea — this appears to be transformative technology. We can finally sequence genomes at scale. Assuming we line every single person up on the planet and they’re holding their medical records in their hand, and we sequence every single individual and we have their full medical records cataloged, how much of human disease can we explain? This is a metric that we call population attributable risk. There’s many ways in which you can do the underlying calculation. We all want that number to be 80, 90, 95%, but in reality that number’s probably on the order 14 to 15 to 16%, which means that even if we fully understand the human genome, there’s a massive amount of information that’s missing. It’s almost akin to solving a puzzle in which you only have 10 to 14% of the pieces. It’s virtually impossible, especially for a complex puzzle such as the human body. I became very interested in that other side of the puzzle, those types of risk factors that are not captured in heredity from your mother and father, but that come from the world in which we live, everything we eat, drink, smell, smoke, and the way we live our life. We know that has huge implications for human disease. That information source is not encoded in your genomic sequence. Your genomic sequence is set from the moment of conception for the most part. Rather, it’s continually evolving or dynamic through life and it’s encoded in small molecule and large molecule chemistry in human blood with the idea that the blood serves as the conduit for the communication trail with the external world as well as the internal world, how our liver communicates with our brain, how our fat cell communicates with our heart, etc., as well as our interaction with that external world, everything we eat and drink and smell and smoke, the microbes in our gut, etc. All of that information is encoded in this type of chemistry. That area of chemistry is captured by technology called mass spectrometry. That was the conception of Sapient. We spent the last 10 years developing these technologies, this garnered some interest from large foundations, government organizations, NIH, large pharmaceutical companies, and this is what gave rise to Sapient spinning out approximately four years ago.

Ross Katz: It’s an amazing story and I want to spend some time on mass spec a little bit later because it seems connected to the entire story. But before we get there, can you walk us through the types of data that you generate at Sapient? I’m imagining we’re talking about that other 84 to 86%, the things that go beyond genomics as you were just discussing.

Mo Jain: That’s exactly right. You can imagine this again includes the world in which we live, both our internal world as well as our external world, that has profound impact on human health and disease over time. Thinking through this data asset, one of the clear lessons over the last 10 years as large data has come online across multiple areas of healthcare and drug development is that not all data is created equal. Having large data in itself is not the solution. It’s about having the right type of data and having that right type of data at scale. There’s several million human sequences that are available in the world, sequencing that one additional person starts to become diminishing in returns. We spend a lot of time thinking about how we complement genomic sequence, what are those dynamic markers that read out our internal health as well as our external world in which we’re interacting. This falls into the small and large molecule chemistry that’s captured by mass spectrometry as a technology. This includes all the metabolites, all the lipids, and all the proteins that are within the human circulation, within our tissues, within our blood and urine, CSF and other bio-repositories within the self.

Ross Katz: Right. That’s the metabolome, the lipidome, and the proteome that we’re talking about here. Can you walk us through why do you generate each type of data and then I’d like to go into the how?

Mo Jain: Absolutely. Each one of these is somewhat orthogonal and there’s slightly different reasons and information streams that are captured within each of these omes. Let’s start with the underlying proteome. The proteome represents the business end of the machinery that’s made by every one of these human cells. If you think about drug development as a whole, 98% of drugs target proteins. Being able to measure the actual proteins tells us what is physically dysregulated in a cell during a process of disease. Being able to identify those and measure those at scale now provides an enormous amount of information regarding the next batch of drug targets. To date, if we think about how target identification has been performed, it’s largely come from genomics information as well as RNAseq and transcriptomics information, which in many ways are just serving as surrogates for the underlying proteome. Now that we can finally interrogate proteins at scale, comprehensively, in a tissue, in blood, in various biospecimens, and do that at scale across thousands of samples at once, we no longer have to use surrogate markers for what is the best drug target, we can actually measure quantitatively those proteins that provide an enormous amount of information regarding the actual targets.

Ross Katz: Yeah, that makes a lot of sense. Okay, can we talk about the metabolites and lipids?

Mo Jain: Absolutely. That’s one core piece of information. When we think about metabolomics and lipidomics, in many ways these are the communication streams that occur between cells and the internal world as well as our external world. Metabolites and lipids are among the most ancient of biological material. This is how bacteria communicate with one another. This is how they sense their underlying environment, going all the way back to single-cellular organisms. In the same way, metabolomics and lipidomics captures a tremendous amount of information regarding our external inputs, everything we eat and drink, the way our microbes work in our gut is captured in these metabolomic markers. The way one cell in our body communicates with another cell in our body is partially through proteins but is actually more communicated through these metabolites and lipids and small molecules. These molecules are made within our cellular compartments and are made to be excreted into our central circulation where they travel around and are sensed by other cell types within the body. Being able to capture that information provides an enormous amount of underlying data regarding the state of health in our various cells as well as interrogating the communication stream between these organs and organelles, and that’s helpful for drug development.

Ross Katz: That makes a lot of sense. When you take all of these things together, you’re getting as much as possible a more 360-degree view closer to that 100% that you’re targeting of what’s actually happening inside of the body, and that gives you more data, more evidence with which to draw conclusions about the mechanisms of drugs, the biomarkers of drugs, things like that. Am I thinking about that right?

Mo Jain: That’s exactly right. The analogy I use is when you go to the doctor every year, Ross, they draw two tubes of blood. In those tubes of blood, we measure 20 different metabolites, lipids, and proteins. That represents a huge portion of our underlying diagnostic capabilities that allow us to tell how our liver and our kidney and our heart and our brain, etc., are functioning as well as what is going to happen to us over time. There’s 20,000 of those molecules floating around in your body. We’ve now just gone from a 20 diagnostic marker to a 20,000 diagnostic marker in a 1,000-xing the information provides an enormous amount of resource now to understand health, understand disease processes, identify new drug targets, as well as help develop drugs that do what they should.

Ross Katz: Awesome. Can we jump to the how now? How does Sapient go about generating these different types of data?

Mo Jain: The idea behind these orthogonal data assets, whether it be metabolomics, lipidomics, or proteomics, as a way to complement genomics for drug development is not a new idea. These ideas have been around for a long time. The underlying challenge has always been how do we go about doing this from a tactical perspective. Mass spectrometry is a technology that’s been around for many decades, it’s unfortunately been too slow a technology to actually apply at scale. There’s two components here, one being able to assay thousands of these molecules, the breadth and depth of measurement, and the ability to scale this to thousands of samples at a given time. This is where Sapient has spent quite a bit of its technology innovation, around developing those technologies now that allow us to go much faster than is typical to be able to assay thousands of samples at a given time and in doing so be able to measure tens of thousands of these metabolites, lipids, and proteins in a single biological specimen and be able to do that everything from preclinical samples to cellular samples for drug screening to large human studies.

Ross Katz: Correct me if I’m wrong, mass spec by its nature is this flexible assay that can measure all of these different types of things, the metabolites, the lipids, the proteins. Am I thinking about that right?

Mo Jain: That’s exactly right. There’s been a number of advances in mass spectrometry technologies over the last 10 years and particularly over the last two where mass spectrometers from a hardware perspective have gotten more sensitive and faster. You can imagine that there’s additional complementary components to this. How we process samples in high throughput, how we do chromatographic separation of samples as they’re introduced into a mass spec, and then how we handle the data on the back end. Having one component be better doesn’t necessarily allow the entire pathway or approach to be better. What we’ve been able to do is leverage that innovation in technologies and support it with innovative aspects around sample processing and robotics, support it with advanced chromatography and support it with advanced software on the back end that allow us to accelerate the entire discovery pipeline from a sample to data.

Ross Katz: Interesting. Now that we understand the data that you’re collecting and how you collect it, can we talk a little bit about the types of problems that you solve? Why do pharma and biotech companies come to you? And if there’s another category of work that you’re doing, we’d be interested in that as well.

Mo Jain: Absolutely, Ross. There’s many particular applications for this. This fundamentally represents a tool the same way next-generation sequencing is a tool that can be applied in many different ways. I’ll give you some potential use cases in which these technologies and our approaches can support biopharma across the entire drug development pipeline. From very early on in being able to take very large sample numbers and identify what are the ideal targets to go after for my drug development program. Whether it be in oncology or immunology, neurodegeneration, inflammatory and fibrotic diseases, cardiometabolic disease. In essence, now pharma has the means to drug many different targets that we never could go after before. Particularly with the emerging therapeutics around ADCs and T-cell engagers and various types of therapeutic modalities, CRISPR, etc. But the real challenge is what’s the ideal target and how do we validate that target. We can finally answer that question at scale, particularly using our proteomics approaches. The second application that we work quite a bit in is around drug screening, being able to do very high-throughput drug screening using protein measures, metabolite measures, and lipidomic measures. Oftentimes many drugs are targeting enzymatic processes or are targeting the interaction with specific proteins and being able to assay that interaction as part of a drug screen in very high throughput is powerful. The next phase is around preclinical validation and safety and tox and being able to apply these discovery approaches to predict which drugs may have potential early toxicity so they can be removed from the pipeline and we can focus on those drugs that are ultimately going to be safe. Then as we enter into the clinic, it’s absolutely essential for most drug development pipelines that we are able to understand the stratification of patients, which patients are likely to benefit from a particular therapeutic, to be able to enroll those patients and then to be able to follow those patients through clinical trials to ensure the drugs are reaching their target and doing what they should be doing. If we can do that successfully across the entire drug development pipeline, you can very quickly imagine how much faster that pipeline can go and how much more successful it can be.

Ross Katz: These are basically the entire spectrum of the drug development pipeline. But I think the drug screening and the preclinical validation seem relatively straightforward. If you wouldn’t mind, I’d like to spend a little bit of time on the ideal targets. How can you use the data that you have available and the methods that you apply to identify what is an ideal target for a given disease?

Mo Jain: Let’s take an actual example in oncology, particularly tumor types, and we’re trying to identify what’s the ideal target for an ADC or a T-cell engager or radio-pharmaceutical-based therapeutic. This is based upon the idea that you need a target that’s expressed in the cancer cell, but not expressed in normal tissues. We can leverage that to then introduce a chemotherapy or a radio-pharmaceutical specifically into the cancer cell. This is based upon the idea of what’s fundamentally different about a cancer cell than a normal cell. Traditionally, we’ve used surrogates such as genomic mutational analysis and RNA sequencing data to try to understand those metrics. We recognize that, particularly in oncology, there’s a great difference between the blueprint and what the actual final product looks like here. Being able to measure the proteome in cancer cells, compare it to normal tissue. Pancreatic tumors versus normal pancreatic tissue. And then also look across a spectrum of other normal tissues within the human body to say, where is this protein not expressed as a way of being able to understand the safety profile of a potential target is absolutely critical. Now we’re finally for the first time able to take large numbers of tumors, to map the thousands of proteins that are present in those tumors, to understand how they relate to normal adjacent tissue and how they also relate to non-tumor tissue. We can tell what’s the expression of that protein in the brain, in the heart, in the liver, in the kidney, where may potential toxicities emerge, becomes a facile way now of finally being able to identify targets at scale for cancer chemotherapeutics. The challenge traditionally has always been, if we have a target, how do we drug it? In many ways, biopharma’s done an enormous job over the last half decade or so in greatly expanding our target spectrum through ideas like targeted protein degraders and ADCs, etc. We can now target, or we, meaning the biopharma industry, can now target thousands of more proteins than we ever could before. Just being able to say, from the entire proteome, this is the one or two or five best targets is where we can help quite a bit in target ID and validation. At the same time, at Sapient, we have very large internal databases where we’ve gone out and collected large tumor samples, collected blood samples from patients, etc., in which we have clinical information, and that allows us to leverage those data assets to say, yes, you have identified this target and we know in other tumor types this is also a great target or this is not expressed in other normal human tissues, etc.

Ross Katz: That’s really interesting. What it sounds like is, coming out of the mass spec-based processes that you have, you’ve got high sensitivity and also high specificity in these. You’re able to capture a broad and very clear picture of what’s going on. That just gives you the opportunity to generate a lot more hypotheses, by comparing normal with disease cell tissue, you can ask the question: what are the differences and which of these differences might actually be connected to the mechanisms that we’re trying to measure here? Is that the kind of question that you’re asking at Sapient or are you providing that data back to biopharma and enabling them to ask those kinds of questions?

Mo Jain: It’s the latter. We operate as a service organization. We are not doing the hard work of drug development but rather we’re in service of the biopharma industry. They will come to us typically with these questions and we can leverage our technologies and our data assets and our biocomputational approaches to answer that question for them. You can imagine that if you can de-risk and validate a target as quickly as possible up front, that just accelerates the entirety of the pipeline downstream from that one point.

Ross Katz: For sure. Should I think about the reason they come to you as what you were talking about earlier in terms of the high-throughput but high-resolution mass spec technology that you’ve developed? Or what are the reasons why they come to Sapient rather than doing this kind of multi-omics-based discovery in-house?

Mo Jain: It’s a great question. Many of the pharmaceutical organizations with whom we work have very large internal mass spec operations. This is not particularly isolated to single external entities. The difference is the way in which we can do this, the scale at which we can operate, our ability to handle that very complex data and from these thousands of measures say this is the one best target or these are the two best targets, and then ultimately to be able to validate this quite quickly using our internal data assets becomes the entire process that allows them to de-risk or identify the ideal target for them. We can do this faster than most organizations can do this simply because this is what we focus on. This is a very different externalization that is typically done by biopharma where when they work with most CROs, typically CRO assays are lower complexity assays that are being done at very large scale and so there’s some efficiency that’s being achieved there. This is the other end of the coin where we’ve specialized in very complex assays that have a very important role in drug development. That’s frankly why organizations come to us. We just have a larger group that can do this better and has much more expertise as a whole.

Ross Katz: When do biopharma companies decide to come to Sapient? Is it that they already have a relationship with you and they just know that you’re going to speed them along that much faster? Or what are the kinds of problems that they’re encountering that lead them to the need for the data that you can generate?

Mo Jain: It’s all of the above. Because we can slot in anywhere within that drug development pipeline, there’s individuals that come to us at very late stage with late assets that are in Phase 3 that perhaps have failed and have not achieved a primary endpoint and they believe there’s a subgroup that would benefit from their actual therapeutic and they’re engaging us to help identify from their Phase 3 clinical trial a marker that denotes those individuals who did benefit from the drug. There’s individuals who come to us very early on in the drug development pipeline and say, help us identify what is the ideal target for our targeted protein degrader system or for our ADC or T-cell engager system. And then there’s everything in between. To your point, when we typically work with a pharma organization, this tends to be quite sticky because there’s so many potential applications. There’s so much need in science to be able to measure proteins, to be able to measure metabolites and lipids. Being able to do that just means there’s many applications and typically we start with a single project and that typically’s then expanded to many projects over time.

Ross Katz: That makes a lot of sense. If I’m understanding correctly in terms of the services that you bring to bear, you have the measurement capabilities, the high-throughput mass spectrometry technology. Then you have computational frameworks that you can put on top of that to understand what’s happening. And then there’s proprietary databases that you’ve developed that serve as points of comparison to enable better analysis of the data that’s coming off of them. Is that a reasonable overview of the services that you provide, or how would you talk about them?

Mo Jain: I think that’s exactly right, Ross. Those three pillars, the first is the technologies that allow you to make measurements at scale with great quality and robustness. The second is being able to handle that data with a biocomputational team that can take very complex multi-dimensional data and be able to identify the key markers, the key targets, the key understandings to bring biological insight from data. The third is underlying data assets which allow us to amplify and accelerate the underlying discovery process. Those three really map onto the three pillars of discovery here at Sapient. In the same way we built our mass spectrometry group, in simultaneous nature we built our biocomputational group and our data assets to be able to fulfill that entire spectrum of discovery for our clients.

Ross Katz: That’s really interesting. We’ve spent a decent amount of time on the data that you generate through the mass spectrometry. Can we spend a little bit of time on the computational framework and then on the proprietary databases? What are the computational frameworks that you’re applying to the data that’s coming off of the machines?

Mo Jain: My bias when it comes down to data handling is there’s no one size fits all. Many times folks who don’t work in this space believe there’s a magic piece of software with a big red button and you press it and the answers just spit out the other end. For many of your listeners, they know this is not the case. Being able to do very high-quality data analysis boils down to a couple critical points. One, you have to have very high-quality data that’s entering into your analysis. If you’ve got bad data, you’re doing statistical gymnastics and rarely what comes out on the other end is valuable. Being able to input very high-quality data is key. Being able to understand the actual biological question is oftentimes key. At least in my perspective, it’s not the case that the math here is the limiting factor or the computational framework, and now with distributed computing, you’ve completely eliminated that bottleneck. Rather, it’s being able to truly understand the clinical or biological question and frame the analysis in a way in which you’re able to answer that question. When you talk about the underlying analysis platforms and approaches, it’s very customized for every client based upon their specific question. That means you have to be able to do everything from very reductionist and Bayesian statistics, regression analysis, all the way through very complex AI categorization and classifiers and be able to be facile across that spectrum, understand the appropriate framework for applying different statistical tools to a particular question in order to be able to answer it, and then be able to execute on that with great robustness. That’s the way we think about it. There is no singular this is how we approach the answering of the question, it comes down to what is the question and how can we use the data that we have, be able to answer it in the most effective way possible.

Ross Katz: It sounds like it’s almost like an internal data consulting services team that’s doing bioinformatics-type work, computational chemistry-type work. Can you just talk a little bit about how that team, if you have an example of a question, when a question comes in, how does that team interact with the client or with the question that they’re working on?

Mo Jain: You’re absolutely right, Ross, in that this goes way beyond just having an isolated team that works in a siloed system that just goes and does an analysis and then returns a result. The way our team works with our clients, we say we’re simply an extension of our client’s team. We’re on the bench over. That bench may be a thousand miles to the left in San Diego, but we are simply an extension of their team. Our biocomputational team will meet with the biologist and with our client even before samples are received here at Sapient and we frame out the key biological questions, we’ll frame out the statistical analysis plan so everyone is aligned on the question and how we’re going to approach it even before the first ounce of data is generated. That’s critical. Without an understanding of the question, without an understanding of the clinical context, it’s virtually impossible to be able to come up with a meaningful answer. This process starts very early on in our engagement with our clients and then as that data emerges, being able to do very strict quality analysis of that data and quality control to ensure the data that’s entering into the analysis pipeline is of the highest level and then being able to execute on that statistical analysis that we framed out with our client and ultimately answer their question.

Ross Katz: That makes a lot of sense. Can we talk about the proprietary databases? You’ve got this arm of your business that’s focused on the problems of your customers, of your clients, and they’re coming to you with questions that require you to generate specific types of data. I’m just interested in how you decide about the development of that proprietary database, what you want to be inside it, and how it enables the capabilities that you bring to bear across Sapient.

Mo Jain: Absolutely, Ross. As you’re well aware and as your listeners are well aware, data is enormously powerful and it can accelerate discovery in so many different ways and we’ve learned this and we continue to learn this lesson. It was very clear early on in Sapient’s inception that we had very fast technologies that could generate data very quickly. Oftentimes the limiting factor now was not the speed at which we could generate the data but rather the scale, meaning the number of samples that we had. You can imagine if you’re a pharma client who has a Phase 1 clinical trial, you may have a hundred samples. A hundred samples is a great number from which you can begin to start a hypothesis, but it’s really hard to have a definitive answer regarding how is this drug going to work in the population across several hundred thousand people when I’m starting with only a hundred individuals. We became very interested in this idea of how we could build data assets that would serve to amplify discovery for our clients. From that, we went through that process thinking through particular therapeutic areas and disease and drug modalities, understanding what data would help our clients best, and then went out into the world and collected those biological specimens, analyzed them using our mass spectrometry systems, homogenized all the clinical information as well as the genomics information, microbiome, clinical outcomes, diet, lifestyle, all these factors that we know have massive implications for disease states as well as drug response, and then homogenized all that along with our mass spectrometry data, our proteomics, metabolomics, and lipidomics in a singular relational database that is searchable internally. That’s the way we’ve built this data asset. It includes everything from hundreds of thousands of plasma samples from individuals in which we have longitudinal and serial outcomes data over many years and so we can follow people using these dynamic markers. It includes thousands of tumor samples that have come from a diverse array of tumors or IBD samples or immune samples. It’s very much based around the areas in which our clients work in and it’s not about just having a certain amount of data, I’ve got more petabytes of data than you have, but rather having the right data in a way that allows our clients to answer their questions.

Ross Katz: That’s really interesting. You mentioned quality control just a little bit ago when you were talking about your computational team. Can you talk a little bit about what are the consistent quality control approaches that you can apply across the studies that you’re doing and how do you think about quality control within the context of a given engagement with a client?

Mo Jain: This goes back to one of my personal biases, Ross, and that is the key to good discovery is not necessarily having the best mathematicians, it’s having the best data for those best mathematicians to work with. This again is a lesson that sometimes is forgotten and we focus on size as opposed to quality, and both of those metrics are important but we oftentimes forget how important the quality of data is. The literature is ripe going back for many years with large data sets that may be of questionable quality, and you can see they’re of questionable discovery potential. We’ve spent quite a bit of time and expense in optimizing not only the underlying technologies to be robust and to be stable over time, but also building in the internal quality metrics that allow us to follow this in real time. To give you an idea, whenever we’re running a biological specimen on our mass spectrometry systems, there’s over 500 parameters we’re monitoring in real time that tell us about the quality of the entire data framework. This includes everything from the quality of the sample itself, and we can monitor for how long a sample’s been left out, if there’s been hemolysis, if there’s lipemia in the sample, if there’s any additive, is there something that’s inadvertently been added to the sample either through the acquisition from the plastic or something that the individual was taking that may be interfering with measurement as a whole. It allows us to monitor for the actual instrumentation itself, the mass spectrometers, the chromatography systems. Everything we do at Sapient is fully robotic from our freezers to our sample excisioning, to our processing, to our pipetting and capping of tubes and extraction of proteins and metabolites and lipids all the way up through our mass spectrometers and our data processing is fully automated. We have metrics at each of these stages that we can follow that allow us to tell if there’s ever been a problem and if there’s been any deviation from what we’d consider acceptable. This is a really important aspect of being able to generate data very quickly is being able to follow it at every single stage in real time and ensure the data’s the highest quality.

Ross Katz: I think some of the approaches that you’re describing can be said to be true of any data-generating process, just having that monitoring system in place to make that possible. Are there any upcoming innovations in biomarker discovery or in mass spectrometry that you’re really proud of or excited about?

Mo Jain: Absolutely. Sapient through our history started working in metabolomics and lipidomics initially and then advanced and evolved into the proteomics space. This is owing to our customers who wanted to be able to make all of these measurements in singular samples. You can imagine day over day, week over week, month over month, these continue to get better. Our ability to measure more things faster with greater quality continues to grow by the hour. Innovation is a continuous cycle as opposed to a step gradient change. That being said, we are very excited about where this field is going and how quickly it’s evolving. It’s literally changing month over month, which is phenomenal. For instance, Ross, what we can measure now in a plasma sample with regards to proteins in circulation is two to three to four X where we were just 18 months ago. Those are pretty big differences in the type of proteins that can be interrogated, the amount of post-translational modifications and proteoforms, and our ability to do this at scale very quickly with great quality. That continues to evolve. There are some upcoming innovations in our metabolomics and lipidomics sector. We’ve synthesized very large commercial libraries of standards, thousands upon thousands now approaching 15,000 that we’ve been able to analyze in our systems that allow us to identify many more molecules now from a biosample than we’ve ever been able to. We have really innovated in our ability to go into what we call the dark proteome or the dark metabolome. Those molecules that don’t map onto canonical genomics and for which we don’t know their identity and be able to dive into that end of the pool and identify those, not only measure them accurately and be able to associate them with disease phenotypes or drug response, but also then on the back end identify those molecules structurally. We continue to advance in the tip of the spear here for mass spectrometry, pushing the technologies in what can be measured, mass spec imaging is an interesting space, doing high-throughput proteomics-based drug screening is extremely interesting, particularly for targeted protein degraders, being able to classify, not only generate the data, but classify targets using AI according to their tractability for engagement as well as for safety is super interesting and an area that we’ve innovated in. There’s quite a bit, once you can start from high-quality, large-scale data, you can imagine it just quickly grows from there.

Ross Katz: There’s so many opportunities, which just leads to the kid-in-the-candy-store problem that you were describing earlier. But some of the parameters that you discussed earlier in terms of how you prioritize makes sense. I assume that those are brought to bear there as well in terms of which capabilities you’re developing on the data-generating side and on the mass spectrometry side. You mentioned AI and machine learning. Can you talk a little bit about how it’s being used today, how you expect it to evolve over time at Sapient?

Mo Jain: You could imagine there’s many flavors to a very simple term called AI/machine learning, but there’s many different applications today and through our evolution here. Given the type of data files that we generate which are mass spectral data files, being able to handle those files, being able to extract the information from them and fundamentally these are image files. They’re three-dimensional images. There’s a whole AI machine learning framework that allows us to interrogate those data files to be able to extract what we call the spectral peaks from this, to be able to QC each of these quality peaks, to be able to remove the noise from the underlying data, and be able to do this across thousands of files simultaneously using distributed computing is absolutely critical. That’s a whole AI machine learning-based framework that’s just around data handling and data extraction. There’s a whole component around AI around quality assessment of data, being able to look for underlying drift and biases in data. Then on the back end in the application, particularly with answering biological questions, being able to use AI and machine learning to be able to classify very complex disease states. What we call Alzheimer’s disease or diabetes that really represent a spectrum of many different disorders that have different mechanisms at play, and being able to subclassify patients into those actual categories in a way in which we can identify the subpopulation that’s most likely to benefit from a particular therapeutic is absolutely critical. There’s much more advanced AI patterning that’s coming online here and will continue to evolve as an industry over the next several years, particularly with regards to diagnostics. This goes back to the earlier comment that instead of measuring 20 molecules in a blood sample, we can measure over 20,000 now. How do we leverage that information? This is where AI becomes very interesting going forward to be able to understand how do we look at singular molecules or singular biomarkers, how do we look at patterns of biomarkers to be able to identify patients who are at risk for particular diseases or who are going to respond to a drug. This is changing as you well know, year over year now.

Ross Katz: That makes a lot of sense. As we head toward the end, I’m interested in, where do you see the market going in terms of demand for Sapient’s services?

Mo Jain: There’s a number of changes that are occurring in real time here that you’re well aware of in pharma services, and the biggest driver for this is the need to be able to develop drugs faster, cheaper, that are more effective. Pharma as a whole is one of the few industries that tolerates a 90% failure rate. It’s fascinating to think about how that sets up economically and that has huge implications for the economics of drug development as well as just the time it takes to develop drugs. Given that demand, it’s clear that data is going to be the key to solving this problem and I firmly believe that it’s not going to only come from genomics data, but it’s going to be from the genomics data layered together with orthogonal types of data information streams like proteomics and metabolomics and lipidomics. We see the demand for this multi-omics data generation as well as analysis of data and being able to leverage external data assets for discovery to only grow over time. Human disease is not going away, we know that, and I wish that was the case but that’s not the case for us currently. We have to be able as a community to develop these drugs more effectively, faster, and they have to be done cheaper and that’s only going to come from leveraging these types of services.

Ross Katz: Awesome. As we bring this to a close, where can people go to learn more about you and Sapient?

Mo Jain: Absolutely. Ross, we have an online presence. Our website is www.sapient.bio, b-i-o. That’s where our website is. You’ll find a full menu of services and approaches, contact information. We’re also available on LinkedIn and on Twitter and on social media. We’re always excited to speak to your listeners and there’s many applications for the type of data we generate and we’d love to bring this type of data to the world for drug development.

Ross Katz: Well, Mo, it’s been fantastic talking with you. Really appreciate you joining the podcast and look forward to connecting down the line.

Mo Jain: Thank you so much, Ross. Really appreciate your time today.

Jason: And that’s it for this episode of Data in Biotech. If you enjoyed the episode, please subscribe, rate, or leave a review in your podcast platform of choice. See you next time.

Frequently Asked
Questions

Genomic sequencing has explained less of human disease than we hoped. What data should biopharma sponsors generate to close the gap?
Population attributable risk from the genome alone sits at 14 to 16 percent. The remaining 84 to 86 percent is encoded in dynamic markers: the metabolites, lipids, and proteins in blood that respond to diet, environment, microbiome, and cell-to-cell signaling. Mass spectrometry now measures more than 20,000 of these molecules per plasma sample, up from roughly 20 on a standard clinical panel. Biopharma teams building target ID, patient stratification, or safety programs should pair genomics with proteomics, metabolomics, and lipidomics rather than treating sequence data as sufficient on its own.
How do we keep data quality high when scaling mass spectrometry across hundreds of thousands of samples?
Sapient monitors more than 500 parameters in real time on every mass spectrometry run: sample integrity (hemolysis, lipemia, contaminants), chromatography performance, instrument drift, and extraction steps. The full pipeline is robotic end-to-end, from freezer retrieval through pipetting, capping, extraction of proteins, metabolites, and lipids, and mass spec acquisition. Data quality is the limiting factor in discovery, not mathematical sophistication. Bad input produces statistical gymnastics rather than biology. CorrDyn helps life sciences teams design the QC monitoring, automated alerts, and pipeline observability that multi-omics workflows require.
How should we structure multi-omics data so it can answer biological questions across studies?
Sapient stores hundreds of thousands of plasma samples, thousands of tumor samples, and longitudinal clinical outcomes in a single relational database, with mass spectrometry measurements homogenized alongside genomics, microbiome, diet, and lifestyle variables. The point is not petabyte counts. What matters is having the right samples, with matched clinical context, queryable across modalities. For biopharma sponsors running a Phase 1 study with 100 participants, a relational reference cohort lets you contextualize findings against larger populations before committing to Phase 3 design.

Need a data partner for life sciences?

CorrDyn helps biotech and pharma companies build the data infrastructure that accelerates research and operations.

Book an intro call