Skip to content
Parul Bordia Doshi — AI and Single-Cell Multiomics with Parul Bordia Doshi
Data in BiotechEpisode 28

AI and Single-Cell Multiomics with Parul Bordia Doshi

Parul Bordia Doshi of Cellarity explores AI-driven drug discovery using single-cell multiomics and building multidisciplinary data teams.

39:55Full transcript below
PB

Parul Bordia Doshi

Chief Data Officer at Cellarity

Overview

Traditional drug discovery fails nearly 90% of the time because it targets single molecular pathways, overlooking the complex biological interactions that drive disease. This oversimplification leads to significant R&D costs and low success rates.

Host Ross Katz speaks with Parul Bordia Doshi, Chief Data Officer at Cellarity, who draws on 17 years of experience at Takeda leading data and technology initiatives from discovery through commercial launch. She explains how Cellarity tackles this by leveraging AI and single-cell multiomics to design medicines that modulate entire cellular systems, an approach that promises to reveal new therapeutic opportunities and improve drug development efficacy. Her insights reveal how a precise data strategy and solid infrastructure directly enable scientific breakthroughs in complex biotech environments.

This episode explores Cellarity’s cell-centric platform, detailing its data lifecycle, the critical interplay between AI models and human scientific expertise, and the architectural components supporting their novel drug discovery process. Parul also discusses strategies for managing vast, noisy biological datasets and the practical impact of generative AI in advancing drug design.

Key Takeaways

AI enables holistic drug design, moving beyond single-target simplification.

Cellarity’s platform uses AI and multiomics to understand and modify entire cellular states, rather than isolated molecular targets. This thorough, system-level view addresses the inherent complexity of human diseases, leading to a 20x increase in hit rates compared to conventional screening.

Data strategy must align directly with scientific objectives to drive innovation.

For Cellarity, the Chief Data Officer’s role is to design data infrastructure and management frameworks that directly facilitate novel biological insights and enable AI model training. This ensures data capabilities advance core scientific goals, enhancing the efficiency of discovery processes and pipeline development.

Human scientific expertise remains essential for validating AI-driven discovery.

While AI models generate hypotheses and predict compounds, human scientists are indispensable for designing specialized experimental assays, interpreting phenotypic readouts, and validating model predictions. Their domain knowledge provides the critical feedback loop needed to refine AI algorithms and ensure clinical relevance.

Reliable data pipelines and LLMs improve data quality for AI models.

Public multiomics datasets are often noisy and lack standardized annotations, hindering model performance. Cellarity addresses this through automated ingestion pipelines, rigorous data validation, and by exploring LLMs to extract reliable contextual information, ensuring high-quality input for their advanced AI models.

Related: CorrDyn helps clients build AI strategies that deliver tangible business outcomes, particularly in biotech and life sciences. We specialize in data engineering and machine learning solutions that transform complex data into scientific breakthroughs.

Full Transcript

Jason: Hi everyone, this is Jason, producer of Data and Biotech. Before we get started, I wanted to let you know about our latest white paper. It’s a comprehensive guide to implementing machine learning models in biotech manufacturing. It’s a complete overview of all the potential problems of ML adoption and, more importantly, how to solve them. To download it, simply visit connect.corrdyn.com/biotech-ml. We’ve also dropped the link in the show notes of this episode. Let’s get into it. Welcome to Data in Biotech, a podcast from CorrDyn where we explore how companies leverage data to drive innovation in life sciences. Every two weeks, we sit down with an expert from the world of biotechnology to understand how they’re using data science to solve technical challenges, streamline operations, and further innovation in their business. In this episode, we’re excited to be joined by Parul Bordia Doshi, Chief Data Officer of Cellarity, a company that’s challenging the traditional approaches to drug discovery with their revolutionary cell-centric platform. Parul dives into how Cellarity is leveraging AI and single-cell multi-omics to design medicines that target the entire cellular system rather than individual molecular targets. She discusses the inefficiencies of conventional drug discovery methods and how Cellarity’s holistic approach, powered by cutting-edge technology, is pioneering a new frontier in drug design. Parul also shares the challenges of managing vast amounts of data and the potential of generative AI in the future of drug discovery. Here we go.

Ross Katz: Parul Bordia Doshi, welcome to the Data in Biotech podcast.

Parul Bordia Doshi: Hi Ross, great to be here.

Ross Katz: Awesome. Well, just to kick us off, could you give us a brief introduction to your background and what brought you here today?

Parul Bordia Doshi: Sure. I’m Parul Bordia Doshi, I am the Chief Data Officer for Cellarity. It’s a biotech based out of Somerville, Massachusetts, and we are pioneering a revolutionary cell-centric platform that leverages AI and multi-cell, single-cell multi-omics to design medicines that modify the disease states. I joined Cellarity about three years ago after spending 17 years in Takeda through their acquisition of Millennium Pharmaceuticals. I’ve led technology and digital initiatives that have played a crucial role in launching multiple specialty drugs to the market. My career has been deeply rooted in all things data, right from building, optimizing, managing workflows, and ensuring efficient data management. I’ve worked pretty hard to make sure that we’re using or deriving the data value right from the discovery through the commercial process. With support from great team members, I’ve consistently worked at transforming how we use data to drive innovation, improve operational efficiencies, and ultimately deliver better outcomes to the patients.

Ross Katz: Taking that as the jumping-off point, I’d love to hear about Cellarity, its mission, and how the organization goes about accomplishing that mission.

Parul Bordia Doshi: Absolutely. As I mentioned, the dominant paradigm currently for drug discovery requires reducing complex biology to a single molecular target. The drug hunters then seek to identify compounds that will bind to that single target or protein in the hope that will alter the activity of this disease. Although many successful drugs have been identified using this approach, a major limitation is this prevents translating to the clinic. Most human disease is caused by complex biological pathways comprising many different molecules and cells in the body. When you target only one protein in this complex system, at times it doesn’t work — actually most of the times it doesn’t work. Almost 90% of the drug discovery programs that start Phase 1 never make it to the market, with many more failing even before that. Cellarity was created to challenge this core tenet — what if we could design drugs that modulate the entire state of the cells rather than focusing on a single molecular target? We aim to unlock new therapeutic opportunities by understanding complex cellular dynamics and cellular systems. To accomplish this mission, there are a few tenets that we use. First, we focus on holistic behavior of the cells. We leverage and integrate multi-omics data to obtain detailed insights into the functioning of these cells and interactions of individual cells across the system. We employ AI and advanced machine learning models to analyze these complex biological data, identify patterns, predict how different compounds may interact or intervene with these cellular behavior, and this helps us in identifying potential drug candidates that can restore healthy cell states across the entire disease. We rigorously test our predictions or hits in the lab to reproduce the target omics signature but also induce a clinically relevant phenotype in our in-vitro system. And ultimately we use generative chemistry tools to achieve the ideal ADME properties — absorption, distribution, metabolism, excretion, and toxicity profiles — so that we can minimize the off-target impact and bring more efficacious and safe drugs to the patients.

Ross Katz: Interesting. Can you walk me through how you’re gathering this data, from the very beginning of the process all the way through — what does the life cycle of the data look like and how it’s used throughout the organization?

Parul Bordia Doshi: We start with hypothesis generation, which is deeply rooted in clinical data that we procure from outside. Based upon our hypothesis, we may have a specific need for generating data in-house that is context-specific, so we would generate that data. This is all fed into our models, and using Cellarity Maps — which is our homegrown visualization tool — biologists and computational biologists can look at the various disease states. We use this signature that comes out of Cellarity Maps to see how, using our proprietary data set of perturbation data on millions of cells, we can identify a compound or a set of compounds that will have similar behavior based upon what we want to see from a Cellarity Maps output. Ultimately we have a prioritized list of compounds that we then run in the labs to see if it’s giving us the right transcriptomic signature but also the right functional readout, which we then take further using our generative chemistry models to optimize the compound. All through this, we generate data in the lab — we have Illumina as our workhorse — and bring the data in using technologies like Nextflow for single-cell data processing, ultimately feeding it into the models. Did I answer your question?

Ross Katz: I think you did. So on your website you have these beautiful visualizations coming out of Cellarity Maps. Can you help me understand what those visualizations are showing and how that’s used to understand the outputs of the data that you’re generating?

Parul Bordia Doshi: Cellarity Maps is a homegrown solution. It is a no-code visualization tool that helps bridge the biologist to a computational biologist. We use this tool after getting data from our internal data sources as well as external ones — we bring these atlases together so that biologists and computational biologists can see the disease state from one state to another, throughout the progression of the disease. These visualizations are a starting point, or the hypothesis, that we then feed into our intervention library and ultimately our drug design studio to generate compounds that will have an impact based upon that transcriptomic signature we saw in Cellarity Maps.

Ross Katz: For those who are listening, you should definitely go to the Cellarity website and take a look at the visualization — as a data person, it’s a beautiful visualization. I think I have a better understanding now of the data you collect and how it works through the system. As the Chief Data Officer, what is your role in overseeing the operation of the system and thinking about the strategy surrounding it?

Parul Bordia Doshi: My primary responsibility at Cellarity is to make sure that Cellarity’s data strategy and infrastructure are perfectly aligned with its mission to revolutionize drug discovery. This involves designing and implementing a robust data infrastructure and data management framework that facilitates the collection, analysis, and integration of these complex data sets. These data sets by design are noisy, so you need to have a robust framework to validate and standardize the data. This data enables scientists to uncover novel biological insights and empowers our computational and machine learning experts to train models and validate our AI algorithms. By aligning our data capabilities with our scientific objectives, we’re driving innovation, enhancing the efficiency of our discovery processes, and contributing to the advancement of our platform and the drug pipeline. My role also involves shaping and executing a data strategy that fuels the company’s innovative approach to drug discovery — what kind of data would we use, the workhorse being transcriptomic, what else is needed — and fostering collaboration between data teams, scientists, and AI teams so that they’re looking at the data holistically and driving insights from it.

Ross Katz: Interesting. You gave us some insight into the components of a data strategy. Can you give us a sense of some of the current objectives you have in place and how that’s manifesting in terms of the initiatives you’re undertaking at Cellarity?

Parul Bordia Doshi: A key asset for the company is data, and we realize that for our models to get better, we need more data, which is context-specific as well. While there is a lot of public data available, it is noisy, so it does take time to standardize, to validate, and follow the ontologies so that we can use it internally. We are spending a lot of time and energy right now on increasing our data set. We have millions of compounds currently in what we call our intervention library, which is a perturbation data set, and we are growing that across multiple cell lines so that we can look at the disease holistically but also go beyond heme and immune, which is our current focus right now.

Ross Katz: Can you give us some insight into how the data teams are structured at Cellarity in order to implement the data strategy that you just described?

Parul Bordia Doshi: I’m fortunate that I lead a team that has very diverse skill sets. Each member has core technical expertise, but they also bring diverse scientific background, enabling us to support a multilingual organization. I have a relatively flat organization and the team collaborates directly with research scientists and the computational biologists and machine learning experts. They sit in other functional team meetings as well so that they can understand the broader context and contribute beyond their core responsibilities. That is something that’s unique at Cellarity — my team has a pretty diverse background, and going beyond what they were hired for, they’re able to contribute much more broadly.

Ross Katz: You mentioned sitting in meetings with other teams in order to understand how their work connects. Looking at the leadership team at Cellarity, you have leaders across science, drug creation, translational medicine, genomics, computational chemistry, platform, and discovery science. How do you partner or interface with all these different departments and leaders at the organization?

Parul Bordia Doshi: Collaboration is fundamental to what we do at Cellarity. One of my favorite reminders comes from a fortune cookie: don’t think alone. We actually follow this at Cellarity. Our approach is inherently multidisciplinary and multilingual. We sit with team members from different functions by design because we want to drive innovation. Project teams are intentionally cross-functional — they bring together specialists from various fields to tackle complex challenges. Leadership within these project teams is not determined by seniority, but by selecting the best individual for that team. The close collaboration between departments is reinforced by our open office environment. We are in the office three days a week, most of the time five days a week. This helps us have hallway conversations even outside of formal meetings to figure out what the problem is and how to solve it. We also sit in different groups, so we hear conversations — maybe a problem that the comp chem team is dealing with — and the team is able to offer support around that. It’s a great collaborative environment. The leaders you mentioned all share the passion that we have to change the way drugs are discovered, and that shared enthusiasm everyone brings to the work every day.

Ross Katz: I want to return to the end-to-end discovery process and dig under the hood a bit about how drugs are discovered and how the platform enables that. Would you give us an introduction to one of the programs you have underway that we can use as an example of how the platform enables drug discovery?

Parul Bordia Doshi: We utilize generative AI including VAE and transformer-based foundation models. This is all trained on our proprietary data set, which we call the Intervention Library. Additionally, we use contrastive learning techniques to expand our capabilities beyond the data that we have generated. In general, we use cloud compute to do this, and as I mentioned, we use either off-the-shelf software when available or we build our own tools. There are three components of our platform I can walk you through. The first is Cellarity Maps. With our cutting-edge genomic facility and advanced ML algorithms, we have created single-cell atlases that illuminate the novel biology underlying a complex disease — the same approach we used to identify for our Sickle Cell program as well. These foundation models learn across the disease representation of the cell and enable us to uncover the cellular drivers of the disease, leading to the identification of novel pathways to treat these complex diseases. These signatures are highly dimensional omics profiles, and we use this as an input to predict compounds that would modify the disease process. We are not simplifying the biology — we are embracing the complexity of the biology. Looking at the disease signature, we tie back to the intervention library, which is a proprietary database comprising millions of chemical transcriptional profiles. We use contrastive learning algorithms to predict compounds that we think would reverse the disease as identified from Cellarity Maps. This leads to about a 20X increase in hit rate compared to traditional high-throughput screening, which is the bread and butter for most pharma right now. We feed these predictions into our drug design studio, the third element of our platform, which gives us a strong starting point for lead optimization. We then use generative AI and quantum chemistry algorithms to learn the critical elements of the compound and design totally new chemistry optimized for safety and efficacy. Using this platform, we were able to identify an entirely new novel pathway for our sickle cell program to induce HbF, the human fetal hemoglobin. That target has never been identified in any hematology or sickle cell program. We are a small molecule company, and we hope we’ll have cell therapy-like efficacy in a once-daily pill for the patients. Pretty revolutionary for folks who are suffering from sickle cell.

Ross Katz: That’s amazing. Throughout the data gathering phase and as the models are being developed, how are the outputs of the models being consumed by people? And how are the outputs being fed into the experiments you run in order to close the loop and bring data back into the models?

Parul Bordia Doshi: Our platform is an end-to-end discovery platform that’s used across the organization in every step of discovery. The output from Cellarity Maps is the disease state signature that’s an input for our next platform component, the Intervention Library. The output from our Intervention Library is a prioritized, actionable list of compounds that are then used in our drug design studio as an input to further design de-novo compounds. Data is integral to what we do. We have AI models that are fed data, but as we test in the lab, we also bring that data back into our models to make them better as we learn from the wet lab experiments.

Ross Katz: How much human intervention should I think about there being throughout this process? It sounds like there’s a decent amount of potential for automation between these different phases — especially once you have the intervention library, you’re quantifying how likely each of these compounds are to be close to the one you’re looking for and then generating based on that. Where do humans come into the loop?

Parul Bordia Doshi: The list of prioritized compounds comes from the models. But when it comes to actually testing those compounds, that’s where humans are integral. They are currently the ones deciding what the assays will be and what the functional output is that we are looking for. We hope that soon we will have model-driven recommendations for even the assays and these readouts. But right now, it’s all human-driven. Ultimately, even as we get to a point where the models are recommending the functional assays, humans will still be the ones choosing which functional assay to take into the lab.

Ross Katz: And when you say the functional assay, you’re saying, how should we even measure the efficacy or the different elements of the compound that we’re trying to get to?

Parul Bordia Doshi: Yes. When we have the disease signature and a compound we think is going to reverse that, we have to test in the lab whether the compounds predicted from our IL — intervention library — are actually giving the right response. First to the signature, and second to the phenotypic readout that we would want, which is actually reversing the disease. We have specialized assays — these are not usually off-the-shelf assays because we are not reducing the complexity of biology, so our assays are complex as well. These in-vitro models and assays that we test in the lab are currently selected and driven by humans.

Ross Katz: This is almost backwards of the way we think about the roles of machines and humans in a lot of scientific processes. What it sounds like — correct me if I’m wrong — is that the models are capable of doing a lot of the hypothesis generation and getting you all the way to potential intervention candidates that you want to test, but then the selection of how you go about testing for the phenotypic results you want requires the creative thinking and scientific background to really understand what’s happening.

Parul Bordia Doshi: Absolutely. It definitely requires the domain expertise that the scientists bring. We critically evaluate the assumptions underlying the AI platform that we’re building — we’re continuously testing, benchmarking, and validating these model predictions against the experimental data from the in-vitro models and assays in the labs. This is where our scientists who are domain experts are advising and figuring out what assays to use and what are the right functional metrics we should be tracking. And this feedback loop — we bring what we learn in the wet lab back to the dry lab to make sure that we are refining our AI models as we go along.

Ross Katz: Given that you have these models in production — evaluating and predicting new compounds and modeling the cellular states — how do you think about enhancing those models and the processes you’re using to identify potential candidates as you look forward?

Parul Bordia Doshi: We are building foundation models across different cell types to enable us to do that. We are generating a lot more data across the cell types — that’s the key component. With the advent of generative LLMs, we are using some of the available models and then training on our own data, context-specific. For us, context-specific data is important, and the volume of that data is important, as we think about improving our models.

Ross Katz: So the data science team members are retraining the existing model architectures using the new data coming in, and it also sounds like there’s an element of evaluating the additional foundation models and new architectures that are becoming available and bringing your existing data sets in.

Parul Bordia Doshi: Absolutely — to see how our models, along with the models that are available, can be used in conjunction to deliver better results. We are really looking at the architecture in a whole different way.

Ross Katz: Should I think of that as ensemble approaches? New models become available, so there’s an opportunity to use multiple models where you once were using one, or add a third model where you once were using two?

Parul Bordia Doshi: Yes. Or using one model where we were using two as well. It all depends upon what we are trying to get the output and what models are available. We think the data that we have is crucial for us, so as we fine-tune or retrain these models, they are giving us better results that we’re already seeing.

Ross Katz: Interesting. What keeps you up at night as the Chief Data Officer? What are the biggest challenges that you face — the big rocks you feel like you have to tackle over the next year?

Parul Bordia Doshi: In my role I’m also responsible for cybersecurity, and that is one component that does keep me up at night at times. On the data side, we are generating massive amounts of data and have plans to generate even more. Being able to drive maximum value out of that data — whether via models, or by making it available for scientists to do cross-model comparison, cross-data comparison — I don’t know if those keep me up, but they’re what I’m most excited about.

Ross Katz: What do you view as the next steps in the development of Cellarity’s AI platform? Where does it go from here — three years or five years down the line?

Parul Bordia Doshi: Right now we are focusing on heme and immune. As we bring in more data sets, we want to expand the therapeutic areas that we go into. We are building on foundation models — they’ll continue to get better, and we’re already using them. Architecture-wise, being able to use multiple components to drive better results is where we are headed. For us, the key is the data, which is context-specific but across cell lines, so that we can widen the therapeutic areas that we tackle.

Ross Katz: You mentioned you’re on AWS and that you’re using Nextflow. Can you give us some insight into the infrastructure components you’re using — where are the workhorses that you’re running your platform on?

Parul Bordia Doshi: We have advanced single-cell sequencing technologies that are used to generate high-dimensional data. We use tools like 10X Genomics for that. For data storage, we have scalable cloud infrastructure. We use Nextflow for our automated bioinformatics workflow — there’s a Seqera tower that provides the Nextflow. We use Python and R quite a bit, with libraries like Pandas, SciPy, and so on. R is a standard tool for data analysis and statistical computing. For deep learning frameworks, we use PyTorch and TensorFlow for training these deep learning models on single-cell data. The data scientists use Jupyter Notebooks and we use Saturn Cloud, which is a managed data science platform for us. We use containerization with Docker and orchestration with Kubernetes to ensure reproducibility and scalability of analysis pipelines and machine learning models.

Ross Katz: Interesting. On the data science side, what does the process look like for bringing a new model to production, or retraining an existing model and bringing it to production?

Parul Bordia Doshi: We set the goal for what the model would do and define benchmarks for what we think the model should be able to achieve, all driven by data. We use Weights & Biases as well as BentoML to make sure that we have different versions of models as we think about bringing it to production. It’s a close collaboration between the data team and the ML scientists to think about how we would benchmark, what architecture to use, and when is the right time for deployment — and then of course containerize and Dockerize these things.

Ross Katz: That makes a lot of sense. Depending on where the model fits inside of the AI platform, the nature of bringing that model to production might be a little bit different. You mentioned benchmarks. What do benchmarks look like on the generative side?

Parul Bordia Doshi: It depends on what we are trying to improve on. Based on what our current model or current process is able to predict, we tie that back to what the new model should be able to do — whether it’s running a POC with a small capability to see whether we’re improving it or not, or building an MVP, tying it back to benchmarks for whatever readout we want the model to improve: the quality or properties of the compounds. The comp chem team works closely with the medicinal chemistry team to determine what properties we are trying to improve and what would be a meaningful impact. As we build and train the models, we measure that back to make sure the models are improving these things. Once the team decides it’s ready for production, we would use it in a program with a functional readout to ensure it is producing the results that we need.

Ross Katz: Right. You have your past data that you’ve collected, you’ve got your holdout data set — the golden data set that it needs to perform well on — and then when the time comes to roll it out to production, the question is whether the new compounds being generated actually demonstrate the improvement in the functional capabilities that you need.

Parul Bordia Doshi: Most of our ML teams have a software development inclination as well, so they work very closely with my team on the software side. As we think about these models, we think about the right architecture — whether it’s parallelization, the right kind of GPUs, compute, and so on. There’s very close collaboration as we think about moving things into production.

Ross Katz: That makes sense. As you start thinking about broadening the different disease categories and the different cell lines that you’re looking at as a result, what are some of the data considerations on your mind for how you lay the groundwork for those capabilities looking forward?

Parul Bordia Doshi: Our starting point is the clinical data set. Ensuring that we are able to get good clinical samples to work on hypothesis generation would be key for us. We are looking at what public data sets are available to support our internal data generation efforts as well. Heme and immune is what we’re focused on right now because it’s easier to get samples when thinking about sequencing in the lab. As we expand into other therapeutic areas, it’ll be based upon the clinical samples we can get, the cells we can get and onboard internally, and then test the hypothesis as well.

Ross Katz: Do you feel like the quality control processes and lessons learned that you’ve applied to the public data sets for the therapeutic areas you’re targeting today will be applicable or replicable, or is it going to require a lot of thinking through the nature of the new therapeutic area?

Parul Bordia Doshi: For the public data sets, we are exploring using LLMs. They’re noisy at times — the relevant annotations or information about the data set might be in a table or buried at the end of the paper — so we are exploring options to use LLMs to pull information from publicly available data so that we have more trust in what we bring in. We now have an automated process to ingest the external public data. We have standardized the annotations so that as we load data into a central repository, we have standardized it, we validate, and anything that doesn’t meet these standards we kick out for people to look at — making sure we are not bringing in garbage. These pipelines that we have built will support different data sets that we bring in. We hope that the LLMs will also improve and reduce the human intervention that’s needed right now to extract relevant information that’s deep in the paper somewhere.

Ross Katz: When you were saying publicly available data sets, it didn’t occur to me that this is information that needs to be parsed out of academic papers in order to be normalized for use — these are experimental results that have been conducted in publicly available studies. Am I thinking about that right?

Parul Bordia Doshi: These are data sets that are available for public consumption, so we can get them. But the challenge is that they are not always clean. They are from different labs following different best practices — there’s not one methodology. At times they say a certain type of file is available, and when you actually dig in, it’s a totally different file set that’s available. Right now the team spends quite a bit of time even understanding what assays were used — things that are deep in the paper, not something you can simply parse out. They are in the table or buried somewhere. And for cell types, there’s no standardization on cell names and things like that. It takes a lot of effort from our scientists, especially the computational biologists and the tech team. By employing LLMs, we hope that we can reduce that. The cleaner the data, the better reliance we can have on the hypotheses we’re generating.

Ross Katz: The quality of the data is essential for the models you’re producing and for developing the capabilities to develop working drugs further down the pipeline, but the nature of the work upfront is grueling and mind-numbing for people who have a lot of expertise and probably don’t want to spend their time copy-pasting or editing text that’s been extracted from different documents. That’s a hard challenge.

Parul Bordia Doshi: Absolutely. We are hoping that we can ease some of this so that we have better hypotheses.

Ross Katz: As we head toward a close, just a couple of final questions. I heard that Cellarity is partnered with Novo Nordisk. Can you tell us about that collaboration and how you integrate with a large pharma company like that?

Parul Bordia Doshi: Earlier this year we announced the expansion of our partnership with Novo to discover and develop a novel treatment for MASH. This partnership builds on the success of a research collaboration that was already in place. We’ll continue to leverage Cellarity’s platform to create a small-molecule therapy for MASH. We are really excited to be working with experts in metabolism at Novo Nordisk as we leverage our platform and bring in life-changing therapy for MASH patients. More to come on that.

Ross Katz: As you look toward the future of the application of data to the kinds of problems that Cellarity is solving, what do you think the next three to five years holds? What excites you most?

Parul Bordia Doshi: The use of generative AI is exciting. Being able to use models that are available outside and integrate them with what we are doing at Cellarity with our data — we’ll be able to take an immense leap from where we are right now. And to me, seeing our sickle cell program enter the clinic soon — being able to see a program that we’ve worked on really get to the clinic and into the patients — that is the most exciting thing for me.

Ross Katz: Does that change your role as Chief Data Officer once you’re in the clinic and you can start thinking about the potential for clinical trials?

Parul Bordia Doshi: We will add on the clinical data side of things, but we have a good robust pipeline, so there will be more programs to keep us busy on the pre-clinical side as well.

Ross Katz: Where can people go to find out more about you and about Cellarity?

Parul Bordia Doshi: cellarity.com — we have several case studies there that highlight our platform as well as the programs that we are in.

Ross Katz: Well, Parul, it’s been a pleasure speaking with you today. I really appreciate the time and I’ll look forward to connecting down the line.

Parul Bordia Doshi: Thanks a lot. Thanks for the opportunity. Have a good one.

Jason: And that’s it for this episode of Data in Biotech. If you enjoyed the episode, please subscribe, rate, or leave a review in your podcast platform of choice. See you next time.

Frequently Asked
Questions

How does AI specifically improve drug discovery success rates?
Cellarity's AI platform uses single-cell multiomics to model and modify entire disease states, rather than single molecular targets. This holistic approach helps identify more effective compounds, resulting in a 20x higher hit rate compared to traditional high-throughput screening methods.
What is the balance between AI automation and human involvement in Cellarity's drug discovery process?
AI models generate hypotheses and predict potential compounds. Human scientists then critically design and conduct complex experimental assays to validate these predictions, interpret phenotypic results, and provide the domain expertise to refine the AI models and ensure clinical relevance.
What infrastructure components support Cellarity's AI drug discovery platform?
Cellarity leverages AWS cloud infrastructure, Nextflow for automated bioinformatics workflows, Python and R with scientific libraries, and deep learning frameworks like PyTorch and TensorFlow. They utilize Jupyter Notebooks via Saturn Cloud and containerization with Docker/Kubernetes for reproducibility and scalability.

Need a data partner for life sciences?

CorrDyn helps biotech and pharma companies build the data infrastructure that accelerates research and operations.

Book an intro call