Listen on
Overview
Host Ross Katz speaks with Pradeep Ravindra, Data Scientist at Lyell Immunopharma. CAR-T cell therapy promises a new front in the fight against cancer, yet many biotech manufacturers still rely on paper batch records for commercial products. This operational gap stalls timely decision-making, impacts product quality, and jeopardizes patient outcomes—directly affecting clinical trial success rates and overall business viability. Data leaders in this sector face the challenge of extracting real-time insights from complex manufacturing processes while managing stringent regulatory demands.
In this episode, Pradeep details how his team confronts these issues. He brings a unique, cross-domain perspective, having applied data science techniques from military intelligence to personalized medicine. He explains how data analytics is critical for balancing the three pillars of CAR-T manufacturing: achieving target cell dosage, ensuring product quality and non-contamination, and enabling on-time delivery.
The conversation explores strategic approaches to data governance and the foundational role of a reliable semantic layer in enabling advanced analytics. Pradeep shares how to accelerate insights by deploying non-GMP analytical systems alongside validated ones, encourages cross-functional data collaboration, and envisions the future of digital twins and simulations for rare disease therapies. This episode offers practical guidance for data professionals seeking to elevate their impact in high-stakes biotech environments.
Key Takeaways
A well-defined semantic layer is essential for meaningful analytics.
Before any advanced AI or machine learning can truly deliver value, an organization must establish a clear semantic layer. This layer translates raw data into meaningful business dimensions, acting as a crucial bridge between technical datasets and human understanding of business outcomes. Without it, data remains siloed and difficult to interpret effectively across functions.
Build rapid analytical tools for insight before formal validation.
Biotech firms can accelerate crucial insights by developing non-GMP analytics platforms for informational purposes. This approach allows teams to explore data, identify correlations, and advise the business quickly. Formal validation for GMP-critical systems can follow, ensuring innovation isn’t stifled by initial compliance hurdles in early clinical phases.
Proactive alerts deliver more value than static dashboards.
The most effective data products proactively deliver critical information, notifying stakeholders when specific data points exceed thresholds or reveal significant trends. Rather than expecting users to constantly monitor dashboards, intelligent alerts ensure timely action and transform data from passive observation into a driver of operational decisions.
Data professionals require deep business and scientific understanding.
To move beyond an “IT support” role, data analytics professionals must actively acquire domain expertise in the specific science and business operations they support. This knowledge enables them to anticipate problems, propose strategic solutions, and truly sit at the decision-making table, translating complex data into corporate goals.
Related: CorrDyn specializes in data engineering for biotech and life sciences companies. We help organizations develop strong technology strategies to ensure data quality for critical operations like CAR-T manufacturing. Read more about how we help biotech manufacturers gain data value.
Full Transcript
Jason: Welcome to Data in Biotech, a podcast from CorrDyn where we explore how companies leverage data to drive innovation in life sciences. Every two weeks, we sit down with an expert from the world of biotechnology to understand how they’re using data science to solve technical challenges, streamline operations, and further innovation in their business. This week, we’re excited to be joined by Pradeep Ravindra, Associate Director of Data Analytics, Manufacturing from Lyell Immunopharma, who shares his story on transitioning from military intelligence to life science analytics, the role of data analytics in the manufacturing and delivery of CAR-T, TCR, and TIL therapies, and how to foster a culture of data-driven decision-making and collaboration within a biotech organization. Here we go.
Ross Katz: In 60 seconds, can you just tell me about your career today? And I won’t hold you to exactly 60 seconds.
Pradeep Ravindra: Sure. Yeah, I’ve been in data my entire career. It started off in the military. I enlisted in US Army intelligence, did a tour in Iraq, where I was doing geospatial visualizations and density plots to assess enemy activity and try to forecast where we think enemy activity would occur the most to help preserve life and our convoys, which were carrying supplies as we were closing down Iraq at the time. Then after I left the military, I was studying pre-med for a while and didn’t go well. I was dealing with some transitional issues from coming out of war and being a college student, it wasn’t the smoothest transition and eventually I went back to computers because I’ve always loved computers and got very lucky as I was graduating with my bachelor’s in a field called informatics, I was recruited into a company called Celgene that was working on CAR-T cell therapy in clinical trials. That’s how I ended up in oncology and specifically personalized medicine and I’ve been working there ever since and I love it.
Ross Katz: Very interesting. The military background really jumps out. It sounds like you were doing some pretty data-intensive analytics in the military, but I’m wondering, what was the transition like from military intelligence and analytics to life sciences analytics? Do you see a lot of overlap between them? Do you find yourself leveraging that experience?
Pradeep Ravindra: Absolutely. Obviously, the problems are different that we’re solving for, but the solutions that you’d use to solve those problems are very similar. For example, in the world of flow cytometry, there’s this technique you apply called gating. I’m not an expert at flow cytometry, but I do know a clustering when I see it. Clustering is something that you use quite often to try to find patterns in some sort of spatial dimensions. So, while the problems that I’m solving are obviously quite different, I’d say the techniques are quite translational.
Ross Katz: That’s one of the amazing things about working in data is that these methodologies that are applied to vastly different domains can have this incredible effectiveness in unexpected places. Can you talk to me about the role of data analytics in the manufacturing and delivery of CAR-T?
Pradeep Ravindra: Before you even start manufacturing for patients, you need to file an IND for CAR-T cell therapy. The FDA has to approve you to go into human trials. Before that, you mess around with mice and in vitro models. When you transition into the patient domain, where you’re actually treating patients who are courageously volunteering themselves to advance a frontier of science, which is essentially what they’re doing when they’re signing up for a clinical trial, that’s my understanding when I talk to clinicians about this, it’s a game changer because each data point that you’re looking at in the manufacturing space is a data point correlated to the performance of a patient’s run. That intimacy that you have to have with the data completely changes. It’s not just some data you slap into R and you produce some correlation plot. This is a human, this is the data that their cells are producing. How can we make this product as good as it can be, the best it can be to give them the best chance possible? When you think about it like that, it sounds like a very complicated problem because there are so many data points, there’s so many different kinds of data that you can ingest. Let’s take a very simple example. One of the foundations of CAR-T cell therapy is that you receive a patient’s apheresis, a collection of their white blood cells. You go through a selection phase where you try to only extract the cells that you’re looking for, specifically T cells, CD4, CD8 T cells. Your goal is to transduce them with some sort of gene delivery, which in most cases is vector for CAR-T cell therapy. Once they’re transduced, you grow them. You’re taking these cells that have now been genetically modified and you’re growing them in a bioreactor or whatever means that your company is using. You have to also understand that each day that the cells spend outside the human body weakens them. It’s this inverse relationship between the quality of the cell versus the population of the cell because we need to grow those cells. We need to hit some sort of target dosage. The dosage here we’re not talking about Robitussin, which is measured in milliliters. We’re measuring actual living products, so this is a cell count. In this case, the cells have to be grown to a certain target to meet the protocol for your IND filing. You said to the FDA, this is the dosage that I’m going to give my cohort, so you have to meet that dosage. But on the flip side, when a patient’s cells aren’t growing so well in the bioreactor, when they’re not responding well, when they’re growing slower, the tradeoff that you have to constantly battle with is that each day that they’re outside of the human body is a risk towards the quality of the cells. Just because you’ve met the target dosage, just because you’ve grown the cells to the cell count that you’re looking for, doesn’t necessarily mean that you’ve succeeded. You could have spent much longer than necessary growing those cells when there were other factors or other variables at play that we could have examined to grow those cells faster. It’s a multivariate problem and we’re constantly looking at ways of establishing at the very least what are the list of parameters that you want to collect on your process. Those parameters are driven by your IND. Your IND says, these are the critical parameters that we’re going to monitor and test for. QC’s going to have these assays, they’re going to look for mycoplasma or any sort of contamination to make sure it’s safe. Those are fundamentals, those are obvious. But then the not-so-obvious ones are measuring cells that are alive versus dead and the way that’s done is through an assay for a family of proteins called C-PARP, which is supposed to be relevant when there is evidence of cell death, apoptosis. Apoptosis is programmed cell death. It’s part of the natural course of the lifecycle of a cell. It’s a live product. What you want out of your live product, one, you want to hit that target dosage. Two, you want your cells to be high quality, which also means you don’t want a large amount of them to have this C-PARP, this family of proteins, to have any indication of this because that means that they’re dying. You want them to be alive and not dying. So no C-PARP and they’re alive. Number three is delivering this therapy on time while balancing those other needs. Because the process is variable, the patient material we’re getting is variable. We can’t just source these cells from one continuous pool. That would make process controls a little easier. Variable starting material means variable process outcome. When you stack that variability against each other, you end up with a very complicated problem and if you’re thinking about this, you might be thinking, are there possibilities of deep learning, machine learning, artificial intelligence that we can delve into? Yes and no. The problem with it is that now we’re talking about volume of data. The volume of data in a phase one isn’t very high. But if you’re talking about a company that’s gone commercial and has all that data they’ve been accruing, unfortunately a lot of that data sits in these paper batch records because not a lot of companies invest in electronic batch records or electronic systems and automation upfront, which makes sense. When you’re in a phase one, phase two, you’re bleeding capital, so you have to be very smart about how you strategically deploy it so you can survive through your key clinical milestones and raise more money. It’s looked at from a fiscal perspective, but long story short, there’s a lot of elements whether it’s supply chain, quality control, manufacturing, on the research side or on the clinical side. If a patient performed really well and you want to replicate that, then you’ll go back and look at the CMC data and say, I want to see all their CMC data and try to replicate this in other patients. What variables here are noticeably different than other patients who may not have performed as well and how can we reproduce that.
Ross Katz: Very illuminating. What I’m hearing is there’s three major elements that you’re trying to impact. There’s dosage or volume, there’s the quality and the non-contamination of the sample, and then there’s on-time delivery. Those are three of the high-level KPIs that you’re looking at. Are there any big initiatives that you’ve undertaken for one of those that you felt were particularly impactful? If so, I’d love to hear about what kind of data you used, what methodologies you used, that sort of thing.
Pradeep Ravindra: Absolutely. As I mentioned, there’s a lot of data points that we collect in the manufacturing process. One of the keys is developing a data model and a data governance strategy. You also have to have the in-house talent to build the cloud infrastructure that you need. There’s multiple angles to this. There’s going the GMP-compliant route, which is totally feasible through cloud computing, but it is a lot more expensive and stressful. It’s not necessary for analytics in my opinion in a phase one. I think you want to be moving fast. You should deploy products that are used for informational purposes, but actual validation, you don’t need that until you go into high-scale manufacturing where your throughput is much higher, exponentially higher than in a phase one.
Ross Katz: Can you just talk briefly about what you mean by validation in that context?
Pradeep Ravindra: One of the key Code of Federal Regulations that everybody in the industry knows: 21 CFR Part 11. You hear it all the time. Talks about the importance of data, if you think about it. It is really just having a well-established data trail of everything you do, breadcrumbs everywhere you go. What do I mean by breadcrumbs? A system where I can’t just go in and edit something and nobody knows. There’s an audit trail that shows Pradeep logged in at midnight and altered this data. That’s logged in an audit trail. What else? Signatories. Signatories that log all the date-time stamps. This person made this change to this data point or this person submitted this data point at this time. All of these things are not rocket science, but they are absolutely crucial for compliance. What I mean by that is in a data analytics strategy, when you have validation, to be specific to your question, when you’re making GMP decisions off of the data, then the tool that you’re using should also be validated as a source of truth. That makes sense. Yes. When we’re talking about data analytics, generally you’re extracting data from a system, maybe even compounding it with data from other systems through some pipeline and smashing it together in some dimensions to produce an angle on that data that wasn’t visible before because they were living in their host systems and never been blended like that before. You had your peanut butter jar here and your chia seeds over here and you’ve been spooning them separately but you never thought what would happen if I combined them. What would happen if I blended them together? Data is much like that. You have dimensionalities and when it comes to GMP validation, you don’t necessarily need to be GMP-validated to draw insights from data because at the end of the day if you look at a product in a non-GMP domain like PowerBI and you’re gleaning some insights from it, you still have to go back to the source-of-truth system to action it. That’s what I mean is there’s this line you can tread where you aren’t making GMP-informed decisions but it is driving insights collaboratively. Even for non-GMP domains, maybe research is interested in this GMP data, but you don’t want to give the entire research team GMP access and have them go through all that training to get access to that system. It seems a little silly. If you could pull it out in a non-GMP domain, that would accelerate cross-functional collaboration without having to impede your GMP capabilities.
Ross Katz: In terms of validation, I’m imagining that’s the value proposition of the modern data stack writ large of getting the data out of your source systems, of warehousing it, of putting a semantic layer on top of it, of visualizing it, of having these diagnostic use cases that deliver insights back to the business. It’s a story that I think as data people we’re used to hearing, but I don’t know if you’re able to share, back to the value that you were driving, I’d love to just close the loop on that story.
Pradeep Ravindra: As I mentioned, one of the key phases in manufacturing these immunotherapies is the ability to grow them to meet that target dosage. To grow them you obviously need the proper instruments. I mentioned a bioreactor. Today there’s a lot of fabulous services on AWS that you can use to pull that data right out of these instruments. Quite literally what we did was we created infrastructure where we could stream live data from instrumentation, dump it into a visualization tool and put it up on a big screen so that everybody in the manufacturing facility could see the data in real time. Across all of the instruments that were being used if there’s multiple batches going on at the same time, you could look at all of them and drill into them. The ability to do that in the immunotherapy world is, and this is the shocking part, I guarantee you there’s lots of companies out there who have commercial immunotherapy products, commercial, not even clinical, not even phase one, phase two, commercial, that don’t have these capabilities, that are still using paper to log all their data. Sure. Imagine if they went with the strategy that had logged everything electronically. How capable would their data analytics team be in order to provide them insights? Maybe even real-time insights. One of the things that you’ll find in the GMP world is that you have a lot of conservative thinkers. You have the quality assurance folks, you have the regulators, and they’re generally conservative, not all of them are, but most of them are. It’s because the nature of their job is they’re so focused on compliance, on 21 CFR Part 11. That’s their job. They should be. But when you bring up things like AI, machine learning, deep learning, they all get a little scared. They’ll be like, what are you talking about? The thing is if auditors haven’t seen it before, they have no idea how to audit it. We’re in this constant tug of war where you have the crazy disruptors, the mavericks and the rogues, and then you have the QA people trying to bring them back into hey, you gotta stay within these compliance boundaries. Because anything that hasn’t been done already or vetted by the FDA is unknown territory and generally you want to shy away from that and you don’t want to be too disruptive, but you still want to enable people to be disruptive without having to constantly say generative AI in the pharmaceutical world, we shouldn’t be doing that. What I’m trying to say here is if you were to throw that all out the window and just start over from scratch with a blank canvas and say let’s focus on the solution first, let’s focus on solving the problem first, and then let’s think about compliance later. Let’s figure out how to make this compliant later. Let’s come up with something really cool and if it’s value adding then you’ll figure out how to make it compliant. Don’t worry about compliance right off the bat because it should be the other way around really is make it fit your needs and then transition it into the compliant world. Basically what we did was we took an array of AWS services and we showed that we could build an analytics strategy harnessing data from GMP systems provided in a non-GMP format that could advise business strategically to make decisions faster on trends that they wouldn’t necessarily see. When you’re talking about the insights that you can get from this data being so early in your clinical development cycle, it really increases the odds of success of your clinical trial, the probability of success I should say, because you’re looking at possibilities that you never even knew existed before. You’re looking at correlations that you wouldn’t even be able to draw because maybe it wasn’t in your IND filing. Maybe you weren’t thinking about a particular feature that might be important. The reason why this is important is because there’s a lot of things we don’t know about cell and gene therapy that we’re still trying to understand and unpack. The ability to bring in new features ad hoc and just say I want to compare these features, I want to see if there’s some correlation here and you’re like this is mind-blowing. We should include this as a metric that we measure. The FDA’s learning, we’re all learning and it’s important to keep feeding that fire, but at the same time you need to dwell within some realm of value and business because you could go really off the deep end here. I already brought up the buzzwords, I already brought up deep learning, machine learning, AI, and we’re not even that far into the conversation. But the reality is we’re talking about selecting features and trying to drive an outcome. One outcome I told you about was growing cells. Imagine if I could come up with a list of features, feed that into a deep learning model and could accurately forecast when these cells would grow and hit their target harvest date. Then I could inform the patient and the healthcare provider when I suspect their dose will be ready for infusion. That pays dividends. We’re talking about a business value that isn’t measured by dollars but the impact that has on the patient that is sitting there in anticipation. One clinician described it as pins and needles, waiting for their CAR-T to arrive. Their own cells genetically modified brought back to them to fight the cancer that ails them. That value is something else that you can’t really measure. The way to get there is by getting all of the data under one data warehouse and a unified data governance strategy so you can drive that synergy across your research, your CMC, and your clinical functions.
Ross Katz: That makes a lot of sense. I heard you talk about the compliance-oriented and conservative thinkers and how that creates roadblocks to this sort of exploratory exercise that you’re talking about. I think a misconception of data work is that it’s a deterministic engineering discipline rather than a creative R&D open-ended curiosity-driven discipline. When I hear you talking about getting the data to a place where your team is enabled to do that kind of exploration, to drive those insights, to answer the most critical business questions, to put those insights in the hands of the people who are making decisions day-to-day and then later on you can talk about how those get connected directly to manufacturing processes or how those get integrated into the product or service directly, but really what you’re talking about is getting the knowledge of what’s possible first from the data and just the ability to answer critical questions, not trying to jump directly to ML or AI or anything like that.
Pradeep Ravindra: Spot on. Ross, you brought this up, you mentioned a semantic layer that perked up my little antennas there because the semantic layer is really your precursor to harnessing any of that technology that’s sophisticated. Without a well-defined semantic layer, good luck harnessing the real potential of any AI whatsoever. Because at the end of the day you have these large language models, you have basically transcribing what does this abstract parameter and this value mean to a human? And what does it mean to a computer? That lexicon, that codex, whatever that is to decrypt that bridge between taking a, for example, a neural network, feeding features into a neural network and producing an output, there’s all that translation that happens in that model comes from a well-defined semantic layer. It’s funny to me when you see companies asking for an analytics guy with AI experience and then you ask them about their data governance and their data warehouse strategy and they’re just like we’re still working on that. You look at their business layer and it doesn’t even exist. It’s just raw data dumps and one of the problems with analytics is people don’t have domain experience in what they’re working on. You could take me, an analytics person, put me somewhere else, copy paste me into I don’t know, the finance industry. I don’t really know a whole lot about the finance industry. The most I know is maybe a little bit of SOX compliance, but that’s it. I will struggle probably for at least a little bit of time until I figure out what they’re trying to deal with. In the cell and gene therapy world, if you’ve worn many different hats and then you’re in analytics, the advantage is you know what the problems already are. You see the problem before it even manifests because you’ve done it before. You actually have a seat at the table of the business rather than having a customer service approach where you’re constantly like, hello, sir, how can I help you today? Instead it’s the reason why you’re having these issues is because you don’t have some visualization and some alerting to tell you this material is expiring in 48 hours or 72 hours and really that should be a Slack message or a Gmail, some sort of alert that should tell you before it even expires rather than having the worst thing that you can do as a data analytics person is build a bunch of dashboards but not make actionable insights out of them. At the end of the day you can’t expect people to stare at dashboards all day. It’s great to see this line graph trending. But what’s even cooler is not having to look at the graph at all and having a bot message you, a Slack message to say just so you know today there was this data point submitted that exceeded acceptable thresholds, click this link to see it. Then you click it and it brings you to the graph and you see what you need to see and then you move on with your day. What I’m trying to say is the best analytics products out there are the ones you don’t need to look at actively. That just tell you when to look at them. That’s the next step. I think the second step. I would say semantic layer, then your visualizations, then make the data talk, and then you can get into foundations of bring in the buzzwords.
Ross Katz: Talk to me about the semantic layer in the context of what you’re doing because there’s so many domains that are involved in your world. You’ve got scientists and you’ve got clinicians, you’ve got engineers, you’ve got operational people and business people. How does an organization develop a semantic layer over time or get yourself to a semantic layer that can drive the business effectively?
Pradeep Ravindra: I think one of the mistakes that people make when they’re talking about their semantic layer is the first thing that we need to do is talk to the business and determine what exactly are the ways in which we would want to understand the dimensions of the data we’re looking at. Right off the bat, because I have domain experience, I know for immunotherapies, one of the dimensions that’s very important is process days. Amount of days or time interval, time, a spatial complexity you can think of as time very simply. Being able to relate all the data that we have on a scale of time, process days, day zero to whatever the final day that we make the drug product. I know right off the bat that I need to be able to relate a particular process with a particular point in time. That’s one example. You should really think about what the final picture looks for your business before you even build your semantic layer because as we’ve discussed, a semantic layer is a translational tool to translate that into the final vision. If you don’t have the vision laid out, your semantic layer is not going to be very successful. The reason being is because a semantic layer’s completely relative. You can scramble a semantic layer many different ways to meet all sorts of needs. At the end of the day without the use cases in mind, without those business requirements, without a fundamental understanding of what the problems are you trying to solve, the semantic layer is just smoke. It’s supposed to transcribe a human value to the data. If you don’t know what value you’re trying to get to, how are you going to translate it in the first place? I would say that’s the big thing that at least I’ve learned.
Ross Katz: What I hear is the guidepost is to begin with the end in mind. You have to know what outcomes your business is trying to drive from the data in order to develop the semantic layer in the direction of those outcomes. Just to imagine these different domains, if the outcome you’re trying to drive is on-time delivery or time to deliver then you focus the development of the semantic layer on the elements of each model that build up that question that you’re trying to answer and then let the outcomes drive that process organically. I’m making some intuitive leaps there, but that’s what I heard from what you said.
Pradeep Ravindra: You nailed it actually, because your semantic layer doesn’t have to be one, it can be separate business functions. You have a semantic layer for supply chain to understand it in the dimensions that they’d like to look at data. As you mentioned, on-time delivery, understanding how long each of your processes take to manufacture the drug product and being able to target longer processes that may be underperforming and say we need some operational excellence here, we need to apply some lean principles and improve this, without having that semantic layer that breaks down each process into a segment of time, without having that decomposition take place, you’re never going to be able to surgically target your lean activities, your operational excellence strategy. That’s a great point. I think what you said nailed it. You don’t need to have the entire value proposition in your hand to start building the semantic layer, you just need to have one of them. At least one thing to get started.
Ross Katz: This is another area where being in the cloud I think helps is that when you’re on-prem in the manufacturing database directly, you can afford to do different things than you can when you’re in a cloud data warehouse where the resource is there to support you in doing those less efficient things that drive the end user experience from an analytical perspective. One of the things that you’ve gestured to that I think is really important is you’ve acted out the conversations that you’re having with your stakeholders. There’s this two-way share of value that happens with stakeholders in an organization like this, where you need them to help you understand what’s happening in these source systems because they’re the domain experts in their part of the business and they need you to translate that into the data model, into the semantic layer so that they can get answers to the questions that they need. I’m just curious how do you go about that process of collaboration, of getting buy-in from all of these different stakeholders that you need to have that two-way exchange of value?
Pradeep Ravindra: That is an excellent question. I know 100 percent Ross that there’s a lot of data professionals that work in a field maybe similar to mine, maybe just like mine, that struggle with this. A lot of times organizations or entities like mine exist within IT and we’re treated like IT business partners. Truth be told I’m basically a product manager if you really want to just lay it on the table. That’s what I am. My product is manufacturing analytics. That’s my portfolio of products that I manage. We can’t be treated just as this separate island of an IT organization which happens. A lot of places you go, IT’s living on their island and nobody really knows what IT does unless their computer’s not working or they can’t log in to their email. Nobody really understands those implications of what IT does. Similar in the data world, a lot of people don’t really understand data analytics, how it works, what it does. Forget even talking about AWS cloud services. If I bring up Lambda and Postgres in a conversation with a business user they’re probably going to go cross-eyed and be like what’s this guy talking about? You want to keep that vocabulary out of your mouth, remove it as far as possible. I think one of the things that’s enabled me in particular to be successful in my role is that I actually genuinely read all the materials I can on the science itself. I’m not an expert at the science, but I can understand it at a high level. In order for you to have a seat at the table, to actually be there for strategy and making decisions, you need to have the domain experience where you’re not just sitting there and people are talking about very specific topics about the business and I’m an IT guy, I’m a business operations business analytics intelligence guy, I don’t need to understand this side of the business. That mindset, that mentality kills a lot of people because they’re just I just need to get really good at my craft. Really what benefits data analytics professionals the most is being a generalist. Is saying I don’t know a single thing about biology but I can learn today. Let me learn what is CAR-T cell therapy? Very simple answer, learn the business. Be able to understand at a high level what your leaders are thinking strategically, understand the risks that they’re trying to manage and develop propositions to help manage those risks, to develop propositions to help achieve those corporate goals. Because at the end of the day that’s what analytics does.
Ross Katz: Just looking toward the future, you touched on it in terms of potential applications for machine learning, but what are the emerging trends or technologies in data and analytics that you foresee impacting your work in CAR-T manufacturing?
Pradeep Ravindra: I think I mentioned the availability of data to be an issue, the volume of data that we have is one of the barriers to using some of these emerging technologies for the benefit of immunotherapy manufacturing. There’s some interesting concepts out there of simulating patient data like real-world evidence for comparing your clinical trial results to. As you might be aware, the idea is you have the standard of care for a drug product and then you have your experimental drug product and you want to compare the two and you want to show my drug product is outperforming some of these commercial drug products. Why shouldn’t it be commercial? That’s oversimplification, but bear with me here. One of the cool things that we’re seeing in the space, and this could be applied to the GMP world, is, and you’ve probably heard this buzzword digital twin, is it possible in the immunotherapy world to create simulations of data? Maybe even I don’t know, metaverse? That’d be really crazy. We’re talking about taking simulations and being able to reproduce them to wet lab experiments. That’s already being done today. There’s a video actually done by my current manager on Levinthal’s paradox, which is talking about the number of possible protein structures that exist in protein design and how they applied high-performance computing from AWS to reverse engineer these proteins. Anyway, the idea that we can simulate these experiments and be able to reproduce them in a wet lab. Think about it. We’re talking about virtual materials, virtual resources versus physical resources. Which one’s cheaper? Not even a comparison. If we could somehow do that on a reliable scale where we could say this model that we’ve built, this metaverse that we’ve built actually ties back to the wet lab. What we’ve done in the metaverse will happen in the wet lab also. That would be profound. I think that’s where we’re heading, is simulations. I don’t know if regulators will be open to the idea of using simulated data to expedite clinical trials, but I do think that’s where we should head especially in cases where it’s a rare disease and there’s not a whole lot of candidates available to thoroughly test your product.
Ross Katz: That makes a lot of sense. Before we let you go, are there any tools or frameworks or books that you’ve found to be particularly useful in your work recently or otherwise that you would recommend to people in the field?
Pradeep Ravindra: Reading material. Honestly I’ve been reading this book right now called Emperor of Maladies, which is an autobiography of cancer. I’ve been reading that to understand, as I mentioned for me it’s about understanding the science and the business side more. If you’re talking about the technical skills, it’s a fabulous time to be alive because if you’re trying to get a job in the industry then I would recommend a visualization tool, like PowerBI, Tableau, Grafana or Quicksight on AWS. Any of those are good. Some statistical analysis tool, R, SPSS, Prism. I think SQL has to be your backbone. You have to have a decent, not an expert, but just be able to navigate. Those three skills, I would say, just pick one of each and be able to demonstrate that you’re able to drive synergy across not only can I write the SQL to create the backend table, I can connect it to a visualization tool and build a dashboard for you. Then you have the ad-hoc analysis, which is in a statistical tool because you’re not going to use a statistical tool for a report, you’re probably going to do some ad-hoc exploratory analysis. You basically want to prove that you have all these capabilities under your belt. Honestly, I don’t really care about degrees anymore. I think we’re heading to a point in time where degrees are cool, I went to blah blah blah and I have a bachelor’s a master’s a PhD whatever. But at the end of the day it’s about the skills that you’ve acquired throughout the course of that journey whether it’s college or self-study online on Khan Academy or Coursera. I think there’s many different ways to learn and I wholeheartedly endorse skills-based interviews rather than credentials. I think you’ll find there’s a lot of smart people out there that will surprise you and they may not have the credentials you want, they may not have that PMP certification, but they’re brilliant in other ways. That’s what I would say.
Ross Katz: They’re able to read and understand the clinical trials that are publicly available on the internet.
Pradeep Ravindra: Yes. Resourcefulness. Get out there and just go learn something that you never thought you could learn. It’s so simple today. It’s the easiest time to do it. You just need the question of time and resources, but if you actually are passionate about it, you can learn whatever you want today. That’s a beautiful thing about being alive right now.
Ross Katz: For sure. Pradeep, I really appreciate the time. Such an awesome conversation, really insightful and look forward to connecting with you sometime down the line.
Pradeep Ravindra: Yeah, thanks so much for this conversation, Ross. I really appreciate the time.
Ross Katz: Of course.
Jason: And that’s it for this episode of Data in Biotech. If you enjoyed the episode, please subscribe, rate, or leave a review in your podcast platform of choice. If you’re a biotech company struggling to unlock a data challenge, CorrDyn can help. Whether you need to supplement existing technology teams with specialist expertise or launch a data program that lays the groundwork for future internal hires, you can partner with CorrDyn to unlock the potential of your business data today. Simply visit www.corrdyn.com to learn more. See you next time.





