Skip to content
Ross Katz — Reflections & Predictions: Year One with Ross Katz
Data in BiotechEpisode 33

Reflections & Predictions: Year One with Ross Katz

Ross Katz reflects on one year of Data in Biotech, sharing key lessons from 30+ episodes and predictions for biotech data science.

39:42Full transcript below
RK

Ross Katz

Co-hosts

Overview

In this solo episode, host Ross Katz reflects on one year and over 30 episodes of Data in Biotech, sharing the key lessons he has drawn from conversations with biotech data leaders. He outlines a future where biotech data becomes democratized, accelerating disease solutions through open collaboration and accessible tools. This episode is a must-listen for data leaders seeking to understand the biotech data landscape and prepare their organizations for a more integrated, data-driven future.

Key Takeaways

Biotech’s Data Landscape Today Mirrors Early 2010s Data Science.

The biotech sector, while rich with opportunity, operates within a nascent data infrastructure and a fragmented information environment. Data leaders must work through highly academic content and evolving open-source tools to build foundational capabilities. This landscape is ripe for standardization and improved accessibility, which will be critical to attracting and retaining computational talent and scaling advanced methods.

Bespoke Biological Processes and Equipment Drive Unprecedented Data Complexity in Biotech.

Each biotech organization frequently develops unique biological processes, purchases specialized equipment, and designs custom workflows, leading to highly complex and disparate data assets. This makes integration and standardization exceptionally difficult, particularly for organizations operating under capital constraints. Prioritizing strategic data architecture over reactive problem-solving is challenging but essential for long-term success.

Cultural Divide Between Discovery and Development Teams Hinders Unified Data Strategies.

A fundamental cultural split exists in biotech: discovery teams prioritize rapid, creative experimentation for insights, while development teams focus on scalable, reliable assays and regulatory compliance. This makes it difficult for data professionals, who rely on clean, consistent data, to contribute effectively across the entire product lifecycle. Mergers and acquisitions often address this economically but rarely bridge the underlying data culture gap without deliberate leadership.

Democratizing Biotech Data and Tools Will Accelerate Disease Solutions.

A future biotech landscape could see mass collaboration, open data platforms, and accessible computational tools like cloud labs. This vision allows decentralized groups to tackle long-tail diseases, drawing parallels to open-source software development. While barriers like data noise, experimental costs, and specialized expertise remain, historical trends suggest that currently specialized scientific endeavors will become more accessible over time.

Related: CorrDyn helps biotech organizations with data assessment and data engineering to build strong data foundations. Read how biotech manufacturers gain data value and consider our guidance on AI strategy.

Full Transcript

Jason: Hi everyone, my name is Jason and along with Ross I’ve been producing the Data in Biotech podcast since it began over a year ago. To date, we’ve hosted over 30 leading data experts from the biotech industry, exploring how they’re driving innovation, streamlining operations, and ultimately delivering more value to end patients. I’m proud to say that we’ve achieved over 13,000 unique downloads since we started, and we consistently rank as one of the top-performing life sciences podcasts in the US. And as we look at approaching the holiday season, we thought we’d try something a little bit different by putting our host, Ross Katz, in the interview seat. So this episode is a chance for you to get to know the person behind the podcast. Ross, welcome to Data in Biotech.

Ross Katz: Thank you. It’s great to be here and deeply uncomfortable with what we’re about to do.

Jason: It’s going to be great. Obviously you have a long career in the data science space and having listened to the last 30 episodes plus, you’ve drawn parallels often between today’s biotech landscape and the early days of data science back in the 2010s. Could you talk a little bit more about these similarities and what you’re seeing in that regard?

Ross Katz: Sure. When I first started getting excited about data science in the early 2010s, it was an arena where everyone very much knew that the opportunities were there and that these were the methods that were going to transform industry over the next 10, 15, 20 plus years, but the rungs of the ladder that needed to be there to help someone who was just starting out in their career or just starting to explore how to solve problems with data science, data science as a field was very nascent. The deep learning wasn’t really in vogue yet, there was still an ongoing conversation of whether R or Python would be the primary language of choice among data scientists. The main applications of data were in business intelligence, which I think was just called data visualization back then. But basically, you could see on the horizon that if you were able to sift through all of the very challenging open source GitHub repos and the very academic way in which a lot of the data science content was presented back then, that you could potentially gain the skills that you need in order to do this for real in the company that you were at. I wasn’t what I would call the best at that, but what I’ve seen over the ensuing decade plus is that the information ecosystem for data science has become so much more rich such that it meets you wherever you are. There’s content and there’s projects and there’s opportunities for anyone no matter what stage of the journey that you’re on, no matter how much exposure you have to data science previously, there are things that you can work on that help you improve your expertise and get you closer to applying the most advanced methods in the field. What I’m seeing right now in the biotech space is the beginning of this opening up of opportunity for people with an interest and with an understanding of data science and data science tools to come into the space and try to apply these methods, the computational methods, to biological problems. Obviously, I’m certain that people inside of the biotech industry have mixed feelings about that where people come in and they don’t know what they’re talking about from a biological perspective and that—and the truth is that in all very complex fields where there are lots of people working together to solve complex problems, you need people who are focused on depth and people who are focused on breadth. Obviously, the best types of people are the people who are combinations of both, but in truth, when you bring a team together, you need team members with different strengths. What I believe is that the industry could use more computational talent coming in and attempting to solve these issues. I also think that the way that the industry is structured right now, it’s undergoing a lot of changes where it’s uncertain how the biotech industry is going to look in five years or 10 years relative to the way that it’s looked historically. With the release of AlphaFold, with the release of Nvidia’s BioNeMo, with the recent Nobel Prize issued to David Baker and to Demis Hassabis, I don’t know if I’m saying his name correctly but from Google DeepMind, what we’re seeing is an opening of opportunity for people who have curiosity and have intellectual firepower and have computing resources, to which most of us have access to some computing resources, to apply ourselves to problems with real societal weight. It’s no longer the highest value that you can bring to society when working with data is by optimizing advertisements to be served to the correct person. On the horizon, you can see the opportunity where individuals in their home can understand a disease, make—collaborate online to make hypotheses about how that disease could be addressed through some sort of therapeutic approach and then computationally design and experiment with the creation of a pharmaceutical or some sort of therapeutic approach or diagnostic device that actually addresses that disease and then work with a cloud lab or partner with an organization that has the capital-intensive infrastructure to run experiments to feed more data into the computational approaches. That’s really what I see when I look at the biotech industry is just this—the beginnings of an ecosystem where maybe right now we only have the TensorFlow 1.0 version of opportunities for people to get involved and there’s just classes of people who are never going to be able to apply themselves to these kinds of problems in that environment, but in the next five to 10 years, could there be ways in—organizations like Hugging Face that make working with large language models so easy that people without expertise can fine-tune a large language model on a weekend, could that opportunity arise for more people across the open source ecosystem? I think the answer is yes. I’m excited to see how the biotech ecosystem evolves and I’m excited to play a role with this podcast in hopefully liberating some of the information from the walls of these biotech organizations where there’s not necessarily an incentive or a platform from which this kind of information can get shared.

Jason: So speaking about this new beginning, we know that with any new beginning, there will always be challenges and CorrDyn as a business, as a consultancy, is not exclusively focused on the biotech sector in your everyday work. Can you talk to us a little bit about some of the unique challenges that you see in working in biotech versus the other industries that you currently serve?

Ross Katz: These are problems that have come up in the podcast episodes. I think back to the Jesse Johnson episode where he was talking about all of the insights that he brings to bear from the scaling biotech substack that he has, and I think back to the Invert Bio episode where they’re talking about doing the unification of data assets on the fly from bioprocessing and what I see is that in some ways every biotech organization is unique in terms of the biological process that it’s applying and importantly in terms of the measurement methods and approaches that it’s using. Every biotech organization also designs its own processes and purchases specific equipment to fit into the processes based on the biological approach that they’re taking no matter where they are inside of the ecosystem of drug discovery, drug development, clinical trials, go-to-market. If you’re buying unique, highly precision-oriented cutting-edge equipment that does specific aspects of your process, if you’re designing that process by hand and that looks very different from the other biotech organization up the street, what you have is a situation where the level of complexity in the data that you collect and how those data assets connect together is just much higher than you would find at other industries. That’s what you heard from Nathan at Ganymede in terms of their approach, and that is in my opinion the most challenging part of working with biotech organizations and of understanding what’s going on in biotech organizations. All of the data is captured for a reason, but trying to figure out how to connect things together in a way that is sustainable, especially in an environment where you’re capital-constrained and where you’re trying to bring a new therapeutic to market or bring a new diagnostic device to market, or when you’re already in market and you’re trying to just manage cost in order to maximize the value of what you’ve already brought to market, it’s really hard in that environment to be strategic about the collection, the organization, the analysis, and the modeling of data so that you have standardized practices and processes and automated practices and processes in place for working with that data. What you see with Invert and what you see with Dave Johnson at Dash Bio is that they’re creating semi-vertically integrated solutions where the data collection and the data analysis and the machine learning and modeling, it’s all happening within the same sort of ecosystem within their domain because then they control all of the interoperability, they control all of the ways that the metadata comes together and there isn’t this need—the working hypothesis is there won’t be a need to reinvent the wheel every time you go into a new biotech organization if they adopt this sort of stack. But because biotech organizations take a lot of capital and a long time to bring to market, what you have is a lot of organizations with this path dependence where they’re always building on the foundation that they created when they were in that discovery phase and so reaching that level of data maturity either requires a lot of investment in tearing it down and building back up or adopting a new stack or just tackling the specific problems inside of the ecosystem that can be solved given the data that’s being collected and being strategic about collecting the new data that you’re going to need in one to two years in order to solve the next set of problems. But having that level of strategic vision is hard. Those are all things that we try to help our clients with.

Jason: You alluded to four or five episodes: Dash Bio, scaling biotech, Ganymede, Invert. We’ll drop the links to those ones that you referenced in the show notes of this particular episode. But I think you’ve also highlighted the kind of breadth of guests that we’ve had on the show over the last kind of 12 months or so and also kind of alluding to a divide between discovery and development cultures in biotech. That’s a theme that’s come up on a number of these interviews. Could you explain what you mean by this divide and then the implications this has for data science?

Ross Katz: I don’t think that I’m the first one to notice this divide. It’s basically the divide between the innovators and the operators, between the people who are creatively exploring a space in order to find—in other industries, you would call it product-market fit, but in this industry, you would call it biological molecule fit. And then there are the people whose responsibility it is to go through all of the hurdles of figuring out how to manufacture the thing that works, of trying to design the studies that prove verifiably that this therapeutic approach works, of organizing the trials and interfacing with regulators in order to ensure that these things work. I’ve always believed that where you stand is a function of where you sit, that your opinions, your viewpoints on the world are a function of your experience and the areas where you spend your time, the problems that you try to solve. On the discovery side, what you want is to be able to get to insight faster and close feedback loops so that you can answer the core question, which is, ‘How can we cure this disease? How can we treat this symptom? How can we reduce the toxicology of the therapeutic that we’ve developed?’ These sorts of questions. On the development side, it’s how can we design the assays that are going to scale, that we can apply in situ so that they’re not going to be extraordinarily expensive but they’re also going to be reliable enough and that we have some sort of statistical bounds around which we understand the reliability of what we’re doing. If you ask those two groups of people to work together—and by the way, data people tend to be more comfortable on the operational side because working in the computational space means that you need to have very clean, very organized, high volume data about the phenomenon that you’re trying to gather. And operating in the creative space where every experiment is different and where there isn’t a lot of consistency between the protocols and the steps that are identified makes it very difficult for someone in my world to do my job of leveraging large quantities of data to identify where the signal is and help to move things forward. As a result, you have this sort of cultural mismatch between the two groups. At the end of the day, these groups have to work together to be successful. I think the greatest biotech organizations are the ones where you’re able to demonstrate leadership, but also, I think this is one of the reasons why you see the industry set up the way it is where you have the creative discovery organizations that start small and then grow to a certain size and then they exit and are acquired by the large operators who know what to do once the discovery is already made. This makes a lot of economic sense for both sides, but it’s never going to change the culture of these things.

Jason: So we’ve started quite high in terms of analyzing the industry and your kind of key takeaways at a high level on the biotech space and the intersection of it with data science. I want to dive a little bit deeper now and shift gears into some of the specific kind of technical challenges and solutions that have come up time and time again during the last 30 or so episodes. And for listeners of the podcast who may have only—this is their first episode or they’ve only listened to a few, we have covered a wide breadth of different topics from Bayesian optimization, real-world data, knowledge graphs, etcetera, etcetera. But the three challenges that I want to talk to you about are feedback loops in drug development, finding that balance between automated evaluation and capital-intensive experimentation, and then finally foundation models. And so the first question I have for you is you and your guests on many occasions have spoken about the challenge of long feedback loops in drug development. Walk us through some of the solutions that you’ve heard from guests about trying to shorten these cycles, because the shorter they are, the quicker patients get help. What are your thoughts there?

Ross Katz: When we talk about feedback loops, the first things that come to mind are Markus Gershater and Synthace in the episode about design of experiments and then Wolfgang Halter from Merck talking about BayBE and open source Bayesian optimization for experimental design. The challenge—and this has been the key insight of software development over the last couple decades—is that the tighter you make those feedback loops, the faster you’re able to learn and the faster you’re able to optimize the process that you’re creating in order to get where you’re trying to go. It’s not as easy to run a biological experiment, to design a biological experiment as it is to run a line of code and check whether it has the output that you expect. But there are methods inside of design of experiments where there is now an increasing availability of automation for running these experiments across 384 wells at a time to learn as much as you can about the space that you’re operating in as quickly as possible. Get as much information as you can at one time using the design of experiments and then using Bayesian optimization, understanding how to leverage the information that you’ve gathered as well as you can to move in the direction of where you expect the optimal place to be—now this is obviously easier in the development side of things than it is in the discovery side. But if you can convert discovery into something that looks more like development in terms of automating the experiments and optimizing the parameter space so that you’re moving in the right direction based on the information that you gather, then you can make the feedback loops much tighter and you can thereby accomplish your goal much more quickly and gain momentum as an organization. What you see in the protein design and development space with Ryan Mork from Evozyne and with Mike Nally from Generate Biomedicines and with the team at Cambrium is they’re leveraging computational methods and a variety of models to molecules, run computational experiments to determine whether the proteins have the correct shape, have the correct characteristics, whether your probabilistic model believes that it’s going to accomplish the goal that you’re trying to accomplish with it, and then they’re able to run a bunch of feedback loops computationally inside of their GPUs to get closer to where they expect to be and they’re also using design of experiments and Bayesian optimization methods to select what is the point at which these experiments should be run and how can we optimally feed the information from these experiments back into the models in order to close the feedback loops still further. That’s really the way that I see all of these processes speeding up is just setting up the problem in such a way that you can move more quickly through the optimization process versus having to experiment blindly or just change one thing at a time, which I know is the bane of Markus’s existence based on what he shared.

Jason: What role do you see foundation models then playing in the future of biological research and drug development?

Ross Katz: I’ll admit this is one of the places where my level of biological background can fail me, but I’ll give you my best version of what my understanding is of how foundation models play a role in this ecosystem and I send an invitation out to the audience that if I’m not thinking about this the right way, I would love for you to educate me on where you think foundation models are going to make the biggest impact on the space. I think foundation models do a couple of things that are really valuable. The first is encoding insights about the biological landscape that you’re exploring into this latent space so that when you’re using a foundation model to understand what’s happening in your experiment or what’s likely to happen with a given molecule that you’ve designed, you can trust it to bring more insights to bear than you could with any of the heuristic-based methods or mathematical model-based methods that you would have used historically. And that allows you to unlock information that honestly hasn’t even really been discovered yet. Nobody knows in their head the entirety of what’s happening in biological space because despite the advances in imagery, most people do not have a lot of intuition about what’s happening at biological scale inside of a human body or inside of our ecosystem. And these foundation models, as a result of encoding a lot of data and a lot of experiments that have been run over a long period of time, are able to codify that knowledge to the extent possible in a way that makes it accessible for specific purposes that enable drug design or drug discovery processes or other sorts of models that you would build for going through the discovery and development process. That’s the most important place where I see foundation models being applied. The second thing that it does is as a result, in all of these models, there’s this embedding space and an embedding is a vector of numbers that represents how the model codifies an entity inside of its space. And what you can do is you can analyze the relationships between the embeddings and attempt to gain insight about the biological process that you would not otherwise be able to gain. Obviously the outputs of the models themselves are really useful, but a model is useful because it teaches you something—if you understand its constraints and you understand its limitations, then it’s often able to teach you something about the world that you’re attempting to model. It might be teaching you about the limitations of the model itself, but it might also be able to show you insights about proteins in a particular multidimensional space having certain characteristics that make them useful or make them have certain properties that would be useful in a given domain. What we’ve been told over the course of our conversations with all of the protein design companies is that the landscape is just so huge that what we’ve actually experimented with is just a drop in the bucket of that landscape. There’s this opportunity to view that landscape through a wider lens and glean information about where we should be exploring in that landscape in order to get answers to the specific questions that we want. The foundation models, the outputs of them and then also the underlying assets that they create like the embedding enable that kind of understanding.

Jason: The final technical kind of challenge that I want to cover with you is this idea around how do biotech companies strike a better balance between automated evaluation and capital-intensive experimentation? I think you’ve touched on it a little bit already, but elaborate on that for us.

Ross Katz: I think that the challenge is always where do you draw the line? Because if you continue to spin your models in their little ecosystem and have them learn from themselves, models evaluating models, then what you end up with is an uncertain bias getting injected into your process. What you need is the ability to ground your model in the real world and learn the areas where it doesn’t have enough information yet in order to fully characterize what’s going to happen when you run the experiment that is driving your models. I don’t know the method for how to strike that optimal balance and once again, invitation to my audience if you have insight into this, I would love for you to teach me. My understanding is basically that you allocate as much experimental space as you can and you pack as much experimentation as you can into the capital that you can allocate and then you maximize the value of each experiment by using the methods that we talked about earlier and then feeding that data back into the model. You want every experiment that you run to be the next most informative experiment that you can possibly run so that the next time you retrain your model or you feed the model the new data, your model is learning the most it possibly can and then you can have less concerns about the biases that the model might create if you’re continually evaluating the model using another model.

Jason: Let’s look a little bit forward because you have an interesting vision on democratizing the biotech data space. Could you share a little bit about what you see that ecosystem looking like?

Ross Katz: As I’ve been doing this podcast, I’ve been thinking about what the biotech ecosystem looks like that’s the analog for the data science ecosystem that we were discussing earlier, and what are the areas where there’s an opportunity for this sort of mass collaboration. I see some really great biotech organizations creating these competitions where it’ll be a protein design competition or they’ll create opportunities for people who want to learn to come in and try their hands at these computational approaches in the context of a specific problem that they’re trying to solve. That’s a great first step. The nature of the space of human disease is that there’s this long tail of diseases that are out there. Just as an example, my mother has primary progressive multiple sclerosis, the most common form of multiple sclerosis and where all of the research goes is remitting-relapsing. If I wanted to wake up one morning and attempt to design a therapeutic for reconstructing the myelin sheath so that my mom’s primary progressive MS was somewhat less detrimental, I wouldn’t know where to start or how to do it. Every disease touches everyone’s life through oneself or through a family member, so there’s always people who are motivated to do something about it. It’s just that there isn’t really an avenue through which that doing something can be done particularly productively other than throwing a donation over to an organization which is investing in the kind of research that you care about. But if you wanted to be involved in actually solving a problem that you’re passionate about, then you would need some way of understanding the nature of the disease. Open publishing platforms like bioRxiv are a really good move in this direction where the academic literature is more accessible than it’s ever been. There’s that as a starting point. And LLMs have introduced this opportunity for digesting really dense and inscrutable material that you couldn’t read previously in a way that can make it more actionable, but then you have the problem of hallucinations and also that if you’re not adept at analyzing research in a field that you’ve never studied, then you’re prone to draw conclusions that are probably not the right conclusion. So there would need to be some way of wrapping your arms around what is the mechanism of a given disease and then what are the sorts of druggable or therapeutic approaches that could be tried—these are hypotheses that are out there that could be tried. And then once you understand that there are things that could be tried, there’s a molecule that could be developed for a particular target, then you need a toolkit and a data set. This is another thing that’s talked about regularly—the Protein Data Bank and the noisiness of the data, there’s smaller subsets that you can use that are more reliable, they’ve been verified, but regardless, you need a highly specialized set of data for each disease category. I’m summarizing how biotech organizations come to exist. Classically it’s somebody has an insight about the disease, they believe that if they develop an approach to this particular disease mechanism, then they can found a company on the back of it and then all they have to do is prove it and then they can scale it up and sell it. But if the data were more freely available and/or if there were an opportunity to pool resources and pool activity around cloud labs or around individualized experiments that could then feed that kind of computational pipeline, then you would have the beginnings of an ecosystem where you could see decentralized groups of people coming together to create a cure for a given disease in the same way that decentralized groups of people come together to create an open source tool for solving some sort of software issue. I don’t know how long that takes. I don’t even know if that’s economically feasible in the biotech landscape that we live in where the cost of generating data and the cost of running experiments and even just the level of expertise required in designing assays to run experiments in a way that leads you in the direction that you’re trying to go, there’s so much expertise locked in there. But what I’ve seen over the course of my life is that what seems like the realm of only the top PhDs today becomes the realm of your hackers at home a decade later.

Jason: As we take our crystal ball out of the cupboard and try and look a little bit further into the future, what emerging topics, trends in biotech and specifically biotech data science are you most excited about to explore in your second year of hosting the podcast?

Ross Katz: We’ve touched on a lot of them in this conversation already. I think the foundation models that have come out recently are really interesting. I mentioned Nvidia’s BioNeMo, which is a package of a variety of models. I’m really interested in the new AlphaFolds that are coming out. I’m really interested in ESM3. There’s these models that are coming out and I’m interested in the differences between how they’re trained, what’s the vision for how they fit into a drug discovery or development workflow, because what I’ve heard from our guests is that it’s an ensemble approach. There’s a variety of models that are being brought to bear. I’m always interested in learning about what is the ensemble, and also within that ensemble, what are the foundations that people are using and how are they being applied? How are they being trained? What are the assumptions that the models make? These are answers that I find difficult to glean as an outsider, and I imagine I’m not the only one. I’m also really interested as an extension in other generative approaches to biology. Most of these are based on foundation models, but among the design-build-test-learn loop, the design is one of the places where there’s still so much opportunity. I’m interested in hearing more about that. There are also new approaches to therapeutics that I’m interested in exploring—cell and gene therapy. We’ve had a couple of episodes with Atara Biotherapeutics and with Lyell Immunopharma talking about these, but these are emerging modalities that I’m interested in exploring further. Any ways that AI is being applied to the biotech space is interesting to me and there are just so many unique ways that these new capabilities that have come about, in particular large language models, but also image-based approaches or even video-based approaches. I’m interested in the intersection of where new abilities to image or measure what’s happening in biological phenomena—so the creation of new types of data—leads to the ability to gain new insights or create new models or create better ways of visualizing something that we had never visualized before. Parul Doshi from Cellarity, we were talking about their cell visualization toolkit. I’m really interested in the ways that these new assay technologies and the new manufacturing technologies, the new experimentation technologies create new types of data that make it possible to gain insight in ways that people haven’t been able to gain insight before. We haven’t talked to anyone from the real-world data world in a little bit, but Vira Mukherjee from Datavant and Lana from Novo Nordisk brought some really interesting insights on the real-world data portion of this and we were talking about feedback loops earlier. I think the most important feedback loop is with the patient. In our conversations with the different manufacturers, what we’ve heard is it would be great if we understood how the patients are responding to the things that we’re manufacturing, but there are all these barriers to closing that feedback loop and real-world data offers that opportunity to close the feedback loop in a way that wouldn’t have been closed before. I’m interested in what are some innovative ways that people are, similar to Datavant, doing an end around all of the privacy and anonymity challenges that you have in terms of getting access to that data, quality-controlling that data and actually applying it to the discovery and development work that they’re doing. We mentioned democratization earlier and one of the areas that I see really being interesting is in cloud labs. My understanding of cloud labs—I would love to have somebody on to give me a better understanding—but basically you think of a cloud lab as the AWS or the Google Cloud or the Azure of laboratory equipment where you submit a set of experiments or an experiment and then the cloud lab does all of the hard stuff and then delivers you the data back and you don’t have to purchase the high capital infra in order to do that work, you can run your experiments without having to own that capital-intensive equipment. If you’re one of the people who can be very parsimonious about the experiments that you run and gain the most possible information from those experiments, then there’s this opportunity to potentially get all the way through a discovery phase without having to own your own experimentation equipment. As the ecosystem evolves, I’m really interested in how CROs and CDMOs are looking different. We talked to Mo from Sapient and back to Dave Johnson from Dash Bio about a different vision for what a CRO looks like. It just strikes me that there are a lot of organizations in this ecosystem who are peeling off a piece of the discovery or development workflow and making it super efficient or making it a lot more data-driven or making it a lot easier for a biotech organization to interface with without having to invest in bringing all those capabilities in-house. Maybe they can get better insights than they would be able to previously, maybe they’re just able to get to scale faster and more economically than they would be able to otherwise. But I’m interested in vendors that are playing that role. And any other interesting vendors in the biotech space who are using data in fun ways to increase the efficiency or improve the outcomes that biotech companies are able to face. We’ve had some really interesting vendors on like Joseph from OmicSoft, Jonathan Eads from Genomenon, just people who are enabling the research process to go more efficiently. There’s just a lot of opportunities to explore how data-driven technology can make biotech organizations work better.

Jason: As we look at finishing up the interview here, what final advice would you give to data scientists that are interested in entering the biotech space?

Ross Katz: I’ll open the curtain a little bit here and just admit that starting this podcast has been an emotional journey for me. I have always been interested in this arena but have not been willing to ask the questions that needed to be asked for me to start to understand what I need to understand in order to grow in the way that I serve biotech organizations or the way that I work in this industry. As I’ve grown as a data scientist, what I’ve discovered is that even though I might have a spike of imposter syndrome or of anxiety around asking somebody really smart a question that sounds really dumb, over the long-term horizon, I’m usually glad that I asked the question. My advice is no matter where you are in your journey, just be willing to ask the question that sounds dumb to you to the person whose intellect you respect because if you have imposter syndrome, that means you’re in the room with somebody who could make you smarter, probably, who has knowledge that they might be willing to impart to you and if they’re not, if it’s not the right time for them, then that’s not your fault. You lose nothing by asking. Figuring out how to find creative ways to check my ego at the door and let my curiosity guide has been the most important insight I’ve gained from this first year of doing this podcast and it’s an ongoing battle but it’s something that I hope I can continue to embody and for all of you who are doing this kind of work out there, I recommend you try to embody that as well.

Jason: Where can people learn more about you and CorrDyn?

Ross Katz: Our website is corrdyn.com, C-O-R-D-Y-N, and the best place to follow me is on LinkedIn. I’m B-dash-Ross-dash-Katz on LinkedIn. Please connect, follow, share resources with me. I’m really interested in hearing from you. If you have suggestions for guests, if you have suggestions for questions that I can ask on this podcast that would be more interesting to you, if you want to collaborate in any way, I’m just interested in understanding what I can do to make this podcast more valuable to you. We care about getting our name out there as a data science consultancy that does this kind of work, but we also just care about building community around the podcast and making this valuable to you. If you have something to share, I think LinkedIn message is the best way to reach me.

Jason: Ross, thanks very much for being a guest on Data in Biotech.

Ross Katz: It’s been my pleasure. Thank you.

Jason: And that’s it for this episode of Data in Biotech. If you enjoyed the episode, please subscribe, rate, or leave a review in your podcast platform of choice. See you next time.

Frequently Asked
Questions

How can our biotech organization overcome the unique data complexity Ross describes?
Begin with a comprehensive data assessment to map your specific biological processes, equipment, and data flows. Focus on strategically collecting new data that anticipates future needs, rather than solely reacting to immediate problems. Prioritize foundational data quality and integration in critical areas to build a sustainable, scalable data architecture.
What is the biggest challenge for data professionals looking to bridge the discovery and development gap?
The core challenge lies in the cultural mismatch between creative, often unstructured experimentation and the need for operational standardization. Data professionals must demonstrate value by translating discovery insights into reproducible, quantifiable metrics, while also advocating for consistent data capture early in the discovery phase to enable later development and regulatory compliance.
How realistic is Ross's vision for data democratization in biotech, given current costs and expertise barriers?
While significant capital, specialized expertise, and data quality challenges currently exist, the trend towards open data platforms, advanced foundation models, and accessible cloud labs suggests increasing feasibility. History indicates that complex, specialized fields gradually become more democratized and accessible to a broader community over a decade or more.

Need a data partner for life sciences?

CorrDyn helps biotech and pharma companies build the data infrastructure that accelerates research and operations.

Book an intro call