Listen on
Overview
Host Ross Katz speaks with Jacob Oppenheim, Venture Partner at RAVen. Biotech companies frequently face a false trade-off: invest in advanced science or build strong data foundations. The issue isn’t a lack of tools, but a profound mismatch between what’s available and what agile scientific R&D demands. Generic enterprise systems, designed for industrial manufacturing, often become operational drag, while modern data stack components, built for software engineers, fail to meet the specific needs of scientists. This gap prevents effective data utilization, slows therapeutic development, and siphons resources from core innovation.
Jacob Oppenheim, a Venture Partner at Digitalis Ventures with a background building data science teams in biotech at GNS Healthcare, Indigo Ag, and EQRX, offers an insightful perspective. He explains how this tooling deficit stifles innovation and prevents companies from extracting real value from their data. His venture role gives him a unique view across the biotech sector, identifying the persistent challenges in data management and the opportunities for fundamental change.
This conversation explores the inherent limitations of current biotech software, the critical need for modular data systems tailored for scientific use, and how strategic data foundations can accelerate discovery. It’s a call for biotech leaders to rethink their data strategy, move beyond piecemeal solutions, and focus on building an ‘operating system’ that truly enables scientists and drives therapeutic progress.
Key Takeaways
Existing biotech data tools create operational drag, not efficiency.
Many enterprise software solutions in biotech, like Laboratory Information Management Systems (LIMS), are designed for the rigid processes of a manufacturing line, not the dynamic, evolving needs of scientific research. This mismatch forces companies into complex, expensive implementations that offer poor user experience and hinder, rather than support, exploration and agility. Teams spend disproportionate time configuring and training on tools ill-suited for their primary function, delaying scientific progress.
The true value in biotech comes from digitized systems, not just AI algorithms.
Focusing solely on advanced AI and machine learning without strong data foundations puts the cart before the horse. The real competitive advantage in biotech stems from the ability to generate, store, organize, and operationalize data effectively. Just as the FinTech industry saw massive growth through digitizing transactional data, biotech’s next leap will come from building digital systems that enable efficient data capture, movement, and analysis at every stage of discovery.
A ‘biotech data operating system’ must be modular and scientist-centric.
The current tooling landscape lacks integration and follows a ‘one app to rule it all’ model, which limits flexibility and functionality. While modern data stack primitives serve software engineers well, scientists require specialized tools for data quality control, analysis, aggregation, and generating operational metrics from scientific data. Building a modular data ‘operating system’ that seamlessly connects various functions and provides intuitive interfaces for scientists is essential for accelerating research and decision-making.
External data expertise helps focus internal teams on core IP and accelerate growth.
Biotech companies should concentrate their in-house talent on their core scientific IP and therapeutic development. Attempting to custom-build complex data infrastructure or manage the rapidly changing vendor landscape internally diverts critical resources. Partnering with data consultancies can provide the specialized technical expertise needed to assess requirements, select best-in-class tools, and implement scalable data systems, allowing scientists to focus on experiments and discovery without the burden of infrastructure management.
Related: CorrDyn helps biotech companies build data foundations and assess their current data maturity. For deeper insights into this sector, explore our biotech and life sciences industry solutions or read how biotech manufacturers gain data value.
Full Transcript
Jason: Hi everyone, this is Jason, producer of Data in Biotech. Before we get started, I wanted to let you know about our latest white paper. It’s a comprehensive guide to implementing machine learning models in biotech manufacturing. It’s a complete overview of all the potential problems of ML adoption and more importantly how to solve them. To download it, simply visit connect.corrdyn.com/biotech-ml. We’ve also dropped the link in the show notes of this episode. Okay, let’s get into it. Welcome to Data in Biotech, a podcast from CorrDyn where we explore how companies leverage data to drive innovation in life sciences. Every two weeks we sit down with an expert from the world of biotechnology to understand how they’re using data science to solve technical challenges, streamline operations, and further innovation in their business. This week, we sat down with Jacob Oppenheim, entrepreneur in residence at Digitalis Ventures, a venture capital firm that invests in solutions to complex problems in human and animal health. During this interview, we discussed the importance of establishing strong data foundations in biotech companies and the limitations of existing tools in the industry. Why the current data tools and systems are more suited for coders and engineers, leaving scientists struggling to extract meaningful insights from the data, how the lack of modularity and integration between different tools and processes is a major challenge, and finally why consultancies play an important role in providing expertise and guidance to biotech companies, helping them navigate the complexities of tooling and data management. Here we go.
Ross Katz: Jacob Oppenheim, welcome to the Data in Biotech podcast.
Jacob Oppenheim: Thanks for having me, Ross.
Ross Katz: Awesome. Well, just to kick us off, can you give us a brief introduction to your career to date?
Jacob Oppenheim: Yeah. Originally, probably half or more of the people in data science and machine learning, I was some kind of physicist. I got a doctorate in biological physics, doing what I think today people would call machine learning to understand both how the brain processes sound and how the structure of blood vessels are orientated within entire organs. I managed to graduate with that PhD right at the end of the AI winter and before deep learning became something people talked about, so I had no idea how to talk about what I was good at in the first place. I then spent about a decade building and running data science, machine learning, and software engineering teams in biotech. First company I jumped into, GNS Healthcare, I was designing novel machine learning algorithms to understand genomic data. The arc of my career is really trying to get more and more at the basics. I started at this machine learning-focused consultancy, and I realized, hey, we don’t actually have any data. It’s just our clients’ data here, and God knows what happens when our algorithms are applied to it. We never see any results back. So I was like, okay, I’ll go join a biotech, and I’ll start with the statistics of their data and see what we can do there. I joined Indigo Ag. I was there for five years. I built probably every data science team that Indigo ever had. That started with just how do we analyze massive amounts of laboratory data, how do we analyze agricultural field trials, which is actually funny as a statistics problem. That’s how modern statistics was actually developed—looking at fields of corn in England. That’s what R.A. Fisher was doing. So it was fun. I got to go back to the basics in some ways. Then computational biology, bioinformatics, remote sensing and image processing, a lot of interesting stuff there. But at the end of the day, what was interesting is a lot of the things that mattered most were just about could people use the data effectively? Could we store it effectively? And could someone who wasn’t a data scientist actually put all the pieces together about—we were looking at microbes mostly to improve agricultural performance. Could somebody put all the data together and draw conclusions from it? The answer to that was usually no because the systems and tools to work with the data were not established from the beginning. We were constantly behind in trying to build a strong foundation for data in this biotech company. That was uncomfortable. At some point I realized I need to join a company even earlier so that we can establish the fundamental predicates of all of this from the beginning and establish the right culture and the right tools. I was at EQRx for three years from not long after it was starting until it wound down, really trying to use data to enhance the mission of EQRx. The idea of EQRx was to lower drug prices by selling directly in volume to insurers, providers of healthcare, national healthcare systems. One way you can justify lower prices is make it up in volume. Another way is to be much more efficient in everything you do. There was a synergy between the management of EQRx and hiring me because the idea was that we would build the right data platform from the beginning that would make everybody much more efficient in what we did. That meant starting at that base systems level and building up from there. That was really exciting because one of the things I’ve always found is that the most interesting questions are the ones that come from building up from the beginning and are oftentimes surprising from the outset. When EQRx wound down about a year ago, I started looking for the next thing and I’d been thinking very deeply about what I had learned across my career so far. I made a jump into venture land. I’m an EIR at Digitalis Ventures. We’re an early stage biotech fund based in mostly Boston and New York, and I am working on novel solutions for technology in biopharma.
Ross Katz: So, you’re EIR—so entrepreneur in residence—at Digitalis Ventures. Can you just talk to us a little bit about Digitalis and what its mission is and how you contribute to that mission in—in the role you’re in now?
Jacob Oppenheim: Yeah. Digitalis has historically been an early-stage biotech fund, half in human health and half in animal health. Very scientifically focused. One of the things that attracted me to Digitalis is, honestly, their portfolio companies say they’re one of the most involved investors and the investor they like the most. Frankly, I think if we’re going to drive change in science and in industry, we need to start at the earliest stages. Early-stage funds which have that collaborative relationship with their portfolio companies and who are deeply focused on the science—that’s who I want to be building technology with. If I want to start up a technology company for biopharma, I want to be working with people who have a good relationship with their portfolio, I want to be seeing the novel science, and I want to be thinking deeply about what’s coming up and what are the primitives that are going to allow them to excel in the future.
Ross Katz: Can you give us a—a few examples of like what that means in the context of biotech for a venture fund to be more deeply involved? So what are the kinds of capabilities or activities that Digitalis is doing that you that some other venture funds might not do in a similar circumstance?
Jacob Oppenheim: Yeah, in early-stage biopharma, there’s a lot of questions about how do you take interesting science—be it from an academic co-founder or stuff that’s right out of the lab—and prove it and turn it into something real. How do you go about picking targets? I think that’s the thing our life sciences team spends the most time working with the portfolio on: how do we pick targets? And then how do we prove out the technology? How do we get to the fundamental milestones that show that we can deliver on something? These translational aspects tend to be very scientific and ones where there’s a lot of back and forth and working closely together. I’d call it scientific strategy overall.
Ross Katz: Yeah, that makes a lot of sense. And as entrepreneur in residence, it sounds like you get exposed to the entire ecosystem of biotech startups and the technological and scientific challenges that they’re trying to solve—where the tooling ecosystem is supporting them, where it’s getting in the way. In the context of that perspective, what are the biggest challenges and opportunities that you see, especially for data scientists and data teams but the startups more broadly, facing in the biotech ecosystem today?
Jacob Oppenheim: Yeah, I think as anybody who’s ever worked in a biopharma company knows, basically all the software that you’re given is totally crap. Has a horrible user experience. It looks like it came out of the Windows 95 or worse yet the DOS era. It’s not at all configured in such a way that things make sense or to encourage exploration and it barely works for what you’re doing in the first place. We have these terrible tools and they work for small molecule drug discovery, but a lot of what we’re doing today is not small molecule drug discovery and there we basically have no tools. You see this everywhere. People are discouraged when they’re a small company without much money from investing in digital tools in the first place—for even tracking what they’re doing or communicating more effectively within a company—because it’s not clear what the value of it is. And even if there is value, it costs so much in terms of time, training, and sheer cash to bring in tools that you can’t even make a logical distinction between well, I’d like to save these data, I’d like to do this a little bit better, let’s bring in a tool here. That calculation almost cannot be made because the configuration, the training, to use any such tools is beyond the scope of what somebody’s doing. I’ve actually sat down with surprisingly large companies and tried to talk about what are the tools that could be really valuable? What’s the lightest, simplest thing that you could put in because something’s finally becoming enough of a pain point that it might be worth dealing with.
Ross Katz: So it sounds like there’s this tradeoff between investing in the tools and the platforms to develop what you’re doing and investing in the core science—validating a scientific hypothesis early in these biotech companies. As a result of that tradeoff where there’s not enough time, not enough resources, not enough talent to focus on both, you have to choose one or the other.
Jacob Oppenheim: I think people get forced into a tradeoff where there shouldn’t be one in the first place, but everything from the economics to the commitment required to use a lot of these tools makes it quite difficult. I’ll give you an example that’s somewhat further afield, but almost everybody with a lab has a LIMS system, laboratory information management system. What is a LIMS? A LIMS is a machine or a set of software designed for doing things like tracking screws in a Ford factory. It wants to know about all the components that you have, it wants to track them, you scan them in and scan them out and follow your processes through. Almost nobody in biotech is at the scale where a system like that makes sense as a core operational driver. Should you have tools to track processes and what’s actually happening in the lab? Absolutely. But they need to be pretty damn flexible because what you’re going to do day to day is going to change as your science evolves, as your processes evolve. When the barrier to entry is bringing in a tool that feels like you’re running a Ford factory, it’s very hard to make rational tradeoffs. It’s very hard for you to say I’m going to invest this much now. That also makes it difficult to bring a higher level of operational order into how people run a laboratory, how people industrialize and translate academic science into something that could be used for therapeutics development. That make sense?
Ross Katz: Yeah, it does. What I’m hearing is there’s a bimodal distribution of the tooling ecosystem where you’ve got these very immature tools developed in academic labs that don’t scale very well, mostly on the analysis side. But then on the data capture and preparation and data strategy side, you’ve got these really heavy tools that require a level of scale that most of the companies in the biotech ecosystem you’re exposed to currently aren’t at—they’re not at the phase where they need this level of tooling. Getting people from the data strategy at the beginning to working their way up that maturity curve is a challenge.
Jacob Oppenheim: Yeah, that’s right to a very large extent. And what’s crazy is that having some kind of system to work with data and operationalize what you’re doing in your laboratory—that should be one of the first things you need to do. You can always hand-code an analysis, but setting up tools to manage data, to manage your compounds and register all the molecules you’re making and what batch they came in—this is fundamental, because this is going to track you all the way through. If you need a laboratory notebook, this should be simple and easy, with all the tracking and some basic automation in there to make a scientist’s life easier. But in most cases, these tools are big and heavy and expensive and feel ancient and archaic. I’ll give you another analogy. When you open up a new Mac computer, there’s this weird version of Microsoft Word. I forget what they even call the word program. I think the spreadsheet is called Numbers. I don’t know if anybody has opened these programs meaningfully in at least a decade. But Apple still occasionally thinks on this old design paradigm of I’m going to own everything you do on a computer, which I thought went away with the Microsoft antitrust suit in 2003. The software in biopharma thinks it’s one app to rule it all. It’s terrible at everything else except for one or two core features. This, too, adds to the expensive, slow to configure, hard to integrate pieces, making it really hard for scientists to accelerate what they do with modern tools.
Ross Katz: Yeah, that makes a lot of sense. Over the past several years, it seems like there’s been this explosion in AI for biotech companies and then a bunch of backend software-as-a-service type applications that on their face seem meant to solve some of the problems we’ve just discussed. From an ecosystem perspective, what are these startups mainly getting right? And what do you think they’re mainly getting wrong, as an ecosystem?
Jacob Oppenheim: My perspective is that on some level these tools are putting the cart before the horse. Some of these things can be very useful in specific use cases. But the key problem is people don’t have data stored and organized in the first place. You say we’re going to do AI and it’s going to be valuable. Well, everybody’s doing something different in biopharma. Their science is somewhat different. The value that you have, your competitive moat in biopharma, is the data that you generate and your ability to operationalize it—and note I’m not saying AI, data science, ML, anything there. It’s the ability to store your data, do interesting science, and use that to inform the next science that you need to do. I’m writing myself out of a job here. There are interesting applications where people have done interesting stuff. But I would be concerned that we’re missing the fundamental value driver here. Let me give you an example. About 15 years ago, there was no fintech industry. If you remember 15, 20 years ago, you still had to go to a physical bank to do nearly anything, and only if you had Charles Schwab or one or two other things could you even meaningfully bank online. You used to have to pay a lot of money to trade stocks. It was crazy. Since then, finance has digitized and moved online in a way basically no other industry outside of tech has. The size of the fintech industry, which basically didn’t exist 15, 20 years ago, is 245 billion dollars. That is an industry entirely premised not on AI, not on machine learning, but on moving numbers around—on capturing, storing, and moving data. That’s fundamentally what’s going on: transaction processing, moving things from place to place. My pushback on some of this stuff is: look at where digitization has happened. Is the value in finance in AI, in fintech AI, or is the value actually in building digital systems that allow this industry to move so much more efficiently and people to do so much better with what they’re doing? I’d argue the value is very much in that space.
Ross Katz: I’m wondering from your perspective, if you were designing that entire ecosystem from scratch, what would be the components of that process and what would it look like?
Jacob Oppenheim: I think this is a great question, because right now we have pieces of an ecosystem and every piece is trying to empire-build and expand too much. But the problem is that we don’t even have all the pieces that we need. We don’t even have an integration layer. It’s like saying you’re building Word and Excel and PowerPoint, but you don’t have Windows. You don’t have an operating system. So it’s hard to even see how the pieces fit together. I can look at categories in the market. ELN—it’s useful to have a record-keeping system especially if it automates tasks for scientists. Why does ELN work in chemistry? Because it automatically calculates a bunch of parameters that you need to set up a reaction. Does your stoichiometry, does your weights, does a bunch of calculations on molecular structures for you. We need a piece to track everything in the laboratory and sign it in and out. We can call it LIMS, but it should be lighter. We need a piece to register all the entities we’re going to be using that are potential therapeutics. That’s our compound registry. Are any of these things going to be interacting with each other? Maybe they’ll touch each other slightly, but all these pieces are touching on data. We don’t even have a place to put data, to organize it, and to operate with it on multiple levels. And that doesn’t mean just dropping it in somewhere, but reporting on it and making it useful to people. On top of that, you need tools around decision-making, project management, collaboration. People are touching at pieces of this ecosystem, but I’m not sure we even know what the entire ecosystem needs to look like in the long run.
Ross Katz: Yeah, that’s really interesting. It reminds me of how, outside of the biotech industry, the so-called modern data stack has tried to become that hub of that system by integrating data in the data warehouse, but the operating system you’re talking about is a deeper level of integration.
Jacob Oppenheim: Part of the issue here is that we’ve got a bunch of great tech primitives that the tech industry has developed over the past 10 years. This is hard to see unless you’re deep in software engineering and data engineering. The expansion of developer tools has really changed our ability to build things very rapidly. You bring people in, okay, we can build a data warehouse on Snowflake, we can use Fivetran to keep our databases in sync, we can put a front end on stuff. That’s great. But those are primitives for people who are coders. They’re not primitives for scientists. What happens is you drop data somewhere and it’s still not useful because nobody can get the data out in a way that is meaningful to them. The number of transformations that you need to do is not like what a software engineer would do. A software engineer, a data engineer—their problems are solved by dbt, and I love dbt. You just automate some SQL transforms in the cloud. But what does a scientist need to do? They need to QC data, they need to analyze data, they need to aggregate it and create new views that are meaningful. A manager needs to create operational metrics on top of scientific data. That’s a complicated calculation. We are missing some of the primitives that would allow us to take this software set for technology and make it fit for biopharma. Because if everybody needs to custom-build on top of Snowflake and Fivetran and dbt and AWS or GCP, we can’t do that. We don’t have the money, the time, the energy. We need better tools.
Ross Katz: Like a tooling ecosystem that is meeting the biotech organization at their level of maturity, but leading them in the direction of the industrialization that they’re going to take on later on. So trying to break that tradeoff that we were describing earlier.
Jacob Oppenheim: Exactly. And even if you had infinite resources, a company needs to focus on their core IP. That is never going to be building technological systems when you’re there to build therapeutics. To the extent that you build technology, it needs to be focused around what you’re doing scientifically and the computational pieces of that. Or suddenly you have some entirely new way to construct gene therapies or CRISPRs or something—sure, you’ll invent some data structures. But you shouldn’t have to build everything by default off of a base database facility and the like.
Ross Katz: Yeah, interesting. It just feels like this is a recurring theme that I’ve heard from people on the podcast and people I’ve talked to elsewhere—the modularity of the processes and the tooling ecosystem. There’s this lack of interfaces between them that allow you to build the stack that your company needs without having to roll your own entire ecosystem. Am I thinking about that right?
Jacob Oppenheim: You are. I love the term modularity here because nothing is modular right now and it drives me nuts. You should have modular components that work together well. That’s how you build good software. That’s how you make money in the tech industry. You don’t create something that empire-builds and integrates with nothing. It’s a broken paradigm. Science is too important for us to waste our time fighting with this.
Ross Katz: It’s a really interesting problem because practically everyone agrees with the premise—the entire initiative around FAIR data is that this ecosystem should evolve in this direction. From your perspective, what are the barriers? Why hasn’t this come to be yet? And what is needed to make that a reality? Is it just the development of that operating system, or is it something more?
Jacob Oppenheim: There are a bunch of changes happening in parallel that have allowed this moment to start to arise in a way that we weren’t ready for even 5 to 10 years ago. First, it’s just way easier to build this sort of software right now. You don’t have to come in with perfect diagrams and everything. If you go talk to people in biopharma who’ve been there for a while and understand data deeply, a lot of them will have beaten their heads on the difficulty of designing a database and designing data systems when 10 to 20 years ago cloud storage barely existed, it was expensive, compute was even more expensive, so every time you had to load data into a database or a data warehouse, you were doing transforms beforehand and throwing away data. Now I can just throw everything on the cloud, build a bunch of different views, operationalize between them and keep them all in sync with dbt and tools like this. You couldn’t do that before, and even if you could set up a manual process, it was very expensive to run all of it. So the first thing is that we’ve been able to have systems that better capture the complexity and mutability of what we’re doing in science. A second thing is that biology in the past 20 to 30 years has gone through a complete revolution in moving from basic science only to being an engineering science. What do I mean by this? If you look at drug development up through the 90s and early 2000s, a lot of it was phenotypic screens—as my friend who’s a psychiatrist very involved in drug development says, it was well, we noticed something weird in rats, let’s go add a bunch of methyl groups to it. This is actually how most drugs in psychiatry were developed. This is not a way of doing science that is conducive to storing data. It’s not meaningful to store data because, even if you went and said I want to drug this target, you would only be able to drug the targets that you could bind in the first place—a small fraction of them, that you had probes to test. It was never a science of what should I do, but a science of what could I do. That kind of science is very hard to store data from and reuse. Now, as biology becomes more engineering-focused, we have the Human Genome Project, we have an immense number of discoveries about RNA that happened in the 90s. We have an explosion in tools and technologies that allow us to modify fundamental underlying biology, create new cell lines, create better model systems. Now, if you have a target, it’s not can I hit this target? It’s should I use an antibody, should I use some other type of biologic, should I use small molecule, should I do a gene therapy, should I do an RNAi? There are so many options. Because of the reusability of what we’re doing and the components and pieces of science that can be reused and we can decide exactly what we want to do, data is important. Data is reusable. I can go talk to biologists and design an assay cascade for a totally new therapeutic target to prove it out. Some of those assays I’m going to reuse again later because I understand the whole pathway and I understand what this assay is trying to measure. I’m not going to say this is a perfect analogy and that these systems we understand every detail of them, but we’re enough at that engineering point where data is fundamentally reusable and data can guide what we’re doing—it’s so much more of a what should I do and how question rather than what can I do other than pray.
Ross Katz: In that environment where you’re moving from the ‘what could I do’ to ‘what should I do’ arena, it strikes me that the problem space has the potential to be so much broader, and that discipline of deciding what questions to ask, what hypotheses to test, what experiments to run, and then how to build on the research you’ve done previously and how to organize it is a fundamental component of that environment. Am I thinking about that right?
Jacob Oppenheim: That’s right. And that totally changes how you think about anything computational on top of it as well. I like to tell people that nobody really knows how to run a data science or machine learning team, especially in biopharma. It’s just too new and a lot of the paradigms people have tried are broken. Back in the day you could say, well, we want to hire some data scientists—that’s just going to be another branch of research science and we’ll see what people can do with data. Today that’s crazy. If you have data, data should be informing everything that you do. The data science team is not another team like the biology team or the chemistry team or the pharmacology team or the toxicology team. The data science team’s job is to analyze data and put it in front of people so that they can make better decisions. It’s to productize the computational analyses that are meaningful. Sometimes that’s going to be really complex ML and sometimes that’s just going to be show people the data. But that too has undergone a shift from being a researchy thing to being a predicate for everything that we should be doing.
Ross Katz: It also strikes me that this is why data science and biotech are navigating these integration points—figuring out how scientists and data teams fit together, navigating the cultural clashes that come from that, the differences in how each side thinks about it, and the tooling conflicts about how each side wants to look at the world. Do you have a vision for how that all shakes out? What does a biotech scale-up of the future look like where they’ve got the right tooling, the right data management systems, the right organizational structure to deliver on the promise of everything you’re talking about?
Jacob Oppenheim: Now that’s a tough question. But seriously—in the future, if we have an operating system that allows people to use the data that’s generated and make meaningful decisions, the point of a computational team is to focus on core IP. It’s to work with the scientists, understand the key scientific questions, bring computational expertise to answer them, and then immediately publish that back out to everybody. I’ve found that teams and organizations work best where everybody knows who their customer is. Every team knows who its customer is and is focused on delivering product to them. The biology team’s job might be to run assays. The data science team’s job is to expose data and meaning from data back to other people. The chemistry team might be to design and optimize molecules. If you know what your function is and who you’re there to support, it allows you to industrialize processes, move much more efficiently, and identify the gaps. I’ve seen companies that do this well, where they all are rowing together. It’s a function of what do I need to do, what do we need to serve, what is the goal here? This isn’t about resources or number of people or amount of time with the automation team. It’s what are we all trying to do here, and what are the critical steps that we need to solve? Because science is always going to be hard. Even if we have an idea of what we should be doing, science is always going to be hard. The faster we can make better decisions and have that transparency, the more effective we’re going to be.
Ross Katz: It reminds me of one of your blog posts where you were talking about how we’re moving to a world where we’re letting the requirements drive the decisions rather than the functions we’re onboarding drive the decisions about the data we gather and the technology ecosystem we’re going to operate in. And that the data and technology strategy component in particular in biotech needs to be thought of at the very beginning in order to set the stage for the industrialization that happens later.
Jacob Oppenheim: That’s exactly right. You should know what your requirements are. That’s what a functional group’s job is—to come with requirements, to know what they need and to lay out what those requirements are. Maybe you don’t know them perfectly, but you can sit down with technological experts. Those experts can go out and find vendors, they can prototype what software should look like and what computational tools should look like. Then they should come back to you and sketch out some possibilities. But this idea that we’re coming as functional groups and knowing what the answer is for something outside of our domain is a very inefficient way to work. It’s no surprise that it oftentimes leads to terrible technology, poorly installed, unable to scale.
Ross Katz: Is that one of the roles that Digitalis is playing with the startup companies as well—helping them think through their tooling ecosystem? What are you seeing at the startup level? Is this more a problem of mature companies, or are the startups in the ecosystem making some of the same mistakes and needing to be guided in the right direction?
Jacob Oppenheim: I see this everywhere. A startup company doesn’t have enough time and energy to pick up what the right best practices are. I have a friend who used to run informatics at EQRx and he got together with the people who used to run IT, and they have an IT and informatics consultancy. Their whole point is we’ll set up your IT and we’ll grow with you. When you need a new informatics system, we’re going to tell you how to license it, we’re going to give you the best-in-breed options, and we’re just going to set it up for you. Because we’ll be able to have that conversation—you’re just not going to have the people who say actually these are the benefits and minuses of every compound registry system. It’s fundamentally going to be very difficult for a company, even at moderate to large stage, to do that because then they need to spend their IT team going out and figuring stuff out and discussing things with their chemistry team and going back and forth. There’s a need for technical expertise and working with people as they grow to combine those pieces together. IT and informatics is a classic thing that can actually be pretty successfully outsourced and we have many years of experience seeing this being done well.
Ross Katz: Yeah, that’s interesting. Since you mentioned it, how do you view the role of consultancies in the biotech space? Is there an openness at different tiers of the market to bringing in external expertise, or is there a fear of exposing your core IP as part of that process?
Jacob Oppenheim: Consultancies are very important. The reason is that not everybody who founds a company or has a company is going to have the requisite experience on staff or even know how to build out a team with the requisite experience. You can see this all the time. If you’re a biologist and you do some sequencing-based assays and there’s some core facility in your academic lab, you may have a great idea about how to use sequencing-based assays to accelerate what you’re doing scientifically. You go out, you raise some money, you stick to the same spit-and-bubble-gum pipeline you used in grad school, postdoc, when you’re a professor. Then suddenly it’s time to run 10 times as many samples and modify your assay to do X, Y, and Z. Suddenly everything breaks and you have no idea how to fix it. There’s a liquid market in biopharma in bioinformatic consultancies up here in Cambridge for exactly that reason. That’s great that they exist because it’s not even clear that those companies necessarily need to have deep experience in translating the results off the sequencer—they just need a pipeline that can be set up and run and somebody they can call when it breaks. That’s quite important because if we’re going to succeed in engineering biology, it doesn’t mean that everybody’s going to do everything. It doesn’t mean that every scientist is going to code, which is a lesson that took me many years to learn. It means we’re all going to try to operate at the top of the stack and bring in external expertise where useful. Especially when you look at how biotech has virtualized the running of assays—people use CROs all the time. That’s fundamentally not so different from calling a consultancy to set up pipelines for bioinformatics, or to help pick informatics systems and get them connected and working with everything you’re doing.
Ross Katz: There are a couple of things that jumped out at me about what you said. I think it applies not just to startups but to companies at basically every level of maturity—you as a company need to focus your people on the competitive advantage you have. What you’re trying to accomplish, where your core IP is. There are always tradeoffs between the skills and expertise you bring in-house versus what you have available. There’s always an opportunity to have a consultant layer on top the expertise you’re missing, bring that expertise for a limited engagement and then unplug so that you can operate at scale more efficiently. What do you think about that idea?
Jacob Oppenheim: That’s exactly right—and ideally help you figure out what the team needs to look like and maybe even help you find the leader of it when you pass it off. That’s quite important because not every scientist needs to learn how to code. Even if you have some biologists in your lab who can do some Python or R, does that mean they know how to hire a bioinformatics or comp bio specialist? That’s hard. If you’re wondering why I say this, think about this: if you are a computational scientist, do you know how to pick reagents in the lab? Do you know how to pick somebody who has good bench skills? Do you know how to pick somebody who knows how to design good biology experiments? I frankly don’t. I don’t think anybody does. We’re deluding ourselves if we think we do. What I want to see is pieces coming together and fitting into properly sized and properly scoped roles. Because that’s how we’re going to be able to accelerate what we’re doing.
Ross Katz: That makes a lot of sense. A couple of questions as we wrap up. The first is, I know that you’ve done some pretty advanced machine learning in your time in the biotech space and you’ve overseen teams that are doing the same. What are you most excited about in terms of the new data that’s becoming available or the new modeling capabilities that are becoming available and how that helps to drive the industry forward?
Jacob Oppenheim: The thing I’m most jazzed about has always been imaging in biology. Sequencing is slow and expensive compared to imaging, even imaging in numerous spectral bands all at once. There’s basically no limit on what you can do with imaging. We used to do a lot of hyperspectral imaging back at Indigo to understand plant health, imaging in the IR as well as the visible bands. You could see all sorts of chemical signals pop out of that. Machine learning means you can start taking an immense amount of data and pulling signal out of it. You don’t have to just start with the signal that you know about. Obviously you need there to be real signal that you can detect and something to train off of. But I think one of the things I’m most excited about is the ability to scale computational algorithms and work with imaging data—that has been incredibly important. I’m a big fan of protein design as well. My opinion there is that we’re much better at predicting off of sequence than going sequence to structure and then trying to predict things off of the structure. We’re in a much better place to do a number of interesting things directly from sequence alone, and continuing to further predict and refine structures is an artifact of the amount of data that we have about them, not the importance and meaning thereof.
Ross Katz: Very interesting. And since you’re on the venture side as well, what advice would you have for scientists or entrepreneurs who are looking to start a biotech venture?
Jacob Oppenheim: The thing I’ve learned the most over the past year is that what you need to be clear about is what your science and your computation—if you’re computational, as I assume most of our listeners are—what does it enable in terms of figuring out biology and designing a therapeutic? Those are the two steps that are missed the most. People come into something with I can do this, this is interesting, but then they end up with a whole bunch of data and no idea about how that informs picking a target and building a therapeutic against it. Or even if you have a better method for target selection and finding novel targets, you have to develop therapeutics off of that and your entire platform tells you nothing about how to design those therapeutics. The way I would point people is: figure out the deep implications of your science and your computation, your machine learning. What does it tell you about biology and how does that map into therapeutics design? If you can figure that out, that’s what we need, because right now I don’t see a lot of that especially on the computational side, and that’s going to limit the adoption and overall efficacy of what we’re doing in machine learning for bio.
Ross Katz: Are there any resources you would point people toward to help connect those dots, or is it just reading the academic papers in their domain and coming up that learning curve the old fashioned way?
Jacob Oppenheim: Talk to some drug developers. That’s good. Unfortunately I can’t emphasize anything other than learning the old fashioned way. I love protein design as I said. There are a lot of interesting questions about what we can do computationally on proteins. What are the rules that we’re learning well and what are the rules that we’re not learning well? Because that’s what we need to do if we’re going to actually design therapeutics. Right now that’s a very hard question to answer and it’s frankly a question that’s not really broached in the academic literature, even the translational literature. And that’s a problem. I need to know: if I computationally design a sequence, am I going to get a bunch of alpha helices out? Am I just mimicking my training data? Where can I design, where can I not? That’s going to fundamentally play into biology. So unfortunately that means a lot of reading and a lot of chatting.
Ross Katz: That makes a lot of sense. Are there any other places where people can find you or Digitalis Ventures that you want to point people toward?
Jacob Oppenheim: You can find me on Medium and on Digitalis’s website. My blog is Engineering Biology and on LinkedIn I try to post a decent bit. That’s really where things end up. And thanks for having me, Ross.
Ross Katz: Jacob, thanks for joining us today. It’s been a pleasure and we’ll look forward to connecting down the line.
Jason: And that’s it for this episode of Data in Biotech. If you enjoyed the episode, please subscribe, rate, or leave a review in your podcast platform of choice. See you next time.





