Listen on
Overview
Most biotech labs have years of assay data scattered across hard drives, unlabeled folders, and one-off directories that only one person can find. Two years ago, organizing that data was not worth the cost: the use cases were narrow, the schema differed run to run, and nothing else could consume it. Bio foundation models have flipped both sides of that equation. The cost of organizing experimental data has dropped because LLMs can do most of the discovery and disambiguation work. The value has risen because foundation models train across heterogeneous assays and cellular contexts to make predictions that previously required millions of dollars of dedicated screening data. The data labs treated as throwaway is now the input that determines whether they capture the foundation model lift.
Jesse Johnson is the founder of Merelogic, a software consulting firm that works with biotech and biopharma organizations on data infrastructure and data operations strategy. Jesse spent his early career at Google, where engineers control the data collection function end to end, before moving into biotech, where the biology does what it wants and the bench scientists generate the data. That dual perspective shapes what he tells clients: the fix is rarely a production-grade pipeline or a cloud architecture. It is lightweight, human-readable standard operating procedures, clear handoffs between computational and wet lab teams, and a data strategy designed for the questions the lab does not yet know it will need to ask.
In this episode, host Ross Katz and Jesse cover the bio foundation model landscape (molecular interaction models, virtual cell models, and patient or digital twin models), why virtual cell models are powerful in theory but underused in early discovery, how Tahoe Therapeutics and Lilly’s TuneLab are placing different bets on data as a moat, and why electronic lab notebooks are not going anywhere even in an LLM-augmented research environment.
Key Takeaways
Bio foundation models have flipped the value-cost equation for experimental data
For most biotech labs, organizing scattered assay data was historically not worth the cost: the schema differed run to run, the data was collected to answer one specific question, and nothing else could use it. LLMs and agentic search have driven that cost down because an agent pointed at a directory can do most of the discovery and disambiguation work that used to take a person weeks. At the same time, foundation models trained across heterogeneous assays extract signal from data that would not have been worth modeling in isolation. The data sitting on hard drives is no longer disposable. It is potentially the input that determines whether a lab captures the foundation model lift.
The biggest difference between tech data and biotech data is organizational, not technical
In tech, engineers control the data collection function end to end. In biotech, the bench scientists generate the data, and the biology does what it wants. The fix is rarely a production-grade pipeline. Jesse’s data operations plans are typically a small set of human-readable SOPs that specify naming conventions, folder structures, and the minimum metadata each assay readout needs. The upfront work is the careful analysis of what to collect; the deliverable is instructions simple enough that following them is less effort than ignoring them.
Virtual cell models are powerful in theory but currently used for secondary applications
Virtual cell models predict downstream cellular response from a molecular perturbation. The architecture is now technically capable, but adoption in early discovery is light because using a virtual cell model on a novel compound requires assaying it first to know what it interacts with, which largely defeats the purpose. The traction today is in target identification, biomarker discovery, and patient stratification, where the compound is already known and the question is about genetic interaction or population response.
The defensible value sits at the intersection of proprietary data and a model architecture that can extract signal from it
Models are commoditizing. Open AlphaFold, OpenFold, and an arms race of bio foundation models means the model itself is increasingly not the differentiator. Tahoe Therapeutics is building a company that is effectively a shell around an acquirable single-cell dataset, betting on data as the core asset. Lilly is giving away access to models trained on a billion dollars of screening data through TuneLab, treating model access as currency for partnerships with startups it might acquire. Both bets only make sense if you believe data utility compounds in directions you cannot fully map today.
Consistency is solvable. Ambiguity is not
LLMs have made inconsistency a non-issue. Mixed capitalization, different file formats, CSVs and JSONs in the same folder — none of that matters anymore because an LLM can resolve it. What still kills you is ambiguity: a column name that could mean two different things, a file you cannot tie back to an experiment, tacit knowledge that is loosely documented or not documented at all. The data operations plan needs to hold the line on ambiguity even where it relaxes the line on consistency.
Electronic lab notebooks are not going anywhere
A common prediction is that ELNs’ days are numbered. Jesse’s view: not until automation can run any assay one-off, which is nowhere close. As long as humans with pipettes are running exploratory and one-off assays, the ELN is the system of record. The more realistic shift is around the boundary between ELNs and flexible documentation tools (Notion-style capture, AI-augmented notes); the question is how much schema rigidity to enforce on the high-volume assays and where to leave room for the long tail.
Related: CorrDyn partners with biotech and life sciences companies on data engineering and AI strategy for the assay capture, ELN integrations, and foundation model groundwork that determine whether a lab can act on the value bio foundation models are creating.
Full Transcript
Jason: Welcome to Data in Biotech, a podcast from CorrDyn where we explore how companies leverage data to drive innovation in life sciences. Every two weeks, we sit down with an expert from the world of biotechnology to understand how they’re using data science to solve technical challenges, streamline operations, and further innovation in their business. Here we go.
Ross Katz: Welcome back to Data in Biotech. I’m Ross Katz, principal and data science lead at CorrDyn. Here’s a question biotechs encounter more and more frequently. You’ve got years of assay data scattered across hard drives, unlabeled folders, and random directories that only one person knows about. Is that data worth anything? Two years ago, the answer was probably not. The cost of organizing it was high and there was no obvious payoff. But biofoundation models are starting to change that in a fundamental way. These models can take heterogeneous data across different assays, different modalities, different experimental contexts, and make use of it, which means all that stuff you might have treated as disposable might turn out to be one of your most valuable assets. Jesse Johnson is the founder of Merelogic, a software consulting firm that works directly with biotech labs on exactly this problem. He’s not building foundation models himself; he’s on the other side of the coin, helping the companies that need to use them figure out how to get their data house in order without over-engineering it. That combination of technical depth and client-facing pragmatism gives him a grounded perspective on what works when the rubber meets the road. Jesse, welcome back to the show.
Jesse Johnson: Yeah, thanks for having me. I’m glad to be back.
Ross Katz: Well Jesse, for listeners who didn’t hear you when you were last on Data in Biotech, as you told us then, you went from doing data quality engineering at Google, where the data’s clean and the systems are deterministic, into biotech where the data’s messy and the systems are biological. How did that transition change your fundamental assumptions about what good data infrastructure looks like?
Jesse Johnson: There are some major differences in the way data is collected and managed in biology that it definitely took me longer than it probably should have to realize. The main difference from an organizational perspective is that in the tech world, engineers have complete control over how data is collected — they can build systems to collect it consistently, make sure it’s clean from the very source, and they have end-to-end control. In biology, in biopharma, it’s the bench scientists who are really collecting the data, and so as an engineer you have to get used to the fact that you don’t have that complete control. It’s as much about working with the biologists to make sure they’re collecting things the way that you want them as it is about what you do with it once you get the data. Then if you bring in the fact that in most tech situations data is coming in one data point at a time — streaming, a user clicking on something, whatever that is — in biopharma, or in any biological setting, data typically comes in batches. An experiment gets run, the plate comes out, you get the readout from it, and that means the way you capture it is very different. A lot of the time is on the biology side before you actually get that data, so it comes sporadically. Especially in early discovery and exploratory work, the way the experiment is set up is completely different each time, which may mean different columns in the Excel sheet — not just because people are naming things differently, but because there are different parameters they’re following. You have to get used to the idea that you’re working in small, often one-off batches. Of course there are assays that get run regularly and get productionized and automated, but that’s only half the picture and in some cases less than half.
Ross Katz: That makes a lot of sense. You go from an environment where humans are in control over the data generation function and in control over how they measure it, to an environment where we’re trying to understand the data generation function as we’re trying to design the method of measuring it most effectively.
Jesse Johnson: That’s right. You have both the biology doing what it wants to, and the bench scientists trading off the requirements in the lab with the requirements coming from the downstream data engineers.
Ross Katz: You’ve been writing about biofoundation models a lot recently and there’s a bunch of different flavors — protein language models, multimodal models, clinical data models. For listeners who haven’t been exposed to the full breadth of biofoundation models or maybe haven’t wrapped their heads around the space, can you walk us through the main categories of foundation models that exist in biology right now and what each type is designed to do?
Jesse Johnson: Absolutely. One thing I find really exciting about foundation models is — when I started, not that long ago, eight or ten years ago, doing software engineering in biology, there was this idea of building multi-scale models and systems where you could work at the molecular scale, then the cell scale, then the organism or organ scale, with multiple gradations in between. It felt very abstract, more about marketing than about anything actual. The folks deep in the research had a conceptual idea of what it meant, but as an engineer trying to write the code in between, it was very clear the actual engineering wasn’t there yet. Foundation models have given me a glimpse of what this could actually look like in a technical, implementable, usable form — because transformer models are a very general and powerful infrastructure that allows you to build models that link these different scales.
The smallest scale is the molecular scale — models like AlphaFold or Boltz, or the OpenFold project, which is an open-source version of AlphaFold. These are essentially trying to predict interactions between individual molecules. We think of AlphaFold as predicting protein structures because that’s where it started, but really the only purpose of predicting the structure of a protein is to understand what it binds to — whether that’s small molecules or other peptides. I think of AlphaFold, or that type of model, as really being about predicting binding between different molecules. There have been other results recently showing you can actually go directly — you don’t really need to predict the structures, you can predict binding from these models directly. Take a protein where you only know the sequence, you’ve never found the crystal structure, and you could actually predict which small molecules or other peptides it would bind with, de novo, having never even synthesized it, let alone assayed it.
That’s clearly valuable if you can make these models accurate enough, but predicting individual protein-small molecule binding on its own is important but still of limited use. The next layer is: how do you go from understanding these individual interactions to understanding the cellular response? This starts with RNA-seq — understanding RNA expression levels in a cell. If you introduce a perturbation, whether that’s knocking out a gene, knocking down a gene, or a small molecule that interacts with multiple DNA, RNA, proteins — if you make one of those changes, what are the downstream effects? Classically this would be a Bayesian model where you mechanistically model the signaling pathways. It turns out that way of thinking just isn’t accurate enough to consistently do this, so what you can do instead is basically use a neural network, take all your RNA sequencing data, look at a bunch of different perturbations, and actually predict fairly accurately what those downstream effects will be without understanding the individual mechanisms. More recently these models are also using the same transformer technology as AlphaFold and the molecular models, which means that in theory — and at least in the academic literature in practice — you can link these two together. You take a de novo protein sequence, predict something about it at the molecular level, and take that directly into what’s often called the virtual cell model, which operates at this interactions level.
Then you can keep going — a tissue with multiple cell types, an organ with multiple tissues, a patient, and then the population level. These foundation models allow us to link all of these, at least in theory. In practice we’re not there yet, but there’s definitely progress. I think there’s a future — maybe ten years off, maybe fifty — where you could take a de novo molecule or protein, something that’s never been synthesized let alone assayed, and actually predict what the population response would be. Virtual clinical trials where you run thousands of novel compounds virtually before any of them have been synthesized. Obviously we’re very far off, but the potential impact is huge.
Ross Katz: The way you segment them in terms of the levels at which the models are being used makes a lot of sense. One of the questions that comes to mind is how do you see biotech organizations navigating that landscape in order to select which model or combinations of models they should be using to accomplish their particular task?
Jesse Johnson: We’re in the very early days. There are companies that have built these models and are trying to sell them. Pharma adoption is still very minimal, which makes sense — this is untested technology, and the economics of running clinical trials means there’s no way to know for sure what’s going to work. Adoption of the molecular-level models is starting to pick up, and there’s an immediate application for that: if you can run binding prediction, biopharma knows how to use it. You’re basically doing virtual screening, which is something that’s done in the lab now. It’s easy to say we’ll use a model that narrows down the list of compounds or peptides we’re screening.
When you get to virtual cell models, adoption has been much lighter. The potential there is that once you can get virtual screening to work — start predicting interactions more accurately for novel targets and novel compounds — the virtual cell model becomes very valuable because it gives you a more holistic prediction. Right now, using a virtual cell model on a novel compound is essentially impossible because you need to assay it in order to know what it interacts with. So they’re typically used for secondary tasks like target identification, and a bit more in clinical applications like biomarker identification and patient stratification, where you can understand how genetics interacts with a known compound. I think it’s going to be a while before they get used more in early discovery, which is where I think the most interesting potential is. The higher-level patient models — digital twins — are getting used more in the clinical space where you already know a lot about the compound, you’ve assayed it, maybe you’ve even run a first-in-human.
Ross Katz: It’s interesting that as the technology evolves, part of the diffusion process is figuring out what the new process looks like with the advent of these foundation models, and what capabilities get unlocked by reorienting your processes around them.
Jesse Johnson: Absolutely. If the direction this is going in continues, I think drug discovery could be radically different in the future. It’s just a question of whether that’s in two years or five years or ten years or twenty.
Ross Katz: One of the things you pointed out is that these models change the value equation for data that most biotechs have historically treated as throwaway. Can you unpack that idea? What kind of data are we talking about, and why did it used to be rational to ignore it?
Jesse Johnson: Classically, at least in early discovery, the way data was collected is: you have a specific question you want answered, you design an experiment to answer that one question, you answer it, and you move on. Until recently there was typically no way to use that data for anything else, because it was collected in a fairly narrow context — for that specific question you were trying to answer. At the same time, if you wanted to make that data reusable, you’d have to store it in a strict format. You’d want to make sure the schema was the same across multiple assays, which is very difficult if the protocol is changing between runs. You’d want to name it in a way that someone later would be able to recognize, which is actually legitimately difficult if you don’t know what the use is going to be. It’s not just about biologists being lazy — it’s a hard problem.
Both sides of that value-cost equation have basically switched. On one hand, the cost of organizing data is going way down because we have LLMs, agents where you can point them at a directory and say, just find the data for me. It’s not perfect — if the information isn’t there or it’s ambiguous, it won’t be able to do that — but the bar for what makes data findable and organizable for an LLM-based agent is much lower than it used to be. Otherwise it would take a person weeks or months to go through all that data and organize it. The cost is much lower.
On the other hand, the reason they’re called foundation models is that you can train a core part of the model that makes multiple predictions. If you have five different types of assays that you’ve only run a few times, you can train the model to predict those readouts and train one core model that predicts all five. That core model is now being trained by roughly five times as much information as if you’d trained five separate models. You see this with the molecular interaction models — they can predict structure, but also binding, and other properties. What makes it a foundation model is that core set of weights, used with some additional layers on top to make these other predictions. Data that previously you wouldn’t have had enough of to build a reasonable model can now be combined with all your other assays, and you can actually get value out of it. The cost of making it accessible to the model is much lower, and the potential value is much higher.
Ross Katz: Yeah. Do you have an example of a place where you’ve seen this done?
Jesse Johnson: With AlphaFold you have these multiple inputs — that’s the canonical example. AlphaFold 2 or 3, I forget which, and then the Boltz model took it even further where once you predict on structure and put in some binding data, it can actually predict binding on novel, unseen proteins. That’s a separate endpoint — a different kind of assay. There are also toxicity kinds of predictions and things like that. But that’s the classical one. I know Eli Lilly’s TuneLab started out with individual models for each of these different endpoints — lots of different small molecule endpoints — and my understanding is they’re also planning to build a foundation model that will be able to use all that data across all those different endpoints to build one foundation model.
Ross Katz: If I’m understanding correctly, there are kind of two ways this can go. You have assays that you’ve developed — labeled data — that you can use to fine-tune models that have already brought in a lot of information about your particular biological domain, so you can use transfer learning and fine-tuning to get better in silico predictive accuracy than you would have previously needed thousands or hundreds of thousands of assays to achieve. And then there’s also the situation where if you’re able to use the data you have as inputs to the model, the model backbone itself can fill in the gaps — your assay can give the model enough signal that it will be able to tell you, for example, that this molecule is close to other molecules in molecular space, or that this predictive endpoint that’s already been pre-trained is able to add signal that you wouldn’t have been able to bring from the information you have yourself.
Jesse Johnson: That’s right. To get a little more into the mechanics: the foundation model learns how to assign high-dimensional vectors to every small molecule, if it’s a small molecule model. Say you have three targets that you have screening data for against thousands of compounds. If you train that core model on those three different endpoints, it basically learns where to place the individual molecules in this core model so that the similarities and differences that matter — the features that matter — are well-defined, and the features that don’t matter are basically ignored. It places them in a way where the geometry of those embeddings is meaningful to how the molecules function. So now if you have a new target and run a smaller set of screening to learn a little about it, because the model already knows the relationships between the small molecules it can use that to predict what unseen molecules will do against this new target.
Ross Katz: To zoom out a little bit, you draw a distinction between two types of biotech companies: ones that are built around a single core assay they’re productionizing the data pipeline for and running over and over again, and ones that are more science-driven and running whatever assays the research demands. Can you outline how the foundation model opportunity looks different for those two types of organizations?
Jesse Johnson: I should add that those two types of organizations sometimes it’s about time — early-stage startups are just building that first assay, and they kind of morph from a very discovery-focused exploratory company into more of a core assay company. And even those core assay companies — just because they have one or two productionized assays doesn’t mean they don’t have that long tail of smaller assays they’re still running as secondary exploratory assays.
For companies with a core assay they’re running and building terabytes of data from, it’s relatively easy to build a model around that because you have that core set of data, tons of it. For a smaller company, or a company that doesn’t have that core assay, where you’re doing more exploratory work, the question is: is there a way to get more out of that investment? Because it can still cost millions of dollars to run all those smaller assays. The foundation model gives those companies the opportunity — and I expect this will become more so in the future — to take all those assays they’ve run in an exploratory setting and bootstrap some kind of model that will be able to use that. For the core assay companies, there’s always the question of whether they have enough secondary data to be more than a drop in the bucket. But I think it’s a big opportunity for companies that don’t have a massive, consistent dataset from a core assay.
Ross Katz: That makes a lot of sense, and it also gives more runway for people to explore the space around the assays they’ve already developed. In your Substack you’ve made the argument that the answer for these labs is not more engineering, it’s more SOPs, standard operating procedures. Regardless of which type companies fall into, you’ve argued that the reason data falls apart at a biotech organization or at science-driven labs isn’t technical — it’s that biologists and data scientists fundamentally disagree about whether structured data capture is worth the effort. When you go into a lab and build a data operations plan, how do you bridge that gap?
Jesse Johnson: Sometimes it’s hard to distinguish between when they’re disagreeing and when they just have a different understanding of the situation. It’s very easy, especially for engineers coming from traditional tech who haven’t had much experience working in a biotech lab, to overestimate how much it’s actually possible to automate and standardize. On the wet lab side, it’s often more about not understanding what’s possible, or expecting it’s going to be too difficult, or just not knowing what the expectations are. Especially if you come out of an academic lab, the expectations for what to do were low — so it may not be that they don’t want to do those things or they’re unwilling, it’s more that they just don’t know what’s expected of them.
There’s a cultural dynamic where the digital folks — computational biologists, data scientists, data engineers — don’t feel like they can ask much of the biologist because they know they’re busy and maybe they’ve had experiences that weren’t so positive in the past. They also may not understand the way the lab works well enough to really know what they would ask for. And the bench scientists aren’t being told what the expectations are, so of course they’re going to do things the way they’ve always done them. I know bench scientists who think of the ELN — the electronic lab notebook — as the golden record, so they write everything down in a doc somewhere and only when it’s perfect do they copy and paste it into the ELN. The computational folks and data engineers would say that’s a terrible idea — it should be in there from the beginning so you don’t forget it. But they may not even realize that’s happening; they just know there’s some breakdown somewhere.
There is actually a lot more common ground than people often realize — it’s just a matter of being more deliberate about saying: these are the expectations, this is why we need you to do it that way. In many cases, especially now with LLMs, you can actually make it less work and less cognitive load, less decision fatigue for biologists, just to follow the instructions if you give them. But a lot of organizations are set up in a way where that just doesn’t happen.
Ross Katz: You call what you deliver a data operations plan. Can you walk us through what that contains and, if I’m a scientist in a lab that just signed up for one of these, what changes about day-to-day workflows?
Jesse Johnson: At its core, the working part of the data operations plan is essentially just a collection of docs with instructions: when this type of readout comes off a machine, this is what you name the folder, this is what files go there, this is what information from the ELN we want — or even if it’s just an ID, this is how we connect it to the ELN entry. It’s just very simple instructions. Behind that, it’s built off a fairly careful analysis of what kind of information needs to be collected, what are the assays, what’s the right naming scheme. There’s a fair amount of upfront work to gather requirements, put together a design, collect all that configuration. But the goal of that upfront work is to make those instructions as simple as possible. If you create a reasonable SOP — a process for collecting that data — it can actually be easier for the bench scientist to follow it, because they don’t have to think: what do I name this? What folder should I put it in? How am I going to remember that it’s here?
As an engineer, I often feel a bit sheepish that this is the solution — it’s not running in AWS, it’s in Notion or Word. But in practice, especially for that long tail of assays that aren’t run often enough to productionize and are changing every time, you just can’t write code to handle that. You have to address the human side of things.
Ross Katz: It strikes me that, given we were just talking about biofoundation models, there’s an element of data strategy here too — navigating the trade-offs between schema detail and ambiguity, the level of burden you’re placing on the bench scientist versus their ability to execute. Am I thinking about that right?
Jesse Johnson: Absolutely. It’s very easy to over-engineer these things. As LLMs get better at finding data and searching through it, it’s easy to think you actually need to be more consistent and precise than you actually do. What’s important is that the information gets there and that it’s unambiguous to an LLM, or to a person using an LLM to help organize it. A few years ago, having a column name where sometimes it’s capitalized and sometimes it’s lowercase was horrific — a huge pain to go through and fix in scripts. Nowadays LLMs can figure that out. But if you have a column name that could mean two different things, that’s still a problem — and that’s a problem even if you had months to go through all the data carefully.
Ross Katz: To think through this a little bit: it’s not about knowing in advance exactly what question you need to answer using the data, or even how that data specifically needs to be structured. It’s about maximizing the information content of the data you’re collecting and thinking about how you might want to expose it to agents or LLMs so that it’s accessible at the point when you ask the questions you don’t know you’re going to ask yet. Am I thinking about that right?
Jesse Johnson: Yes, exactly. And moreover, knowing that LLM technology is probably going to keep improving — at least slowly, you could argue we’ve hit a plateau, but it is still improving — you want to design both for what’s realistic today and for what’s likely to be realistic in a year or two.
Ross Katz: Last question before I let you go. The emergence of biofoundation models, companies focusing more deeply on the data they’re collecting, and the availability of LLMs and agentic AI are all changing the dynamics of the biotech software ecosystem pretty deeply. What are your opinions on ELNs in particular, and how do you think the software ecosystem surrounding a growing biotech organization is going to look?
Jesse Johnson: A lot of people have prophesied that ELNs’ days are numbered. If you’re in an automated lab where robots are doing everything, then sure — you don’t need an ELN if everything is automated and it’s all metadata. But I don’t see a future where we completely move away from scientists with pipettes in labs, because there’s always going to be a lot of exploratory work, always going to be assays you don’t run often enough to automate. Maybe at some point automation gets capable enough that it can run any assay one-off, but we’re nowhere close to there now. As long as there are humans in labs with pipettes, the ELN is the tool they need. As long as there are humans with pipettes in labs, ELNs will be a necessary part of that.
Ross Katz: The analog for industries at large is: as long as there are humans picking up phones to dial customers, there will be a CRM, regardless of whether agentic AI can take care of everything. It’s interesting to think about which of these software systems are sticky and why, based on where the value of those systems actually is.
Jesse Johnson: Right. There’s also the question of whether every lab will just build their own ELN from scratch using Claude Code or Cursor or whatever — which I also don’t think will happen, but that is a much longer discussion.
Ross Katz: I’ve seen some labs I respect saying we’re just using Notion now. The extreme flexibility of a back-office software platform serving as your documentation record, being more comfortable with everything being a little unstructured — versus having an understanding of your workflow — it’s going to be interesting to watch people continue to navigate that schema rigidity versus free-flowing tension.
Jesse Johnson: Yeah, maybe that’s more likely to happen than ELNs going away completely. It’ll be interesting to watch that space.
Ross Katz: Jesse, thank you so much for coming today.
Jesse Johnson: Thanks so much for having me. This was fun.
Ross Katz: Jesse, thanks for coming back on. I want to pull on a couple of threads from this conversation. The first thread is that your data strategy has to account for the fact that the value of your data is going to keep changing in directions you can’t predict. As we discussed, the data you collect to answer one specific question gets repositioned by foundation models to answer questions you never thought to ask. Tahoe Therapeutics, an example Jesse mentioned, is building an entire company that’s essentially a shell around a single-cell data set, betting that the data is the core asset. As we discussed in a previous episode, Lilly spent a billion dollars generating screening data and is giving away access to models trained on it, partially as currency to build relationships with startups they might acquire. Two completely different bets on how data creates value, and both only make sense if you believe the utility of that data is going to compound in ways that nobody can fully map today. And the second thread is the distinction between consistency and ambiguity, which sounds small but changes how you plan. LLMs have made inconsistency a non-issue. Things like mixed capitalization, different file formats, CSVs and JSONs in the same folder, that doesn’t matter anymore. What still kills you is ambiguity. A column name that could mean two different things, a file that you can’t tie back to an experiment, or tacit knowledge that is loosely documented or not documented at all. That’s the line your data operations plan needs to hold. And based on what Jesse is proposing, it’s much more achievable than many realize. You can find Jesse on LinkedIn and on his Substack, Scaling Biotech, where he writes about this every week. If you know someone at a biotech sitting on scattered assay data wondering whether it’s worth the effort, send them this episode. And please subscribe to Data in Biotech wherever you listen, and we’ll see you next time.
Jason: And that’s it for this episode of Data in Biotech. If you enjoyed the episode, please subscribe, rate, or leave a review in your podcast platform of choice. See you next time.






