Listen on
Overview
Host Ross Katz speaks with Jesse Johnson, Founder of Merelogic. Many early-stage biotech startups find themselves at a critical juncture: their scientific advancements are rapidly generating data, but without a coherent strategy for managing it, critical decisions stall. This isn’t just about technical infrastructure; it’s about the fundamental tension between a biologist’s process-driven work and a data team’s need for structured information, hindering everything from R&D efficiency to investor confidence.
Jesse Johnson, founder of Merelogic, brings a unique perspective to this challenge. With a background spanning academic mathematics, software engineering at Google, and hands-on data science roles in biotech startups, Jesse now helps early-stage companies build strong data foundations from the ground up. He understands both the scientific imperative and the engineering rigor required.
This conversation explores how to bridge the cultural gaps between lab and data teams, the specific data infrastructure needs of biotechs at different growth stages, and the often-overlooked role of lab automation in preparing for the true promise of AI. Jesse offers a clear diagnostic arc for how companies can move beyond reactive data cleanup to proactive data management that truly accelerates discovery and development.
Key Takeaways
The Core Data Problem in Biotech: Process vs. Schema
Biologists naturally view each experiment as a unique, complex process with many exceptions, while data teams seek standardized templates and clean data schemas for analysis. This fundamental difference in perspective creates friction and makes it challenging to consolidate experimental data into a unified, actionable view needed for strategic decisions. Bridging this cultural divide is more critical than just implementing new tools.
Scale Dictates Data Infrastructure Investment for Biotechs
Early-stage biotechs often do not need expensive, heavyweight data warehouse solutions. For limited, often self-contained datasets, simple file-based storage or in-memory systems on laptops are sufficient. Significant investment in scalable databases becomes necessary only when data volume and complexity increase, typically around Series B funding rounds, to support more frequent updates and cross-assay analysis.
Lab Automation Is the Gateway to Effective Biotech AI/ML
While AI and machine learning hold immense promise for biotech, their true impact relies on consistent, clean data. Lab automation, by standardizing experimental execution and automatically logging detailed metadata, is a prerequisite. Manual lab work inherently produces messy, inconsistent data, limiting the effectiveness of even advanced ML models, much like the ImageNet revolution depended on vast, labeled datasets.
A Digital Twin of the Lab Drives Operational Visibility
A ‘digital twin’ for the lab environment means a complete, queryable database of every experiment, from initial concept to final data output. This provides data teams with proactive insight into data availability and enables real-time operational visibility. Implementing such a system requires designing software that incentivizes bench scientists to record steps by making their work easier or enabling automation.
Related: CorrDyn provides data assessments and data engineering expertise for clients in the biotech and life sciences industries. Learn more about how biotech manufacturers gain data value.
Full Transcript
Jason: Welcome to Data in Biotech, a podcast from CorrDyn where we explore how companies leverage data to drive innovation in life sciences. Every two weeks, we sit down with an expert from the world of biotechnology to understand how they’re using data science to solve technical challenges, streamline operations, and further innovation in their business. This week, we’re excited to be joined by Jesse Johnson, founder of Merelogic, an independent consulting firm that helps early-stage biotech startups get their data under control so they can focus on the science. During this interview, Jesse highlights the unique challenges of working with biotech data, introduces the concept of a digital twin for the lab environment to better integrate data entry and experiment tracking, discusses the potential of machine learning to enhance efficiency in biotech research, and speculates on the role of large language models like GPT-3 in biotech as potential learning tools to assist in generating experiment protocols. Here we go.
Ross Katz: Jesse Johnson, welcome to the Data in Biotech podcast.
Jesse Johnson: Thanks for having me. I’m excited to be here.
Ross Katz: Awesome. Just to kick us off, can you give everyone a brief introduction to you and your background?
Jesse Johnson: Yeah, I started off my career as an academic mathematician. Did that for probably longer than I should have, and then transitioned to software engineering. I was a software engineer at Google for a couple of years, and then basically by following people that I worked with who had moved on to interesting other jobs, I followed them into the biotech world. My first biotech startup I joined about four years ago, a company called Cellarity. After that, spent a couple of years at Dewpoint Therapeutics. Both roles were data science and engineering, so mostly software engineering, a little bit of data science. Just in the last few months, I started as an independent consultant. My goal is to help biotech startups get their data under control so they can focus on the science. That’s the tagline.
Ross Katz: Amazing. What was it like transitioning into biotech from your deep math and software engineering background? I imagine there was some culture changes and also a lot of terminology getting thrown around that it took a while to get up to speed on.
Jesse Johnson: Yes, definitely. The vocabulary takes a while. You have to be willing to ask stupid questions, which I’m willing to do. Not everyone always appreciates that. If you can find the people who are okay with that, you’re good. The culture—I really liked the tech software engineering culture of leading by motivation and explaining why. In biotech, you get some of that. There’s definitely leaders like that. There’s also folks who have come from more academic labs where it’s more about pumping as much as you can out of grad students and postdocs. There’s a little bit of a culture clash of how much freedom do you want to give to people, how much do you need to tell them to just make them get things done. It’s been interesting, but it’s an adjustment.
Ross Katz: You made that transition several years ago now, so I feel like you’re fully indoctrinated into it. But another question I had based on your background was just that you’ve been in both the bigger company and bigger pharma environment, but also smaller biotech and startup ecosystem. Why did you decide to focus on data problems of biotech startups?
Jesse Johnson: A couple of reasons. The primary one is I really like that early-stage research where it’s very exploratory. There’s a lot of flexibility. You’re not weighed down by regulatory issues—the closer you get to clinical trials, the more of that there is, and the more focused you get, so less room for exploration. The biotech startups, especially those early-stage ones, are pretty much just doing research. That felt like the right place to do it. That seemed like the right opportunity there. I also like coming into a messy room and cleaning it up to where you can see the before and after picture, versus coming into an already mostly clean room and cleaning it up some more. The big pharma, they tend to have a lot of things in place where either they’re good and it just needs to be maintained, or it’s a mess and there are deep structural issues why it’s always going to be a mess. I like being able to come in early on and get from zero to one, as they say.
Ross Katz: More problems of data exploration and getting insights from the data, but also less existing infrastructure to compete with in trying to get the problem solved. That makes a lot of sense. When you’re starting to work with a biotech startup, what do they look like when you come in? When do they realize that they need you? How does that start?
Jesse Johnson: My thinking on that has actually evolved in these months since I started. My primary focus has been very early-stage startups where they’re not yet ready to hire that VP of Engineering or even maybe a CTO. My goal is to come in and help them build something that’s very lightweight, low overhead, and will get them to until they start getting that burst of data and they need to suddenly scale. I want them to be in a good shape to bring in a larger, more high-overhead tool that will get them through the next step. Often, that happens around the stage where most biotechs get to where they say, okay, we’ve now done three or four assays on our compounds of interest. We want to see them all in one table, and for some reason we can’t. It seems like it should be easy. That’s where I can come in. The other type of project that I’ve seen more of as I’ve had more conversations is even platform-based companies—or farther on where they actually have an engineering team. They’ve put some effort into this. Often those engineering teams get so bogged down in the day-to-day just trying to put out fires, keep up with the moving train, that there’s often projects that they really want to do. They know how to do. They just don’t have the bandwidth to do it. There, I can often come in and do a strategic project that frees them up a little bit, just do enough of the project to where they can then take it over and run with it. I’m exploring both of those types of projects right now.
Ross Katz: Do the startups that you’re working with typically have a data team in place when you come in, or are you helping them stand up their new data team that they’re introducing to the organization?
Jesse Johnson: For those early-stage startups, they often will have computational biologists, data scientists, folks who are really good at the analysis side of things and know how to work with data, but don’t necessarily know that engineering side of it, the platform below that analysis. They can do things on their laptop, maybe they can upload things to an S3 bucket, but there’s a layer below of how do you organize that data? How do you build the infrastructure that makes sure that the next time someone comes in, or someone leaves, that they can find the data, they can re-run it? Making it reproducible and all of that. Making the data more of an asset that is valued by the company versus just an operational thing that just gets you to the next analysis.
Ross Katz: That makes sense. I imagine that there’s this aha moment where they encounter some feeling of pain where they’re just like, oh no, we’re not going to be able to continue operating this way. We need to find someone who can help us overcome this. What do those problems look like where the teams start to realize that we need somebody like Jesse in here to help us get to the next phase of what our data maturity looks like?
Jesse Johnson: My goal is to get in before it gets to that level of pain. But often one thing that most startups end up with is they want to see a single dashboard, a single view, a single table of all of their assets, whether it’s compounds or biologics in each row, and then each column is an assay. Basically a decision-making tool that allows them to say, okay, which of these, if you look across all the data we’ve generated, which are the ones that we want to invest more in, what do we want to take to the next assay, next step? That seems to be a universal need. Often that’s something that’s surprisingly difficult to actually build internally.
Ross Katz: There’s that portfolio prioritization problem that’s at a level above the biology that’s being done, and so the data that they have is in all these different places and that unified view is hard to get to. What makes that so hard?
Jesse Johnson: Often if you’re doing analysis on a single assay, you’re thinking about it in terms of, okay, I just want to get to an answer of what did this assay determine. Merging all of that together is surprisingly hard, especially if you’re using different formats for different assays that work well when you look at that one, but not the other. Often data ends up either on people’s hard drives or it’s in a folder for that experiment where if you know about that one experiment, you know where to find it. But if you’re looking across multiple experiments, multiple assays, you don’t know. There’s a certain amount of forethought that you need to put into, how do we want to organize this in a way that someone who isn’t deep in the science can come in and understand across all the different experiments and assays what’s going on.
Ross Katz: I’m assuming that these assays are happening on different pieces of equipment, right? The data that they’re generating is coming off of the equipment in different formats and so munging it together is hard.
Jesse Johnson: Typically you’ll have a primary assay which is telling you, does this work, does this do the unique thing that we care about? It could be a binding assay, is a classical, it could be a phenotypic screen or some other sequencing-based assay. Then you’ll often have secondary screens where you say, okay, we know that this does the thing we want, but is it toxic? Can we run another screen that validates that it actually did what we wanted? These are going to be coming out of different instruments, using different biological models. They often end up in different formats on different people’s laptops or special folders.
Ross Katz: The workflow on each of these pieces of equipment that drives the assay has a GUI that allows you to export your one experiment, and so experiment by experiment the biologists are doing this analysis, but then you’ve got this data team that wants to look at things at a higher level, and also a leadership team that wants to look at things at a higher level.
Jesse Johnson: Yep, exactly.
Ross Katz: Interesting. Starting out from that place and getting to a place where you have these repeatable data processes that are producing consistent data without getting in the way of the biologists that are trying to do the real important research work of the organization is a fundamental trade-off that I’ve seen you write about before. How do you navigate that?
Jesse Johnson: The way I’ve written about that has actually evolved as I’ve gone from working with Series B and C startups to spending more time talking to very early seed-stage. If you’re a Series B, C startup and your goal is to scale and you want to basically run as many assays as you can as quickly as possible, you really need to make sure that’s done in a consistent way so that you’re not spending extra time going back and cleaning things up. If you’re at a very early stage and you’ve maybe run that assay two or three times over the course of many months, putting in those processes is actually more investment than just cleaning them up as you go along. It does really depend on the rate at which you’re collecting data and the rate at which you expect to collect data in the future. The important part is that you have a standard that where it ends up is in a clean, consistent place. If it takes a lot of manual work to get there, that’s fine. Once you have that in place and you start to get to a point where that cleanup is starting to be a pain, that’s when you start thinking about, okay, how do we automate that? How do we go upstream and make sure that the scientists are running the assay in a way that it’s easy for us to get it into the right form? And then hopefully by then you’ve built up enough goodwill with the bench scientists that they’re willing to make those changes because they can see the value downstream.
Ross Katz: I’m imagining that there’s also what’s the level of organizational urgency around this particular problem versus the problem of just finding the targets that you want to go after as an organization. Is it the idea that there’s a funding round that the startup needs to go after that’s driving the urgency around bringing the data together, or where does that come from?
Jesse Johnson: It does coincide, to some extent, with the funding rounds. Depending on the market, the expectations at each round may change. Maybe in six months, you’ll have to adjust the expectations accordingly. But typically, if you’re in that seed stage where you don’t really know if the science works yet, you don’t need to put too much thought into this because either it’s not going to work and you’re going to throw everything away and move on to the next idea, or if it does work, you’re going to want to validate all of those experiments with better assays. You’re going to throw away the data anyway. At that point, it’s fine to be a little bit more disorganized. It’s really when you get to Series B that most companies, they want to ramp up screening, ramp up their experiments. Now it’s really not worth putting that time in and it’s better to invest in an automated pipeline.
Ross Katz: This conversation takes me back to what you were saying earlier, that ideally you would be there in the room before all of these decisions that have organizational momentum are made, because you can help organizations to do a lot more with a lot less if you’re there at the beginning to help them design how they’re going to collect and organize the data, versus having all of the data lying around and having to collect and organize it and automate it later on.
Jesse Johnson: That’s right. The longer you wait, the harder it is to go back and clean it up, as with many things. If you wake up one morning and realize that you need to raise another round, you’re in a much more difficult position if you didn’t do that strategic work up front. Both when you start building the data room for your investors, but also there’s more and more VCs who are going to start asking about thinking of your data as an asset.
Ross Katz: That makes a lot of sense. I know that one of the things you focus on is the way that cultural components and the different stakeholder groups that you have in biotech organizations contribute to the systems and process problems that we see. Can you just talk a little bit about what you see on the ground in biotech organizations?
Jesse Johnson: One of the things that I’ve found most interesting is when I come in, I come in from living in the physical world where there are rules of physics, and they almost always work. If they don’t work, then you get really surprised. Biologists typically live in the world of biology where the rules are really more suggestions. The exceptions are more common than the things that follow the rules. When I think of an experiment protocol, my inclination is to think of it as a series of steps that you’ve swapped things in and out and they’re all going to work and it’s just going to work. Whereas biologists, when they see an experimental protocol, they think step one isn’t going to work, and here’s three reasons why. If you’re going to make changes to this protocol, well, you’re going to have to make all these other changes to the steps for that cell line, for that compound, for whatever it is to make that work. Biologists tend to think in terms of exceptions and they see each experiment as a unique snowflake, because it is, whereas the data folks tend to assume that it’s all just templates that you can swap things in and out of. That is at the root of a lot of the clashes, the way that people communicate. It’s very rare to look into your motivations and your underlying assumptions to the point where you can dig down to those differences and often you may not even realize that those differences exist until you come back two weeks later and you realize that the thing that you thought you had agreed to, they thought you had agreed to something completely different, even though you both agreed to it at the same time. The more you have these conversations, you can find ways of teasing out those differences and change your expectations accordingly.
Ross Katz: Interesting. From the data team’s perspective, they’re just trying to get at the metadata. Let’s have a set a bunch of columns that we’re going to fill out for every experiment, and every single experiment is going to have these columns filled out with consistent values that we can use to build our models, do these analysis, produce these tables. From the biologist’s perspective, the columns—the fields—are changing every single experiment, and also you’re constraining their creativity about the experimentation process by forcing them into this. Is that right?
Jesse Johnson: No, exactly. The biologist’s view of an experiment is this process view of, here are these very detailed steps that we went through in order to make the experiment work. Someone who knows how to interpret that can extract those columns, those values, that the data scientist needs. But the data scientist isn’t necessarily going to feel confident doing that, and the biologist, they’ve gone through that process, they’ve gotten what they needed out of it, they don’t necessarily want to go back and pull that out again because for them it may not be worth their time. You end up with this translation problem of this detailed process view into that column view that you described. Right now, the tools available don’t make that easy to do, and it’s unclear if there actually will be that tool.
Ross Katz: That tool that spans all of the biotech R&D use cases and all of the data uses is a high bar to ask for. It strikes me that the things that make biologists good at their work, which is the ability to follow these complex path-dependent processes through convoluted and hard-to-discern patterns, also make it difficult for them to see the world from the data team’s perspective. But the data team’s desire to organize everything and look at everything globally and just have everything fit neatly into the way that it’s supposed to work for the system to produce the result that you want—it’s fundamental to the type of work that these people do.
Jesse Johnson: Yes, absolutely. That’s right. It’s easy when you’re in the midst of it to think, oh, those guys are wrong, whichever side you’re on, and get very frustrated. But ultimately, you hit the nail on the head. It’s fundamental to the sort of personality, the sort of mind-frame you need to have to do your individual job is very different depending on which side you’re on.
Ross Katz: You mentioned the idea that a tool can solve this problem. Also I think that in the tech world, for the technologically inclined, we tend to lead toward what are the tools that we can bring in that resolve these tensions for us without us having to deal with the real people problems that exist on the ground. To what extent is the data tool ecosystem enabling the process, and to what extent is the specialization of these data tools in biotech hindering the process of getting these biotechs up to scale?
Jesse Johnson: Great question. There’s been a lot of startups recently that have been founded. The software has evolved from 10 years ago. A lot of what was important was ELNs and LIMSs types of systems that were solving the problems that you had today. Now as we’re getting into more AI/ML use cases, the types of tools that we need are evolving. We’re seeing a lot of startups that are trying to build these tools. One tendency is to try and make them broad and have a lot of functionality because it’s very difficult right now to put together different modular pieces, which is the idea, the dream that everyone has, but those modular pieces don’t exist. There isn’t that standard yet. A lot of the startups tend to try to build the everything platform, which from one perspective makes sense because that way the individual startups don’t have to worry about integration. But in practice, often the components that they need is just some fraction of that. It ends up unintentionally creating this sort of lock-in which—I think their goal is not to lock in. They’re trying to make it so that they’re doing good for everyone. I think they’re doing that as well as they can. We’re not yet to a point where we have these modular pieces that we can put together. I’m hoping that we can continue to evolve there just so that there’s less overhead involved in bringing in a big solution and it’s easier to iteratively bring in small solutions that add to your overall solution.
Ross Katz: This conversation takes me back to what you were saying earlier, that ideally you would be there in the room before all of these decisions that have organizational momentum are made, because you can help organizations to do a lot more with a lot less if you’re there at the beginning to help them design how they’re going to collect and organize the data, versus having all of the data lying around and having to collect and organize it and automate it later on. From my perspective, there’s this basket of modular tools that have been put together in consistent patterns across industries. You’ve got your storage layer where everybody’s landing their data in S3 or Google Cloud storage, and then you’ve got your compute layer where maybe it’s a data warehouse, maybe it’s something like Databricks, or maybe it’s more of a streaming-type system. And then you’ve got your output layer where you’re doing visualization and/or driving value from the data. I’m just wondering, is the biotech ecosystem recreating these tools for biotech? And where are those tools that are at scale in the data ecosystem falling short of the use cases in biotech?
Jesse Johnson: A lot of the components from the modern tech stack make a lot of sense. Certainly cloud storage is generic enough that you don’t need to make a specific version of that. You could argue that something like Egnyte where there are some features of that are designed for sharing data with CROs, that can make sense. But it’s more on the periphery. There are a lot of components from the modern tech stack that make sense. There’s also specific aspects of biotech data, particularly the early types of assays that startups tend to run, that are different from if you’re making another social network or an e-commerce site. Things like Snowflake where it’s an at-scale relational database, I think that often makes less sense for these early-stage assays where there’s fewer references between tables. It tends to be more each table is self-contained, relatively small usually, small enough that you can either put it in a relational database if you really want to or just leave it as a file in your cloud storage somewhere. A lot of these heavyweight solutions that have become the standard in the modern tech stack don’t make as much sense. On the other hand, usually the things that make these tools specific to biotech are on the periphery. If you have a visualization tool that can automatically draw compound diagrams from a SMILES string, that’s huge. But it’s small, probably someone could add it in a day, but for the end user, being able to do that versus having a generic BI tool that can’t do that is a big deal. Then you have other software that doesn’t actually need to be specific to biotech, but for commercial reasons it makes sense, where Plotly, which is another great tool that started off in the tech space, realized that most of the adoption was in bio and has doubled down on that and started to really build the features that their biotech customers are looking for, which again may not be specific to bio, but it works, and turned out to be a really good strategy for them.
Ross Katz: Interesting. There’s a variety of things in there that I would just want to unpack. There’s the file format issue up front, the formats of these data are not formats that are commonly used across industries. Then there’s also the computational processes that are being done over those file formats are atypical computational processes, so they’re not typically built in to a lot of the at-scale data warehouses, for example, are not going to allow you to do complex genomics analysis within that environment. Right, right. The last thing I heard was at the output layer you’ve got data visualizations that are unique to the type of analysis that’s being done. There’s this neat connection between each slice of file format, computational approach, and visual output that there’s just lots of these that get stacked together. Is that right?
Jesse Johnson: That’s fair. Let’s start with another—the contrast is if you’re building a social network or e-commerce site, a lot of the times you have tables that need to connect to each other. If it’s people versus topics that they care about versus their connections. On the one hand, if you want to scale to Facebook scale, you need to do that for billions of people. You need to know all these connections and so on the one hand, you have both the connections between tables and then the scale. If you’re looking at especially early-stage biotech startups, the data that you can generate is fairly limited. If you’re doing sequencing maybe you can do hundreds of thousands or millions, but typically once you get down to compounds that we care about, it’s going to be in the hundreds, maybe in the thousands. It’s at a scale where you don’t need a huge hammer to manage data at that scale. It can fit in memory in your laptop. Even for what I’ve worked with like single-cell data sets before, even those data sets they can fit in memory in your laptop. Having a large-scale solution that’s going to work for billions of users, billions of data points for a social network is just way overkill for that. Each data set tends to be its compounds, maybe cell line, and then numbers. The structure is very simple. You don’t have these complex connections between tables. Once you get into clinical real-world evidence then you start to get that. But if you’re talking about early-stage research, it tends to be much simpler. Having a table saved in a file somewhere in your cloud storage is—especially when it’s a handful of developers, handful of data scientists, that’s usually sufficient. It’s only once you get into much more data, much more complex data that you start to have to worry about these other solutions. Though even then, these massive cloud warehouses, data warehouses, are often still overkill.
Ross Katz: What does that architectural pattern look like then, when you’re moving from doing this exploratory research where you’ve got a team of scientists who are working on a laptop with files that all fit in memory, to you’ve got these repeatable analyses that you want to do over and over again? Is it basically just Dockerize the thing that you were doing on your computer and then run it over a bunch of files in cloud storage, or how does it go?
Jesse Johnson: Certainly that can work up to a certain scale. At some point you’ll need to introduce some kind of proper database. Usually, that’s when you want to start looking at assays with each other. Files work if you’re just updating it a few times a week or maybe even every few weeks, you can just rewrite the whole thing. Once you start getting enough updates that you don’t want to blow up and re-create your table every time you do the analysis, that’s when a database allows you to do that incrementally. Then it’s worth that investment. But depending on how fast you’re moving. Of course, if you’re a platform company, the calculus is very different because you know you’re going to get to that stage, and so you might start building that sooner. That’s roughly the calculus.
Ross Katz: Awesome. I want to play a little game of Fact or Fiction. This is the first time we’ve ever done this on the Data in Biotech podcast. I’m just going to throw some ideas at you that I hear in the biotech data ecosystem and get your takes on them, if that’s all right. The first one is that AI and software are going to automate most biotech R&D processes within 10 years. Is that fact or fiction?
Jesse Johnson: I try not to make predictions about the future, or at least not publicly. What’s going to come before the AI and data is going to be automation. People who know more about automation than me have predicted that most of the labs are going to be automated in five years, which maybe that’s optimistic, certainly they would probably say in 10 years. Once you have lab automation, it’s much easier to automate the data because this metadata that you mentioned earlier, if a robot is picking up the plates and all of that, you can extract that all from the logs, whereas it’s much harder to extract that from the mind of a person who sat there with a pipette. If lab automation is the standard in 10 years, then getting that relatively clean data into an automated pipeline is much easier, and so I think it greatly increases the chances of that automation.
Ross Katz: The path to the data ecosystem that we’re looking for is a path of automating a lot of the research processes up front so that the data flows naturally in a format that is easily consumable by the data systems the data teams will tend to manage. Okay, interesting. Data is more important than machine learning models in biotech.
Jesse Johnson: I think that’s true. If you look back at machine vision, image recognition, the convolutional neural networks that are now the standard for doing those, or at least maybe they’ve been replaced now, but for a while they were the standard. Those existed since the 80s, at least on paper, but there wasn’t enough data to actually make them more effective than other methods. It was only when you got ImageNet with millions of images that those models that had been around for decades suddenly started out-performing other models. I’ve seen this a lot with biological data sets, which are currently much more constrained than the kinds of data sets in the tech world, where you basically get to a point where even very simple machine learning models do as well as the very sophisticated ones, sometimes even better, because you’ve reached a constraint on the data. In general, models are fun, but if you don’t have the data, they don’t work.
Ross Katz: The next one, fact or fiction: biotech is a useful industry descriptor for vendors that are building software and data apps to target as a unit.
Jesse Johnson: I think that’s true because even if the software is underlying mostly the same, for a vendor to understand the problems and understand the terminology well enough to be able to tell you how their solution fits in, that’s really important. The software itself may not be that different, but it is a valuable distinction.
Ross Katz: Interesting. What made me think of that one was the idea that a lot of the work that these biotech organizations are doing is very different depending on the types of biological processes that they’re working with, the types of assays that they’re using, the biological specializations that they’re operating in. But it sounds like you, based on your experience, view the problems of these very different biotech organizations as being relatively consistent.
Jesse Johnson: On that side, I guess if you’re comparing it to general software for any industry, it’s a more useful category than industry in general. Once you get into individual biotechs, you find that there actually is a lot of heterogeneity between what different biotechs need, which is a lot of the reason that it’s very hard to build a tool that serves a lot of different biotechs and is very specific. Every company that’s building software for biotech has to make that decision of: do we go deep and build something that’s really useful for a very small number of potential clients, or do we go wide and build something that’s slightly better than generic?
Ross Katz: That makes a lot of sense. There’s the micro-targeting approach where you get deep with a particular segment within biotech or within a particular problem space, but then there’s also the broad approach that you can take as well. The last one we’ll play for fact and fiction is—fact or fiction: hardware and software vendors in biotech are more interested in locking data in than enabling analysis of it.
Jesse Johnson: I have to be careful because a lot of these are folks that I want to work with. In my experience, it’s a mix. Hardware vendors are more likely to try and lock in data because investors are telling them that hardware is not a great investment, and if they can build software and become a software play, then their valuation will go up. In order to become a software play, the best way to do that is to lock your clients into your software system. I’m not going to name names, but I do think that is a motivation for hardware companies. For software companies, it’s much less common and in fact, I can’t think of any software companies that are intentionally trying to lock people in. Especially the newer companies are really great about saying, we’re going to use whatever, we’ll at least give you the data in a way that you can pull it out of the system or we’ll let you use it in your own cloud wherever you want it. That’s because they first of all don’t want to have customers be afraid of lock-in, they know that they’ll get more customers by telling them they can go any time, and keep them longer because if they tell them they can go any time. They want to—they’re already doing that software play, so if they can give more value, then they think they can do best that way. I think some hardware companies, not all hardware companies, maybe one or two software companies, but not most of them.
Ross Katz: Awesome. Well, thank you for experimenting with fact or fiction with me. I enjoyed your responses. I noticed that you started on a project called Foam for data apps in biotech. Could you just talk a little bit about that project and why you started it up?
Jesse Johnson: This was part of this effort that I mentioned earlier, of trying to find a way to make more modular components for building data and apps. One of the things that I found was that if you want to build analysis pipelines, that’s pretty easy. But when it comes to collecting data from users, that’s where it gets tricky because ELNs and LIMSs are not really set up for AI and ML applications. If I just want to let users enter sample IDs and give me some metadata about the samples, there isn’t really an easy way to just spin up that app and have that be integrated. The idea of Foam is it’s what I call low-code for coders. There are a lot of these low-code app development frameworks where you go in with a mouse and draw the windows that you want and then you design the underlying schema and then you click go and you have that app. That’s great as a way of quickly spinning up these apps, but the problem is that if you’re a coder, it’s awful because you want to work with your keyboard, you want to be able to copy and paste. If you’ve got to do the same thing dozens of times, you want to be able to use your fancy IDE magic to do that. The idea of Foam is that it’s similarly configurable except that it’s all based on YAML files. YAML is a version of JSON that lets you write these structured configs. The idea is in a text editor you very concisely define what you want the backend schema to look like and what you want the different windows and the workflow to look like. Then in a CLI, you say compile, and now it generates the code using Django for the backend and React for the frontend, so all fairly standard open source tools. That code you can run locally, you can put it in a Docker, you can push it to GitHub, you can do whatever you want. It’s meant to be very developer-friendly but allow developers to very quickly spin up these frontend applications, maybe even if you’re a backend developer and don’t want to write React code, you can write one of these configs and get your app going. It’s in very early stages, it’s a lot of cardboard and duct tape, it’s at a point where it was enough to convince me that it works and I’m planning to put a lot of energy over the next few months, maybe years, to get it to a good place. If anyone wants to try it out and see if it works for your startup, send me an email, I’d be happy to, for free, jump in and give you some advice and help you work through any issues.
Ross Katz: Is the ideal application for Foam a data entry application that keeps track of the metadata as we were talking about earlier, or what do you expect to be the end user problems that you would solve with apps that are built using Foam?
Jesse Johnson: Great question. I’m probably being very ambitious, but I think it can start off as data entry. It’s designed in a way where you can start building tables that you can use to get insights into what’s going on and maybe even share those with leadership. It’s anything that you can build with data entry forms and then tables. If you want to build custom React code to build very custom views, it also allows you to insert that. I think of it as filling in—the reason it’s called Foam is it’s meant to fill in the cracks between the other pieces of software that you have. Ultimately I’d like it to get to a point where you could build the core data management infrastructure for your whole organization. Maybe it someday could replace the ELN, maybe replace the LIMS—definitely not today. But I’d like it to be able to do anything that involves either entering, finding, or looking up data.
Ross Katz: It’s just interesting to me that you took this approach to solving that problem based on your experience with biotech startups, because it strikes me that it’s an approach that has a lot of respect for the diversity of issues in the data ecosystem from company to company and the diversity of data sources and interfaces that companies are working with. Foam is filling in the gaps of whatever the other tools are that are on the table. Think about that right?
Jesse Johnson: Yes, that’s the goal. It’s still very much an experiment. I don’t know if this will actually work, but I figured it doesn’t hurt to try, and it’s something that I’d been thinking about and had built other versions of enough times that I knew at some point I was going to need it again, so might as well build it open source.
Ross Katz: I think the idea that this is the solution that came to you from your experience is as interesting to me as the actual code that’s written to solve the problem in this way, using Django and React. So that’s cool. An idea that I’ve heard you tossing around is this idea of a digital twin for the lab environment. What does a digital twin of the lab environment mean to you?
Jesse Johnson: For me what it means is I want to know where everything in the lab is at any given time. It doesn’t just mean in terms of inventory, but also in terms of: what are the experiments that people are in very early stage discussions about all the way through the ones that are in the lab or the ones that have come out already. I want to be able to look and see for any question I have about the experiment flow, the lab, the pipeline, any of that, I want to be able to query the database and find that answer. Very operational. The problem that you often run into is data scientists on a digital team in a biotech, every once in a while they get an email: hey, did you analyze that data yet? And they’re like, what data, I don’t know what you’re talking about. Well, because we need it in an hour. For the data scientist to be able to know when the data is going to come off the machine so that they know they have time in their schedule to analyze it, whatever it is, all those questions all the way back to the very early ideation. I want that all in a database.
Ross Katz: Interesting. It’s a combination of the metadata of every experiment, every input, every output, but also bringing data science into the production process of the research that’s happening, allowing them to integrate more seamlessly into it by understanding everything that’s happening. Other than just a database that has all of that in there, what would that look like to you?
Jesse Johnson: What it would mean is having software that helps the wet lab, the bench teams, do what they need to do in a way that it’s recording all those steps as you’re going along. Once they say, oh, I really want to get this answer to this question about these compounds, they can go into the tool and it encourages them to say, okay, here’s an experiment that I want to run. I don’t know when I’m going to run it, but I’m just going to put a placeholder there. As they flesh out that experiment, it makes it easier for them to define that. Now you’ve got records of all these experiments that may happen, and the data scientists maybe can go in and say, oh, if you change these parameters, it’s going to be better in the end. Then continue to facilitate that process all the way down to where they’ve put the protocol together, they’ve put all their plate maps together, it’s ready to go into the machines. Having software that encourages them to use the software and gives them an incentive to record all of that.
Ross Katz: What I hear you talking about is the collaborative elements of it, of both giving the data team and scientists in the wet lab formative input into what’s happening, but then also ensuring that both parties are getting what they need from the initial phase and what they need from the outputs, beginning with the end in mind.
Jesse Johnson: Yes, absolutely. Because the only way that you’re going to get users to do this—to put in the effort that makes sure that data is there—is if they’re getting something out of it. If they can enter the experiment details in a way that allows a robot to do the experiment for them, maybe that’s the incentive. Or maybe it’s something else.
Ross Katz: It wouldn’t be a Data in Biotech podcast if we didn’t at least touch on AI and ML. What role do you see AI and ML playing in the biotech startup space right now and how do you expect to see that role evolve?
Jesse Johnson: It’s starting to be everywhere. From target discovery to interpreting results from instruments, in particular, high-content imaging, HCI, phenotypic discovery. There’s a lot of really interesting stuff there. It’s coming in in analyzing sequencing data, looking at NGS data. We’re already seeing it being applied pretty much every place in the pipeline. I don’t think that’s going away. With ChatGPT, LLMs, harder to see where that fits in. The way I see those tools being used—or the way I would imagine them being used—is almost as a learning tool, because you want to check what they do. If you’re designing an experiment, you could maybe type in some freeform text that says, here’s how I want you to do the protocol, and maybe it writes it out for you. But you’re still going to want to check that. At some point, writing it out longform is going to be more effort than just going in and designing the details. As a bridge to those more specialized—and maybe beyond that, too, but that’s my current thinking is that’s where there’s a pretty good use for it and need for it.
Ross Katz: That makes a lot of sense. There’s the traditional machine learning that’s being applied at every phase of the drug development process and then there’s LLMs that are helping to reduce the manual effort of generating these longform pieces of content but also try to learn what best practice is and embody that best practice in the way that you’re applying it there.
Jesse Johnson: That’s right. We’ve already seen some POCs of this. I know Synthace a few weeks ago, maybe a month ago now, released a blog post about this, and a couple of others have as well.
Ross Katz: As we bring the conversation to a close, I’m just curious if there’s any pieces of content or anything that you’ve listened to or read recently that you think has been really interesting or helped to understand the space a little bit better.
Jesse Johnson: That’s a great question. Kaleidoscope just put out a blog post on the different phases that biotech startups go through, where they asked VCs, okay, what are you looking for at each phase of that. That was really interesting, especially the parts about data assets and how important that was. There’s this guy Ben Stancil who has a newsletter about modern data stack in general, and he had this really interesting one just a week or two ago about why big solutions are often hard to adopt and why corporations often only need 10% of the big solution. I thought that was really interesting. Also not for biotech, but in terms of the question you had asked about software adoption and what are companies doing right or wrong.
Ross Katz: Interesting. Jesse, I really appreciate you joining us today. We’ll post all the places that people can find you on the blog post that we associate with this, so we don’t need you to say it online. But thanks so much for joining. It’s been a pleasure and we’ll talk down the line.
Jesse Johnson: Thanks so much for having me. This was a lot of fun.
Jason: And that’s it for this episode of Data in Biotech. If you enjoyed the episode, please subscribe, rate, or leave a review in your podcast platform of choice. See you next time.





