Skip to content
Data in BiotechEpisode 2

Why Experimental Design in Biotech Is Broken

Markus Gershater of Synthace discusses what is broken about experimental design in biotech and the opportunity for automation and multifactorial experiments.

41:00Full transcript below
MG

Markus Gershater

CSO and Co-Founder at Synthace

Overview


Biological systems are inherently complex, evolved through a messy process that leaves them highly interconnected and interdependent. Yet, many biotechnology labs still rely on manual, one-factor-at-a-time experimental methods. This fundamental disconnect creates inefficiencies, inflates R&D costs, and generates fragmented data, severely limiting the pace of discovery and the ability to derive meaningful insights. For data leaders and executives, this translates directly to delayed product launches, wasted resources, and a lagging competitive edge.

Markus Gershater, CSO and Co-Founder of Synthace, a digital experiment platform, brings a unique perspective from his multidisciplinary background spanning plant biochemistry, bioprocess development, and synthetic biology. His deep understanding of biological complexity fuels a strong conviction that current lab practices must change. In this episode, host Ross Katz talks with Markus about the limitations of traditional, single-factor experiments and the advantages of multi-dimensional design, detailing how Synthace’s platform automates the entire process from experimental design to structured data capture. The conversation highlights the critical shift needed: from merely accumulating data to generating targeted, high-quality datasets that accelerate learning and lay the essential groundwork for effective AI applications in biotech.

Key Takeaways

Biological systems demand multi-dimensional experiments, not single-factor approaches.

Traditional single-factor experiments are ill-suited for the inherently interconnected nature of biological systems, which evolve messily without regard for simplicity. This approach misses critical interactions between variables. Adopting multi-factorial experimental design reveals the intricate interplay of variables within complex biological systems, accelerating the identification of optimal conditions and driving deeper insights in areas like assay development and bioprocess optimization.

Automation’s true value in the lab is generating structured, contextualized data, not just speed.

Automating lab processes extends beyond merely increasing throughput or reducing manual errors. A digital experiment platform like Synthace captures extensive metadata alongside experimental results, creating a structured, contextualized dataset. This inherent data quality and organization are vital for improving feedback loops, enabling organizational learning, and making subsequent analysis, including AI applications, far more effective.

Targeted, designed data is essential for efficient R&D in biotech due to high generation costs.

Given the high cost of generating experimental data in biotech labs, the focus must shift from simply acquiring large datasets to intentionally designing experiments that yield highly targeted, insight-driving data. This ‘medium-sized, directed data’ is crucial for efficient learning, allowing organizations to answer specific biological questions more rapidly and cost-effectively than through broad, untargeted data collection.

Effective AI in biotech requires specific applications and often benefits from human-AI collaboration.

Discussions around AI in biotech often suffer from over-generalization. True value emerges when specific AI methodologies, like active learning for understanding multi-dimensional experimental spaces or LLMs for documentation, are applied to particular problems. Furthermore, the optimal approach often involves a strategic partnership between AI and human expertise, where AI augments human intuition rather than completely replacing it, depending on the system’s complexity and existing knowledge base.

Related: CorrDyn helps biotech and life sciences organizations build the data foundations discussed in this episode. We specialize in data engineering to create reliable systems and implement data quality frameworks essential for experimental integrity. For organizations exploring advanced applications, we also provide AI strategy to ensure focused, impactful deployments.

Full Transcript

Jason: Welcome to Data in Biotech, a podcast from CorrDyn where we explore how companies leverage data to drive innovation in life sciences. Every two weeks, we sit down with an expert from the world of biotechnology to understand how they’re using data science to solve technical challenges, streamline operations, and further innovation in their business. This week, we’re excited to be joined by Markus Gershater, Chief Scientific Officer and co-founder of Synthace, a digital experiment platform built for high-performance life science R&D teams that help them run more powerful experiments and accelerate scientific progress. During this interview, our host, Ross Katz, speaks with Markus on what’s broken about how laboratories across the globe run biological experiments at scale, the opportunity that exists for researchers in striving to achieve experimental design automation processes, why as an industry we must move towards implementing infrastructure that enables multi-factorial experiments versus one factor experiments, and the role of AI in making sense of complex systems. Here we go.

Ross Katz: Markus Gershater, welcome to the Data in Biotech podcast.

Markus Gershater: Thank you very much. Really glad to be here.

Ross Katz: Appreciate you joining us. So, just to get us started, in a minute or two, could you just give us your background and an overview of your career to date?

Markus Gershater: Sure. I’ve had a pretty varied career. The thing that ties it all together is essentially biology, a passion for biology. I actually started out in plant biochemistry, although to be honest, it could have been any area of biology. Pretty much every stage of my career, I’ve been quite opportunistic. In this case, it was an opportunity to work at Kew Gardens for a year here in London, which is absolutely incredible botanical gardens for anybody that’s not been there, recommend it very highly. I just jumped at it. Since then, I’ve done bioprocess development, I’ve done synthetic biology. More lately, I’ve learned an awful lot about how therapeutics are discovered and developed. The other common theme, I guess, is that I’ve always been at the interface of different disciplines, but those interfaces with biology. Chemistry, maths, computer science. I think it’s always that interface where the most interesting stuff tends to happen. I don’t tend to be as interested in where things get really pure and super deep. I’m all about big concepts coming together, big ideas, different mindsets, and what you can learn from people who’ve essentially trained in a very different discipline to your own.

Ross Katz: So that multidisciplinarity and the way that the different disciplines feed into each other and get synthesized together.

Markus Gershater: Right. And I wouldn’t call myself multidisciplinary, really. I am still very much a biologist. I guess I’m part biologist, part businessman these days, as much as it almost makes me slightly weird saying it. But I guess it’s that skill of being able to talk across disciplines to really try and understand what someone’s saying even though they might be saying it in a really unfamiliar way, unpacking things, that kind of thing. I think it’s a lot about communication more than anything. I was chatting with someone the other day about how communication can be a real superpower, but I think it’s one that’s sometimes underestimated.

Ross Katz: So your career has taken this interesting journey at the intersection of these different disciplines. Why have you gone on this journey? What’s motivated you to take the path that you’ve taken?

Markus Gershater: Biology, in a word. I’ve always wanted to be a biologist. I haven’t ever thought of being anything else. My dad’s a biochemist by original training, and I’m sure that rubbed off on me. But biology is the most incredible phenomenon you can imagine. I think once you start learning about it, probably outside of a school environment, because frankly the way that it’s taught at school, certainly in the UK, is just dreary, no offense to my biology teachers, it’s more the curriculum. It was just lists of facts to learn. But when you actually understand these hugely complex systems and how they evolved and what they’ve evolved into, then it never ceases to delight and surprise. And I’m continually delighted by how wrong I was yesterday about biology. It’s continually fascinating. But it’s also one of the most frustrating disciplines. It’s carried out in a hugely manual, step-by-step artisanal way. The tools that most scientists are using for doing biology, it’s pipettes and a copy of Excel and frankly incredible amounts of hard work and patience and stoicism. And when you contrast that with that huge unfathomable complexity of the phenomenon that we’re trying to study, you just take a step outside it and look at it from the outside and it’s utterly baffling. So that frustration is a motivation of the sort as well.

Ross Katz: Right. So the- that frustration, I’m imagining, is part of the reason why Synthace came into being in the first place, is that right?

Markus Gershater: Right. I think it’s a combination of the enthusiasm and the frustration. I’ve got a completely unbounded belief almost in the good that biology could do if we could work with it as effectively as we possibly can. I guess that just makes the frustration even more so. Because you’ve got all of these solutions just there. We have this absolutely incredible phenomenon that we can work with that’s digitally programmable. But it feels like it should be something that we should just be able to turn to any of the big problems that are facing humanity. And yet at the same time, anybody who’s actually tried to do it knows how incredibly hard it is. It’s an interesting phenomenon if you look at synthetic biology startups, because often they’ll start out in this very starry-eyed way of, “Hey, we can program biology and we can solve the world’s problems.” And then very quickly you realize, “Hang on a second, this system doesn’t lend itself to easy engineering and predictable function.” As any bioprocess engineer will be very glad to tell you.

Ross Katz: So biology is just by its nature a noisy process with lots of variables coming into it and also the methods that people are using to experiment with that process are sort of very manual in nature and error-prone in nature, which adds additional noise to this already noisy system. And am I understanding you correctly?

Markus Gershater: Yeah. And it’s not just noise. I think there’s a really fundamental feature of biology which is under-considered and to try and compensate for this I bang on about it all the time. That is that biology becomes a really fascinating thing when you think about the process by which it came into being. This process of evolution. Evolution is a fascinating phenomenon in its own right in that it has absolutely no regard for how complex a system is that it’s making or how incredibly intertwined and interdependent that process might be. So long as it gives some kind of fitness advantage then evolution will select for it. And it doesn’t matter how messy it is. What you end up with then is this ridiculously interconnected filigreed system which nonetheless has these absolutely amazing properties. But it’s not clean and it’s in no way designed, and it’s not particularly easy to design with. When you accept it as, “Okay, it’s this emergent system, it’s something which is many things stacked upon another and relying on each other,” then you realize that it isn’t just the tools that we’re using that need to change, but it’s also the way we’re experimenting which needs to change. What I mean by that is we can’t just be going in with classical experiments where you are fixing everything but the one factor that you’re most interested in at that particular moment because the way that that factor affects your system will be dependent on how other factors change as well with that system. It’s all interlinked. We’ve seen over many years now and I’ve seen prior to Synthace and now at Synthace, how way more powerful it is if you do properly designed multi-dimensional experiments to explore biology. We talk a lot about the tools. It shouldn’t be manual. It shouldn’t be copies of Excel. It shouldn’t be people spending their incredibly educated lives 14-hour days in the lab pipetting stuff. It absolutely shouldn’t be that. But at the same time, we don’t want to be doing dumb automation either. It needs to be stuff which is doing really elegant structured experiments which give us the most insight possible into this system that we’re trying to learn about.

Ross Katz: So, could you just provide an overview of Synthace’s platform and how it helps R&D teams working for biotechnology companies to experiment in smarter ways and to drive that broader learning that you’re talking about driving?

Markus Gershater: Yeah, I’ve been talking about some very high-level concepts and some very high aspirations. What does it look like when the rubber meets the road? Synthace is a digital experiment platform, which is not a term I’d expect anybody to understand what it means because we came up with it ourselves to try and describe what it is we do, because we don’t fit into a particularly neat bucket. We get a lot of questions of, “Oh, so you’re like an ELN?” and we’re like, “No.” That’s electronic lab notebook. Or, “Oh, so it’s like LIMS?” “No.” “Oh, you’re automation software?” “Sort of.” So we call it a digital experiment platform. What that means is essentially a set of digital tools and capabilities which help the scientist through that process of running an experiment. If you look at the process of running an experiment, it’s a hugely involved process of calculations, logistics, experimental design, just mapping things out about which liquids they have to put in what places, stock concentrations, booking the equipment, making sure the equipment’s actually working, making sure they’ve got the reagents they need in the cupboard that a colleague hasn’t just used. There’s just a huge amount of tedious stuff. We look to essentially give people the digital tools that allow them to alleviate a lot of that tedium. Essentially what a scientist can do in Synthace is they can map out what they want to do in their experiment in a no-code drag-and-drop interface, which basically lets them say, “Okay, I’ve got these samples and I want to dilute them and then I want to mix them together,” so they’re thinking about what happens to their samples in their experiment as they go through it. Then from that definition, the core of the power of Synthace is our planner. That takes that definition and it converts it into all of the details of exactly how that experiment is going to be carried out. This is everything from what the stock concentrations might be, how much of each liquid you need, what kind of plasticware you need, how many tips you need, and then every single liquid handling action, every single pipetting action that’s required to carry out that experiment. Then you have a fully detailed map of what has to happen in that experiment to do what the scientist just defined in the system. Then what Synthace can do is convert that map into automation instructions. We interface with the most common liquid handling automation in the lab, and you can then send it to that automation and it’ll carry out your experiment for you. Then it can gather the data that comes from the end of that experiment. But this is where one of the hidden benefits comes in. You’ve had all of this benefit as a scientist, it’s planned stuff for you, it’s programmed the automation for you, you haven’t had to do all the pipetting, wonderful, maybe you’ve done a more complex experiment because it’s suddenly lifted the burden of all that planning and detail from you so that you can think about higher-level things, which is absolutely what we want to enable. But also by doing this, we have created a map of exactly what’s happened in that experiment. So when we get the data, we can associate it with exactly how that data was produced, i.e., all the metadata which describes how that data point was produced. And that’s something which you essentially get for free because you’ve been working in the digital world throughout. You’ve been using digital tools every step of the process.

Jason: Are you a biotechnology company looking to unlock the potential of your business data? CorrDyn can help. We’re an enterprise data specialist that helps companies working in life sciences make smarter, strategic decisions. From developing the right data strategy that starts with our data maturity assessment to building and delivering bespoke technical solutions, we are equipped to tackle the most complex data challenges. We have partnered with dozens of high-growth organizations, from manufacturers of custom oligonucleotides to molecular diagnostic companies to achieve data competence. Whether you need to supplement existing technology teams with specialist expertise or launch a data program that lays the groundwork for future internal hires, you can partner with CorrDyn to unlock the potential of your business data today. Simply visit connect.corrdyn.com/biotech to learn more. Now, back to the show.

Ross Katz: It seems like there’s a lot of things that the Synthace platform is accomplishing as part of this lab and research automation process, experimental design automation process, and it feels like at a high level that data that you’re capturing at the end, the metadata along with the experimental results is trying to improve the feedback loop that drives the research process, that drives organizational learning when I think about dozens of researchers researching things in parallel. You want them building on each other’s insights, not thinking along their own paths individually and chasing down wherever they’re going. Is that part of the picture that you see emerging from the data?

Markus Gershater: Absolutely. That picture that you’re painting there is one that I think everyone aspires to. This beautiful record of exactly what went on so that if someone’s doing something similar they can hopefully even have the system just say, “Oh hey, one of your colleagues did something similar earlier, why don’t you try this?” which, to be clear, isn’t in the Synthace platform yet. That’s the kind of direction we’d want it to go. Because as you’re building up ever more of these structured and fully detailed data and metadata sets that describe all these experiments then that’s exactly the kind of foundation of data that we need in place as organizations to then make much more rapid progress. How is artificial intelligence going to understand what makes a good experiment unless it’s got access to those kind of data which describe a lot of experiments in full?

Ross Katz: Yeah, that makes a lot of sense. Some of the benefits that I could imagine from this kind of platform would be increasing the speed to new discoveries, increasing the quantity of experiments that you can run in parallel through the multi-factorial design. Are you seeing customers of Synthace getting these sorts of benefits? And how have you seen it transform their organization?

Markus Gershater: Yeah, it’s been really quite cool because we make these tools and then you give them to scientists and they do super cool things with them. It never gets old seeing what people do with things. It’s interesting what you say though, in terms of running more experiments. Actually, we see people running fewer to get to a particular goal. Essentially, if you’re looking at one thing at a time and you’re not parallelizing stuff in a multi-factorial experiment, then you’re doing iterative experiments and you actually have to do more experiments than if you can just parallelize everything. So it’s a much more complex and higher throughput experiment that you’re doing and it’s got all of these multi-dimensions to it, but it’s basically answering the process a lot quicker than you would otherwise. We found particularly in drug discovery areas like assay development. When you’re first trying to work out the assay for high throughput screening or for screening for a new therapeutic, then that process of assay development can be really tough. It’s highly complex and there’s lots of factors involved. What we’re finding is that when people are doing these high-dimensional experiments then they’re getting to the answer a lot quicker because when you’re working at that scale of biology, these biologists are used to working in 384-well plates or 1536-well plates. That’s 1,500 wells in an area like this big, for people who are listening and can’t see me, I’m just holding up my hands in the shape of a 96-well plate. It’s really a tiny scale. But what that means is they can do huge numbers of runs. If you can then use those runs to comprehensively cover a multi-dimensional landscape, then you can map out that landscape in exquisite detail and you can get the answer to what are the best conditions for my assay or what are the best conditions for me to grow these stem cells or what’s the best way for my biology to work. You can get that answer even within a single experiment. I was not expecting the platform to be able to be used for things of such power because I wasn’t thinking of doing 1,500 runs. Personally, my biology has not been at that kind of really small scale. So really cool scientists working with a platform and frankly working with some really cool automation as well, that can take things down to that kind of miniaturization.

Ross Katz: So under the umbrella of a single experiment, you’re able to do all of these micro experiments that provide information to each other in really intelligent ways so that you can see the full picture of how what you’re working with responds to the different treatments that you’re providing in that controlled environment. And I hear you talking about that as both enabling discovery, but also enabling optimization of processes as well. Can you provide a little bit of insight into, like, how it works in those different contexts?

Markus Gershater: Yeah, it’s mostly about when you have a system that you’re looking to learn more about. There’s some areas of biology where I couldn’t see it applying quite so obviously. If you’re looking for a new molecule that hits a particular target, then you’re going to have to do some kind of medicinal chemistry or high throughput screening or whatever. This kind of high-dimensional experimentation of the sort I’m talking about won’t necessarily apply. But the methods we’re using as biologists are often highly complex. And we need those methods to be exceptionally effective because they are the foundation upon which the data’s being built. If we have poor methods, then we’re not going to get good data. Whenever there’s a method that needs to be developed and understood and optimized, that’s when these methodologies really help out. I mentioned a couple there: assay development is one. For a biochemical assay there’s lots of different components you might want to put into the liquid of that assay to make it optimized, and this is a brilliant way of working out the optimal mixture. Or similarly for growing cells. And then you go into actually producing the drug substance, so bioprocessing, and this kind of thing’s used all over the place. Again, optimization of processes can be very powerful.

Ross Katz: Another thing that I’m hearing is that this changes the relationship between a research and development org and their data. It’s asking them to think differently about their data. One of the insights that I appreciate is that there’s a difference between data that’s specifically designed to drive insight versus data that is just accumulating and then mined for insights. How does the platform facilitate that kind of relationship?

Markus Gershater: I think that’s a really key observation. If we just zoom out for a second and think about biology as a space and the kind of data that we need to understand biology, I think we’ve got a major disadvantage and a major advantage when we’re talking about running experiments. A major disadvantage is that getting data from experiments is always going to be expensive. You compare it to, I don’t know, getting data from social media to run some kind of AI to work out how to market something better or whatever. You could just harvest all that. But if you have to go in a lab to create every data point that you’re going to make, that is going to be very expensive. The big advantage we have is we get to determine every single data point that we produce. We can choose which data points we’re going to produce in order to try and learn about a system. That then gives us very different opportunities for the kind of machine learning, the kind of techniques we might use for exploring those data sets compared with when you have these really big more amorphous sets of data. We have those more amorphous sets of data in biology as well. Very typically the kind of data we come across is these very big multi-omics data sets. I was talking with someone the other day and they were talking about the millions of different genomes they have sequenced. That’s a very useful type of data, but that’s the one that’s already quite well understood and people understand the value of it. So that’s why I focus more on this other much tighter, probably medium-sized data. It’s not big data in that respect, and it’s not small data, but it’s somewhere in between, and it’s highly directed because it’s data which results from an exceptionally well-designed experiment which is there to specifically drive insight. I do think that is something a bit different.

Ross Katz: So it’s a different relationship to data. You’re not trying to acquire as much data that exists in the world as possible, you’re trying to curate a very targeted data set that is measuring the phenomena that you care about in order to accomplish the goal that you’re working on and designing the data effectively to accomplish that goal. And I-

Markus Gershater: Right. And it’s because the data are expensive to make. Where the Synthace platform comes in then is we’re trying to take away some of that expense. Hopefully 90% of that expense, because a huge amount of it is planning how to run the experiment, sitting there pipetting, trying to copy and paste all your data set together. There’s just so much tedium which goes into your average biology experiment, which is frankly unnecessary. Can we make those data points cheaper? Can we make them more useful? Because often you’re limited by what you can feasibly do by hand. You’re limited by the complexity of what you can have in that 384-well plate because if you’ve got stuff changing every single well, that’s incredibly difficult to keep track of. And then someone comes up and taps you on the shoulder in the lab and says, “Oh hey, did you order in those Falcon tubes?” And you’re like, “Oh yeah, I did. They’re on this shelf in the store cupboard.” And then you go back to your plate and you’re like, “Where did I get to?” So there is just a limiting complexity of stuff that you can do. What we’re trying to do is lift a lot of that burden from the scientist so they can do the experiments that really let them get that insight into the biological system they’re working with.

Ross Katz: I can imagine it’s hard as a research scientist to zoom out from when you’re having to do the manual repetitive labor of pipetting to keep the strategic experimentation in mind of what you’re trying to accomplish and how you’re going to get there. I would imagine that you’re freeing up a lot of mental capacity among research scientists to think bigger about what they’re doing and the direction that they’re taking their experimentation. My understanding of the platform is that there’s the design of the experiments and then that gets outputted. I’m curious where experimental outcomes enter the picture for Synthace because I’m imagining after the experiments are run there’s a process by which quality or purity or some outcome that you’re trying to drive inside of the plate gets in there. Can you just give a little bit of insight into how you see those outcomes entering the picture?

Markus Gershater: That’s a really good question. If we think about what enables that outcome, it’s essentially the data and the metadata. If you look at a lot of XY plots of an experiment then it’s essentially metadata along the bottom in the form of the conditions that have been run or whatever and data along the side, Y-axis data, metadata is the X-axis. So long as you are collecting those things then actually you can give scientists a window into the experiment they’ve run really quite easily and then they can see the outcomes. Sometimes that’s easier than others. We’ve got a partnership with Tecan, which is the biggest liquid handling manufacturer in the world, and they’ve got this fantastic system for doing multiple purification runs simultaneously, it’s called Robocolumns. The way this runs is you have these columns on deck and there’s this liquid being injected into the top of them and then it has a plate on the bottom that’s catching the drops that come out of the bottom of the columns. The physical outcome is you end up with a load of plates with clear colorless liquid in them, and those plates themselves are highly anonymous just sat there on the deck, on the robot. You’ve got these anonymous plates, but each one of those wells has material in it which pertains to a very specific part of that overall experiment. In the physical world these are all completely anonymous and if you’re not careful in the digital world they’re also going to be completely anonymous. You have to have the metadata that understands, “Oh, this particular well pertains to this particular step that’s eluting off this particular column which had this sample applied to it.” And when you have that, you can just concatenate those data points for that particular column really easily and generate exactly the kind of output which a scientist would expect to see in order to then understand the outcome of their experiment. That’s basic data processing, data structuring. The next layer on from that is, “Okay, if I’ve run a really sophisticated experiment then how do I understand the very sophisticated outputs of that experiment?” We also have capabilities in the platform which are specific to running high-dimensional experimentation, specific to running design of experiments, this multi-factorial experimental design that I alluded to earlier. And there we actually go all the way from helping the scientist to generate that design in the first place, all the way through to building the models that come from the data out the other end. That’s about as far as we go. Most of the platform is just about generating those highly structured data and metadata sets which can then be exported into other things for analysis. We want to be the experiment engine, if you like, the thing that’s producing the data. We don’t have to be the place where all of that data’s analyzed, we don’t have to be the place where all that data’s stored. This is way too big a problem for any one company to be dealing with and so it’s inevitably going to be an ecosystem of all sorts of different tools that come together to solve the overarching problem of data within any one of these big companies.

Ross Katz: That makes a lot of sense. If I’m understanding you correctly, the process if they want to do multiple outcome assessments of a 384-well plate that comes out of this process then they’re exporting the experimental parameters from Synthace and then marrying those up in their other systems with all the other tests they ran on the plates in order to understand the relationships between the responses? Or are you typically seeing that the response that the user needs is in the liquid handler that you’re interfacing with directly or being passed back into Synthace for that kind of analysis?

Markus Gershater: It tends to be passed back in. When you’re talking about that specific experiment, anything that’s really close to that experiment that’s just been run, then that all resides within Synthace. If they’ve done an analysis on a different bit of kit that we don’t automatically take the data from, then they can just upload that and associate it with the experiment. And then automatically, because again, it’ll tend to be 96 data points or 384 data points, we know what’s in every single one of those wells because we put it there in the first place. Once again you get this advantage, even if it’s not something that’s directly integrated in Synthace, it can be manually uploaded and associated. But where it becomes broader is, what about the thousands of experiments being run across the whole organization? Some of these experiments, I think, you don’t necessarily need Synthace for them in all honesty. If you’re going to do some kind of genomic study, so it’s mostly sequencing or whatever, then the benefits of Synthace are probably going to be lower there. But we still want our clients to be able to take all the data from those experiments and the data from the experiments they’ve done on Synthace and put them into a much larger and broader data warehouse. That’s what I’m referring to.

Ross Katz: Yeah, that makes a lot of sense. So it’s the point at which multiple experiments need to be married together with other things that are happening across the research organization that you expect the data to go outside of the Synthace system and for more customized, more bespoke, more organization-specific analytics to happen on top of those experiments.

Markus Gershater: Yeah.

Ross Katz: It seems like there’s a certain amount of methodological evangelism that you have to do to convince people that the design of experiments, that multi-factorial experiments versus one factor at a time is worth learning, is worth doing, is worth committing your organization to. Is that correct? And if so, what are the kinds of barriers that you face in trying to convince people of the value of that?

Markus Gershater: Yeah, it’s all too correct. I think it just isn’t commonplace enough that people just accept that it’s a method. Everybody’s been trained in a very different way to do science. Originally I was trained in a very different way to do science. You can get frustrated at people thinking, “Oh, why don’t they just get it?” But it took me five years from when I first heard about DOE to when I actually started using it. So I can’t exactly say that I’ve got any kind of amazing prescience. I do think it does require a bit of evangelism. It requires data. Scientists will always need data and quite right, too. We have some absolutely brilliant data that’s been presented by some of our customers at conferences over the years, running 22 factor experiments to work out the best media for stem cells or running single experiments to optimize an assay to a huge degree. There’s all sorts of stuff that’s been done in big pharma or small biotechs with our platform. That really helps. But also, I think there is a bit of a wave of change. When we first started Synthace, people had never heard of design of experiments at all, really, in my experience. When you tried to explain it to them they were like, “Yeah, that’s your theory.” They didn’t have any sense that this is actually an exceptionally well-established and highly effective branch of mathematics that’s been around since the 1930s. But we ran a series of webinars last year on design of experiments because we realized that there was this growing interest and we just thought, “Okay, we’ll put out this educational stuff, barely mention Synthace at all, it’ll all be about DOE.” And we got more registrants for the first webinar in that series than for the entirety of the previous year put together. What that’s showing is that there is a bit of a sea change. People are realizing that we need to be updating our ways of working, and whether that’s by using more automation or digitizing things or updating our experimental designs or using more AI, more people are feeling the same frustration. More people are impatient with this attitude that actually things are fine, we just need to work harder. Which seems to be the almost unspoken thing in biology, people really pride themselves on the fact that they’ve spent 12-hour days, 6 days of the week pipetting in the lab. That is just changing and it’s real exciting to see that change. But it’s still going to take a lot more evangelism, and it’s still going to take a lot more persuasion, but I enjoy it, so it’s all good.

Ross Katz: But at the same time in the ether is all of the discussion about AI, machine learning, and how that’s valuable for biology. I’m just curious on your take on how AI factors into the mix of what’s happening in research organizations, the type of research organizations that would be using Synthace, could be using Synthace.

Markus Gershater: First off, it’s worth stating, I talk about DOE a lot. Because actually that is a great step change in the way people work. Automated DOE even better. That’s another step change all over again from there. But actually, I see the future as AI. Synthace as an engine is not a DOE engine. It’s just an engine that allows you to run whatever experiment you need to run to really understand biological complexity. That experiment can be designed by an AI just like it can be designed through a stats package. There’s no fundamental conceptual difference there. But for AI to actually have the effect it could have, it’s going to need these highly structured, very well-curated, fully contextualized, high-quality data sets. That’s the way we can unlock the main potential for it. I think it’s going to be used all over the place in very different ways. One of my frustrations a little bit, actually we’ve just been guilty of it in this conversation, is just talking about AI as this big blob. It occurred to me the other day that talking about, “Oh, we need to use AI,” is almost like saying, “Oh, we need to use electricity.” It’s only meaningful if you actually say, “Oh, we should use this particular form of artificial intelligence for this particular purpose. We should use large language models in order to help us to record our scientific intent a lot more effectively and then to write up subsequent reports. We should use active learning to very effectively explore high-dimensional biological landscapes.” When we start talking in that way, then it becomes a lot more meaningful and we understand the data we’re going to need to make it run. We understand the kind of models that we need to be running to actually get benefit and we can understand the kind of change management, frankly, that we’re going to need to put in place to actually achieve these outcomes. It becomes a lot more tractable the moment that you’re actually talking specifics. So I’m a keen proponent of that. Let’s stop saying, “Hey, let’s use electricity.” Let’s start saying, “Hey, this part of the room needs more light, so we’re going to use electricity and we’re going to put it through a filament and that’s going to give out light.”

Ross Katz: Yes. As I’m hearing you talk, the way that I hear machine learning or AI being targeted is as being part of the feedback loop that helps to understand the data that you have, determine what new data you need to gather, and then design the experiment to gather that data and then understand the data that you have and inserting it as part of the feedback loop between the research scientist and the experimental data gathering that they’re doing to accelerate discovery.

Markus Gershater: Yeah, and you just said something super critical there, because you still kept the scientist in there, which I think is really key. Certainly for some biological systems, and again, I think we need to be talking about the specifics of the problem we’re trying to solve. There are some biological problems where, frankly, a human might think they know what they’re doing, but they’re actually going to be actively detrimental. I might offend a lot of people by saying this, but if you’re doing protein engineering, then a protein is such a dynamic and incredibly complex thing that’s working at the quantum scale. You are not going to be able to intuitively understand that system by yourself. That’s just not going to happen. I don’t care what you say. Humans should just get out the way and let machine learning drive those improvements. On the other hand, if you’re developing a bioprocess, help, I want that really wizened bioprocess guy who’s been doing it for 30 years because they’ve got a load of mental heuristics and models which are actually wholly valid. Depending on the system we’re looking at, we should either be looking to how can we unleash AI in its full, or how can we have AI as the way that then we combine with human intelligence to actually get the most understanding from the experiments we’re running. I really like that you had the scientist in there, because I think for a lot of these systems, it makes sense to have that combination of artificial intelligence and human intelligence, and that’s going to be where we get the most benefit.

Ross Katz: That makes sense to me. As I look at Synthace’s take on DOE, with the multi-factorial studies and understanding where the optimal regions are in that multi-dimensional space for a particular biological process to happen, it reminds me of parameter optimization methodologies in machine learning. You’ve got this grid search approach where you’re randomly trying different things, but then Synthace, as I understand it, is providing statistical tools for scientists to look at that and understand where are the areas of this multi-dimensional space where I could explore further and do more experiments. But there’s this opportunity to look at the gradients, to look at the slope of these regions and understand, okay, these are areas of this space that it might be beneficial to push your experimental space out. Does that make sense?

Markus Gershater: Yeah, absolutely. It’s all about building that model which can help us to get insight. Whether that’s artificial intelligence insight or human insight, ultimately it all has to be human insight. But it’s absolutely essential and those methods that you’re talking about are highly analogous to exactly the kind of approaches that we advocate. With every iteration that we make, it’s another experiment and that has cost to it and we’re just trying to drive those experiments as efficiently and as effectively as possible.

Ross Katz: As we bring this conversation to a close, just curious what does the lab of the future look like to you? 5 years, 10 years down the road? What do you expect to see?

Markus Gershater: I’ve never been able to just conjure in my mind an image of the lab of the future. I’m sure it’s got some automation in there and it’s got digital tools and probably iPads and whatever. Maybe people going around with AR headsets on. But I’m a lot less interested in technological solutions than the methodologies that they’re enabling. I’m a lot more interested in what the experiment of the future is. How those hypotheses are generated, how the experimental design’s generated, how it’s executed. Actually, how it’s executed is just the mechanics of it, really, isn’t it? It’s really about what will it look like to drive maximum insight and drive our ability to work with biology the most effectively. For my part, it’s going to be quite large-scale multi-dimensional experiments where you’re measuring every potential outcome from that that’s relevant to your system, those very large structured high-quality deep data sets then automatically piping into machine learning which can then help to distill the huge amount of information that’s there into the key aspects which humans can then interface with well and interrogate before then having that sort of teamwork between man and machine, scientist and machine to then work out what the next experiment is. That’s the thing that I have a clear vision of and that gets me excited. The actual lab itself? I don’t know.

Ross Katz: Before we let you go, are there any books or papers that you’ve read recently that have stuck with you, any material you would recommend to people as they think about this arena?

Markus Gershater: That’s a fantastic question. To be honest, in my day-to-day work I am dealing with so many different things that I don’t have a specific thing I can talk about. I will recommend there’s some great people writing in the space who are trying to unpick a lot of these things. They’re not so much books, but there’s some blogs and some people that are posting regularly on these things. There’s a chap called Nico McCarty who writes a kind of blog-stroke-newsletter called Codon, and that’s always got some fascinating insight. He’s got some really deep thinking about the nature of science and the kind of biology we should be doing and the biology that’s actually happening. There’s someone on the data science side of things, a guy called Jesse Johnson. He’s got a fantastic series as well. And then there’s a scientist at the Crick here in the UK, Erika DeBenedictis, and she actually has a lot of discussion about the nature of experimentation and how we should be driving this forward and she’s super fascinated with cloud labs and how we can drive experimentation in the future. Those are the people and their output that really spring to mind when you ask that question.

Ross Katz: Markus, thank you so much for joining us today. It’s been a fascinating conversation and look forward to continuing it at some later date.

Markus Gershater: Yeah, it’s been a real pleasure. Thanks a lot, Ross.

Jason: And that’s it for this episode of Data in Biotech. If you enjoyed the episode, please subscribe, rate, or leave a review in your podcast platform of choice. See you next time.

Frequently Asked
Questions

How can we justify the investment in a digital experiment platform when R&D budgets are tight?
The platform accelerates discovery by enabling multi-factorial experiments, requiring fewer total runs to reach a goal than iterative, one-factor approaches. This leads to faster assay development and bioprocess optimization, significantly reducing time to market and overall R&D costs by making each experiment more informative.
What kind of data does a platform like Synthace generate, and how does it integrate with our existing data infrastructure?
Synthace generates highly structured data with comprehensive metadata, detailing exactly how each data point was produced. This high-quality, contextualized data can be exported and integrated into broader organizational data warehouses for further analysis, complementing other 'big data' sources like multi-omics and centralizing experimental outcomes.
How does automating experimental design affect the role of our research scientists?
Automating experimental design frees scientists from tedious manual tasks like pipetting, calculations, and detailed planning, allowing them to focus on higher-level strategic thinking and experimental interpretation. By offloading complex logistics, the platform enables scientists to design and execute more sophisticated experiments that yield deeper biological insights, effectively enhancing their capacity for scientific inquiry.

Need a data partner for life sciences?

CorrDyn helps biotech and pharma companies build the data infrastructure that accelerates research and operations.

Book an intro call