Listen on
Overview
Biology is a multifactorial system, yet most labs still run one-factor-at-a-time experiments on paper-era tooling. The consequence for R&D leaders is predictable: slower discovery, higher cost per experiment, and fragmented data that cannot support downstream AI or transfer learning. When methods cannot account for interactions between variables, the resulting data cannot either.
This episode is an abridged pairing of two earlier Data in Biotech conversations, re-released while the CorrDyn team was at its annual retreat. Host Ross Katz first revisits his interview with Markus Gershater, CSO and co-founder of Synthace, on the digital experiment platform that converts a scientist’s no-code experiment definition into liquid-handling instructions and structured metadata, including the Tecan RoboColumns partnership for parallel purification. Ross then revisits his conversation with Wolfgang Halter, who leads data science and bioinformatics at Merck Life Sciences, on the open-source BayBE library and how Bayesian optimization delivers 30-50% time savings (up to 95% with transfer learning) on cell-culture media, viscosity-reducing excipients, and other multi-objective problems. Together the segments cover the methodology, the tooling, and the data standards biotech R&D needs to make experiments more informative.
Key Takeaways
Multifactorial experiments produce the answer in fewer runs, not more.
One-factor-at-a-time experimentation feels careful, but it forces teams into long iterative cycles and misses the interactions that determine outcomes. Gershater points out that Synthace customers run fewer experiments to reach a goal because a single 384- or 1,536-well multifactorial run can map a response surface that would otherwise take months of sequential work. The result is faster assay development and bioprocess optimization, not just more data.
Bayesian optimization turns each experiment into information about what to run next.
Halter’s team built BayBE (Bayesian Backend) because classical DOE designs a grid and stops learning. Bayesian optimization incorporates priors, updates them with every result, and balances exploration against exploitation — and it reports how much additional information the next experiment is likely to produce, which gives teams a principled stopping criterion. Merck Life Sciences sees 30-50% time savings consistently, and up to 95% when transfer learning warm-starts the model from related past campaigns.
Structured metadata is the prerequisite for both multifactorial DOE and AI.
Both guests land on the same point: the data model matters more than the algorithm. Gershater describes getting metadata “for free” because Synthace plans every pipetting step digitally and knows exactly what went into each well. Halter identifies the lack of a universal data model — and the prevalence of Excel and paper-style ELN use — as the single biggest barrier to cross-lab insight generation and transfer learning across Merck’s labs.
Full Transcript
Ross Katz: Welcome to the Data in Biotech podcast. I’m your host, Ross Katz. It’s a holiday week here in the US and the CorrDyn team is currently getting together for our annual company retreat, so we decided to do something a little different. This week, we’re re-releasing abridged versions of two of our most popular episodes. These are dedicated to a critical workflow at the intersection of data and scientific research, design of experiments. Across these interviews, we brought together two leading experts in the field to give you a comprehensive overview of how best practice DOE works in the biotech industry, Wolfgang Halter from Merck Life Sciences and Markus Gershater from Synthace. I also wanted to thank you for continuing to listen to Data in Biotech. Since launching last year, the show has continued to rank among the top 50 life science podcasts in the US and it’s been downloaded over 13,000 times. We’re always looking for ways to improve the show, so if you have any feedback or recommendations on how we can make Data in Biotech better, please connect and send me a message on LinkedIn. I’d really love to hear from you. Now, onto the show.
Markus Gershater, welcome to the Data in Biotech podcast.
Markus Gershater: Thank you very much. Really glad to be here.
Ross Katz: Appreciate you joining us. So just to get us started, in a minute or two, could you just give us your background and an overview of your career to date?
Markus Gershater: Sure. Yeah. I’ve had a pretty varied career. The thing that ties it all together is, essentially, biology, a passion for biology. So I actually started out in plant biochemistry, although to be honest, it could have been any area of biology. Pretty much every stage of my career I’ve been quite opportunistic, and in this case, it was an opportunity to work at Kew Gardens for a year here in London, which is absolutely incredible botanical gardens for anybody that’s not been there, I recommend it very highly. And so I just jumped at it. Since then, I’ve done bioprocess development, I’ve done synthetic biology, more latterly I’ve learned an awful lot about how therapeutics are discovered and developed. The other common theme is that I’ve always been at the interface of different disciplines, those interfaces with biology. Chemistry, maths, computer science. I think it’s always that interface where the most interesting stuff tends to happen. I don’t tend to be as interested in where things get really pure and super deep. I’m all about big concepts coming together, big ideas, different mindsets, and what you can learn from people who’ve essentially trained in a very different discipline to your own.
Ross Katz: So that multidisciplinarity and the way that the different disciplines feed into each other…
Markus Gershater: Right.
Ross Katz: …and get synthesized together.
Markus Gershater: Right. And I wouldn’t call myself multidisciplinary really. I am still very much a biologist. I guess I’m part biologist, part businessman these days as much as it almost makes me slightly weird saying it. But it’s that skill of being able to talk across disciplines to really try and understand what someone’s saying, even though they might be saying it in a really unfamiliar way, unpacking things. I think it’s a lot about communication more than anything. I was chatting with someone the other day about how communication can be a real superpower, but I think it’s one that’s sometimes underestimated. Could you just provide an overview of Synthace’s platform and how it helps R&D teams working for biotechnology companies to experiment in smarter ways and to drive that broader learning that you’re talking about driving? Synthace is a digital experiment platform, which is not a term I’d expect anybody to understand what it means because we came up with it ourselves to try and describe what it is we do. Because we don’t fit into a particular neat bucket. We get a lot of questions of, ‘Oh, so you’re like an ELN?’ and we’re like, ‘No.’ That’s electronic lab notebook. ‘Oh, so is it like LIMS?’ ‘No.’ ‘Oh, you’re automation software?’ ‘No.’ We call it a digital experiment platform. What that means is it’s essentially a set of digital tools and capabilities which help the scientist through that process of running an experiment. If you look at the process of running an experiment, it’s a hugely involved process of calculations, logistics, experimental design, just mapping things out about which liquids they have to put in what places, stock concentrations, booking the equipment, making sure the equipment’s actually working, making sure they’ve got the reagents they need in the cupboard and their colleague hasn’t just used it. There’s just a huge amount of tedious stuff. We look to essentially give people the digital tools that allows them to alleviate a lot of that tedium. Essentially what a scientist can do in Synthace is they can map out what they want to do in their experiment in a no-code drag and drop interface, which basically lets them say, ‘Okay, I’ve got these samples and I want to dilute them and then I want to mix them together’ and so they’re thinking about what happens to their samples in their experiment as they go through it. From that definition, the core of the power of Synthace is our planner. That takes that definition and converts it into all of the details of exactly how that experiment is going to be carried out. This is everything from what the stock concentrations might be, how much of each liquid you need, what kind of plasticware you need, how many tips you need, and then every single liquid handling action, every single pipetting action that’s required to carry out that experiment. Then you have a fully detailed map of what has to happen in that experiment to do what the scientist has just defined in the system. Then what Synthace can do is it can convert that map into automation instructions. We interface with the most common liquid handling automation in the lab, and so they can then send it to that automation and it’ll carry out your experiment for you, and then it can gather the data that comes from the end of that experiment. But this is where one of maybe the hidden benefits comes in. You’ve had all this benefit as a scientist, it’s planned stuff for you, it’s programmed the automation for you, you haven’t had to do all the pipetting, wonderful. Maybe you’ve done a more complex experiment because it’s suddenly lifted the burden of all that planning and detail from you so that you can think about higher-level things, which is absolutely what we want to enable. But also by doing this, we have created a map of exactly what’s happened in that experiment. When we get the data, we can associate it with exactly how that data was produced, i.e., all the metadata which describes how that data point was produced. That’s something which you essentially get for free because you’ve been working in the digital world throughout. You’ve been using digital tools at every step of the process.
Ross Katz: It seems like there’s a lot of things that the Synthace platform is accomplishing as part of this lab design, research automation process, experimental design automation process, and it feels like at a high level, that data that you’re capturing at the end, the metadata along with the experimental results, is trying to improve the feedback loop that drives the research process, that drives organizational learning when I think about dozens of researchers researching things in parallel. You want them building on each other’s insights, not thinking along their own paths individually and chasing down wherever they’re going. Is that part of the picture that you see emerging from the data?
Markus Gershater: Absolutely. The thing is, that picture that you’re painting there is one that I think everyone aspires to. This beautiful record of exactly what went on so that then if someone’s doing something similar, they can hopefully even have the system just say, ‘Oh, hey, one of your colleagues did something similar earlier. Why don’t you try this?’ Which, to be clear, isn’t in the Synthace platform yet. That’s the direction we’d want it to go because as you’re building up ever more of these structured and fully detailed data and metadata sets that describe all these experiments, that’s exactly the kind of foundation of data that we need in place as organizations to then make much more rapid progress. How is artificial intelligence going to understand what makes a good experiment unless it’s got access to those kind of data which describe a lot of experiments in full?
Ross Katz: That makes a lot of sense. Some of the benefits that I could imagine from this kind of platform would be increasing the speed to new discoveries, increasing the quantity of experiments that you can run in parallel through the multifactorial design. Are you seeing customers of Synthace getting these sorts of benefits, and how have you seen it transform their organizations?
Markus Gershater: It’s been really quite cool because we make these tools and then you give them to scientists and they do super cool things with them. It never gets old seeing what people do with things. I think it’s interesting what you say there, in terms of running more experiments. Actually we see people running fewer to get to a particular goal. Essentially, if you’re looking at one thing at a time and you’re not parallelizing stuff in a multifactorial experiment, then you’re doing iterative experiments and you actually have to do more experiments than if you can just parallelize everything. It’s a much more complex and higher throughput experiment that you’re doing and it’s got all of these multi-dimensions to it, but it’s basically answering the process a lot quicker than you would otherwise. We’ve found particularly in drug discovery, areas like assay development, when you’re first trying to work out the assay for high throughput screening or for screening for a new therapeutic, that process of assay development can be really tough. It’s highly complex and there’s lots of factors involved. What we’re finding is that when people are doing these high-dimensional experiments, they’re getting to the answer a lot quicker because when you’re working at that kind of scale of biology, these biologists are used to working in 384-well plates or 1536-well plates. That’s 1,536 wells in an area like this; for people who are listening and can’t see me, I’m just holding up my hands in the shape of a 96-well plate. It’s really a tiny scale, but what that means is they can do huge numbers of runs. If you can use those runs to comprehensively cover a multi-dimensional landscape, then you can map out that landscape in exquisite detail and you can get the answer to what are the best conditions for my assay or what are the best conditions for me to grow these stem cells or what’s the best way for my biology to work. You can get that answer even within a single experiment. I was not expecting the platform to be able to be used for things of such power because I wasn’t thinking of doing 1,500 runs. Because personally my biology has not been at that kind of really small scale. Really cool scientists working with a platform and frankly working with some really cool automation as well that can take things down to that kind of miniaturization.
Ross Katz: Under the umbrella of a single experiment, you’re able to do all of these micro-experiments that provide information to each other in really intelligent ways so that you can see the full picture of how what you’re working with responds to the different treatments that you’re providing in that controlled environment. I hear you talking about that as both enabling discovery but also enabling optimization of processes as well. Can you provide a little bit of insight into how it works in those different contexts?
Markus Gershater: It’s mostly about when you have a system that you’re looking to learn more about. There’s some areas of biology where I couldn’t see it applying quite so obviously. If you’re looking for a new molecule that hits a particular target, then you’re going to have to do some kind of medicinal chemistry or high throughput screening or whatever, and so this high-dimensional experimentation of the sort I’m talking about won’t necessarily apply. But we see this all the time, the things that we’re using as biologists, the methods we’re using are often highly complex. But we need those methods to be exceptionally effective, because they are the foundation upon which the data’s being built. If we have poor methods, then we’re not going to get good data. Whenever there’s a method that needs to be developed and understood and optimized, that’s when these kind of methodologies really help out. I mentioned a couple there; assay development is one. For a biochemical assay, there’s lots of different components you might want to put into the liquid of that assay to make it optimized. This is a brilliant way of working out the optimal mixture. Or similarly, for growing cells, and then you go into actually producing the drug substance, so bioprocessing, and this kind of thing’s used all over the place. Optimization of processes can be very powerful.
Ross Katz: Another thing that I’m hearing is that this changes the relationship between an R&D org and their data. It’s asking them to think differently about their data. One of the insights that I appreciate is that there’s a difference between data that’s specifically designed to drive insight versus data that is just accumulating and then mined for insights. How does the platform facilitate that kind of relationship?
Markus Gershater: I think that’s a really key observation. If we zoom out for a second and think about biology as a space and the kind of data that we need to understand biology, I think we’ve got a major disadvantage and a major advantage when we’re talking about running experiments. A major disadvantage is that getting data from experiments is always going to be expensive. You compare it to, I don’t know, getting data from social media to run some kind of AI to work out how to market something better or whatever, you can just harvest all that. But if you have to go in a lab to create every data point that you’re going to make, that is going to be very expensive. The big advantage we have is we get to determine every single data point that we produce. We can choose which data points we’re going to produce in order to try and learn about a system. That then gives us very different opportunities for the kind of machine learning, the techniques we might use for exploring those datasets compared with when you have these really big more amorphous sets of data. We have those more amorphous sets of data in biology as well; very typically the kind of data we come across is these very big multi-omics datasets. Talking with someone the other day and they were talking about the millions of different genomes they have sequenced. That’s a very useful type of data, but that’s the one that’s already quite well understood and people understand the value of it. That’s why I focus more on this other much tighter, probably medium-sized data. It’s not big data in that respect and it’s not small data, but it’s somewhere in between. But it’s highly directed because it’s data which results from an exceptionally well-designed experiment which is there to specifically drive insight. I do think that is something a bit different.
Ross Katz: So it’s a different relationship to data. You’re not trying to acquire as much data that exists in the world as possible. You’re trying to curate a very targeted dataset that is measuring the phenomena that you care about in order to accomplish the goal that you’re working on and designing the data effectively to accomplish that goal.
Markus Gershater: Right. And it’s because the data are expensive to make. Where the Synthace platform comes in then is we’re trying to take away some of that expense, hopefully 90% of that expense because a huge amount of it is planning how to run the experiment, sitting there pipetting, trying to copy and paste all your dataset together. There’s just so much tedium which goes into your average biology experiment which is frankly unnecessary. Can we make those data points cheaper? Can we make them more useful? Because often the data points you can produce by hand, you’re limited by what you can feasibly do by hand. You’re limited by the complexity of what you can have in that 384-well plate because if you’ve got stuff changing every single well, that’s incredibly difficult to keep track of. And then someone comes up and taps you on the shoulder in the lab and says, ‘Oh, hey, did you order in those Falcon tubes?’ and you’re like, ‘Oh yeah, I did. They’re on this shelf in the store cupboard’ and then you go back to your plate and you’re like, ‘Ah. Where did I get to?’ There is just a limit in complexity of stuff that you can do. What we’re trying to do is lift a lot of that burden from the scientist so they can do the experiments that really let them get that insight into the biological system that they’re working with.
Ross Katz: I can imagine it’s hard as a research scientist to zoom out from when you’re having to do the manual repetitive labor of pipetting to keep the strategic experimentation in mind of what you’re trying to accomplish and how you’re going to get there. I would imagine that you’re freeing up a lot of mental capacity among research scientists to think bigger about what they’re doing and the direction that they’re taking their experimentation. My understanding of the platform is that there’s the design of the experiment and then an understanding and then that gets outputted. I’m curious where experimental outcomes enter the picture for Synthace because I’m imagining after the experiments are run, there’s a process by which quality or purity or some outcome that you’re trying to drive inside of the plate gets in there. Can you just give a little bit of insight into how you see those outcomes entering the picture?
Markus Gershater: That’s a really good question. If we think about ‘Okay, what enables an outcome?’ It’s essentially the data and the metadata. If you look at a lot of XY plots of an experiment, it’s essentially metadata along the bottom in the form of the conditions that have been run or whatever, and data along the side. Y-axis data, metadata is the X-axis. So long as you are collecting those things, then actually you can give scientists a window into the experiment that they’ve run really quite easily and then they can see the outcomes. Sometimes that’s easier than others. We’ve got a partnership with Tecan, which is the biggest liquid handling manufacturer in the world, and they’ve got this fantastic system for doing multiple purification runs simultaneously. It’s called RoboColumns. The way this works is you have these columns on deck and there’s liquid being injected into the top of them and then it has a plate on the bottom that’s catching the drops that come out of the bottom of the columns. The physical outcome is you end up with a load of plates with clear colorless liquid in them. Those plates themselves are highly anonymous just sat there on the desk on the robot. You’ve got these anonymous plates, but each one of those wells has material in it which pertain to a very specific part of that overall experiment. In the physical world these are all completely anonymous and if you’re not careful, in the digital world they’re also going to be completely anonymous. So you have to have the metadata that understands, ‘Oh, this particular well pertains to this particular step that’s eluting off this particular column which had this sample applied to it.’ When you have that, you can just concatenate those data points for that particular column really easily and just generate exactly the kind of output which a scientist would expect to see in order to then understand the outcome of their experiment.
Ross Katz: Mhm.
Markus Gershater: That’s basic data processing, data structuring. The next layer on from that is, ‘Okay, if I’ve run a really sophisticated experiment, then how do I understand the very sophisticated outcomes of that experiment?’ We also have capabilities in the platform which are specific to running high-dimensional experimentation, specific to running design of experiments, this multifactorial experimental design that I alluded to earlier. There we actually go all the way from helping the scientist to generate that design in the first place, all the way through to building the models that come from the data at the other end. That’s about as far as we go. Most of the platform is just about generating those highly structured data and metadata sets which can then be exported into other things for analysis. We see this as we want to be the experiment engine, the thing that’s producing the data. We don’t have to be the place where all of that data is analyzed. We don’t have to be the place where all that data is stored. This is way too big a problem for any one company to be dealing with. It’s inevitably going to be an ecosystem of all sorts of different tools that come together to solve the overall problem of data within any one of these big companies.
Ross Katz: That makes a lot of sense. If I’m understanding you correctly, the process if they want to do multiple outcome assessments of a 384-well plate that comes out of this process, then they’re exporting the experimental parameters from Synthace and then marrying those up in their other systems with all the other tests they ran on the plates in order to understand the relationships between the responses? Or are you typically seeing that the response that the user needs is in the liquid handler that you’re interfacing with directly or being passed back into Synthace for that kind of analysis?
Markus Gershater: It tends to be passed back in. When you’re talking about that specific experiment, anything that’s really close to that experiment that’s just been run, that all resides within Synthace. If they’ve done an analysis on a different bit of kit that we don’t automatically take the data from, they can just upload that and associate it with the experiment and it then automatically, because again, it’ll tend to be 96 data points or 384 data points, we know what’s in every single one of those wells because we put it there in the first place. Once again you get this advantage even if it’s not something that’s directly integrated into Synthace, it can be manually uploaded and associated. But where it becomes broader is, what about the thousands of experiments that are being run across the whole organization? Some of these experiments I think, you don’t necessarily need Synthace for them in all honesty. If you’re going to do some kind of genomic study or whatever, mostly sequencing, then the benefits of Synthace are probably going to be lower there. But we still want our clients to be able to take all the data from those experiments and the data from the experiments that are done on Synthace and put them into a much larger and broader data warehouse. That’s what I’m referring to.
Ross Katz: That makes a lot of sense. It’s the point at which multiple experiments need to be married together with other things that are happening across the research organization that you expect the data to go outside of the Synthace system and for more customized, more bespoke, more organizational-specific analytics to happen on top of those experiments. It seems like there’s a certain amount of methodological evangelism that you have to do to convince people that design of experiments, that multifactorial experiments versus one factor at a time is worth learning, is worth doing, is worth committing your organization to. Is that correct? And if so, what are the barriers that you face in trying to convince people of the value of that?
Markus Gershater: It’s all too correct. I think it just hasn’t been commonplace enough that people just accept that it’s a method and everybody’s been trained in a very different way to do science. Originally, I was trained in a very different way to do science originally. You can get frustrated at people thinking, ‘Oh, why don’t they just get it?’ but it took me five years. It took me five years from when I first heard about DOE to when I actually started using it. I can’t exactly say that I’ve got any kind of amazing prescience or anything like that. I do think that it does require a bit of evangelism. It requires data. Scientists will always need data and quite right too. We have some absolutely brilliant data that’s been presented by some of our customers at conferences over the years, running 22-factor experiments to work out the best media for stem cells or running single experiments to optimize an assay to a huge degree. There’s all sorts of stuff that’s been done in big pharma or small biotechs with our platform. That really helps. But also I think there is a bit of a wave of change. When we first started Synthace, people had never heard of design of experiments at all really in my experience. When you tried to explain it to them they were like, ‘Oh yeah, that’s your theory.’ They didn’t have any sense that this is an exceptionally well-established and highly effective branch of mathematics that’s been around since the 1930s. But we ran a series of webinars last year on design of experiments because we realized that there was this growing interest and we just thought, ‘Okay, we’ll put out this educational stuff, barely mentioned Synthace at all, it’ll just all be about DOE.’ We got more registrants for the first webinar in that series than for the entirety of the previous year put together. What that’s showing is that there is a bit of a sea change. People are realizing that we need to be updating our ways of working, whether that’s by using more automation or digitizing things or updating our experimental designs or using more AI. I think more people are feeling the same frustration. More people are impatient with this attitude that actually things are fine, we just need to work harder. Which seems to be the almost unspoken thing in biology; people really pride themselves on the fact that they’ve spent 12-hour days, six days a week pipetting in the lab. I think that is changing. It’s really exciting to see that change. But it’s still going to take a lot more evangelism and it’s still going to take a lot more persuasion. But I enjoy it. It’s all good.
Ross Katz: I can imagine it’s really challenging as a research scientist to zoom out when you’re having to do the manual, repetitive labor of pipetting to keep the strategic experimentation in mind of what you’re trying to accomplish and how you’re going to get there. I would imagine that you’re freeing up a lot of mental capacity among research scientists to think bigger about what they’re doing and the direction that they’re taking their experimentation. As we bring this conversation to a close, just curious, what does the lab of the future look like to you? Five years, 10 years down the road, what do you expect to see?
Markus Gershater: I’ve never been able to just conjure in my mind an image of the lab of the future. I’m sure it’s got some automation in there and it’s got digital tools and probably iPads and whatever. Maybe people going around with VR headsets on. But I’m a lot less interested in technological solutions as the methodologies that they’re enabling. I’m a lot more interested in what the experiment of the future is. How those hypotheses are generated, how the experimental design’s generated, how it’s executed. Actually how it’s executed is just the mechanics of it, really, isn’t it? It’s really about what will it look like to drive maximum insight and drive our ability to work with biology the most effectively. For my part, it is going to be quite large-scale multi-dimensional experiments where you’re measuring every possible outcome from that that’s relevant to your system. Those very large structured high-quality deep datasets then automatically piping into machine learning which can then help to distill the huge amount of information that’s there into the key aspects which humans can then interface with well and interrogate before then having that teamwork between man and machine, scientist and machine, to then work out what the next experiment is. That’s the thing that I have a clear vision of and that gets exciting. The actual lab itself? I don’t know.
Ross Katz: Markus, thank you so much for joining us today. It’s been a fascinating conversation and look forward to continuing it at some later date.
Markus Gershater: Yeah, it’s been a real pleasure. Thanks a lot, Ross.
Ross Katz: Wolfgang Halter, welcome to the Data in Biotech podcast.
Wolfgang Halter: Hi Ross. Thanks for having me.
Ross Katz: I know that your team has been working recently on a tool related to design of experiments, DOE. Previously we had Markus Gershater from Synthace on to talk about DOE, so I’m interested in the work that your team has done related to DOE. I know that you’ve been working on an open-source tool called bayes. Am I pronouncing that correctly?
Wolfgang Halter: Yeah, I think we call it bayes.
Ross Katz: Okay. Yes. If you wouldn’t mind just talking a little bit about how your team integrates with design of experiments more broadly at Merck Life Sciences and how that led you to the development of bayes.
Wolfgang Halter: That’s a good question. Bayes stands for Bayesian backend. That already tells what it is about. At the core it’s a toolbox around Bayesian optimization. Bringing that together with design of experiments is already showing you a little bit what our way of thinking is in that area. We started our first design of experiments use cases four years ago when I also started the company and there were probably some more of them around before I was there even. My colleagues and I, we were coming more from the scientific and academia world. So we were also thinking, why are people still doing these old-fashioned design of experiments things? There are so much better methods or better approaches available these days. But when you look into the public offerings, vendors and professional software for that, you don’t find too much going beyond classical DOE. That’s where we started developing more probabilistic-based approaches, more iterative approaches where you take into account all the information that you have about your system that you’re trying to optimize with each optimization step. I think that’s at the core of bayes.
Ross Katz: Interesting. I think it’s worth spending a little bit of time on what were the classical approaches to DOE and why is the Bayesian approach the right alternative?
Wolfgang Halter: For me the classical approaches are the ones where you look at maximizing your information content and you basically design your experimental space just based on the parameter space that you have. But you usually don’t take into account what kind of experiments you already did in that area. And if you do, you always do a full factorial or something like that. Either way with a limited or new parameter space that you’re looking at. Now with bayes it’s different because we really incorporate all the knowledge into the priors that we have about the parameter space. That really gives us a predictive model about where we expect the highest outcome at what probability. The key here is really including probability measures in your design because that allows you to do a really nice balance between exploitation and exploration where you say, ‘Okay, I really expect something to be much better than that I have experienced so far in that area, but another area also looks promising so I’m going to focus on those two’ while those where it’s just very unlikely, you don’t do any experiment at all. It’s really this incorporation of uncertainty and previous knowledge about your system. That’s the difference here.
Ross Katz: Interesting. In this context, going iteratively one experiment by experiment would somehow be optimal in terms of resource utilization, but there’s also a speed component here too where you might want to parallelize a lot of these experiments. You mentioned a 96-well plate; you’ve got machines that only do 96 experiments at the same time and so you might as well leverage that capacity in order to optimize your throughput. Are there other considerations that you bring to the table when you think about advising different units within Merck Life Sciences about how to approach Bayesian optimization and when to parallelize versus when to be more sequential in the experiment?
Wolfgang Halter: It really depends on how long it takes for one experiment to look at. If one experiment takes a few minutes then it’s not necessary to do ultra-parallel. However if you’re looking at, I don’t know, stability experiments where you put something in a freezer for three months and then test it afterwards, that’s your limiting factor in the end and then you have to treat those differently. You can’t do that sequentially in the end. That’s I think the main deciding factor for how you set up your sequential versus parallel design. In the end the upsides are really mostly about speed, getting to better results faster with less iterations. That’s the main point about this. As a side note, it also helps to save a lot of resources. As you said, ideally you also spend less chemicals and less materials in your experiments. That’s also something that would define your approach. How much importance do you want to put on resource saving versus time saving? You have to choose your weights in a multi-objective optimization problem basically.
Ross Katz: Interesting. In order to clarify things before I ask more about the benefits and how it’s being used internally, could we take an example experiment? I know that you use the construct of a campaign, an experimental campaign. An example experimental campaign that might be done at Merck Life Sciences or some other place and how a person or a team goes about thinking through how to structure an experimental campaign in bayes and where the insights come from through that process?
Wolfgang Halter: The general points are always the same. In the end it starts with an optimization problem. You want to make something more optimal. Sometimes that could be the formulation of a therapy or it could be even in a digital space like fitting the parameters of a digital twin model or something like that. It could be either. That’s your starting point. Then the first step is really figuring out the search space. What are your parameters? What are your variables that you can vary to influence your target? Let’s take cell culture media for example, that’s a very nice example. You have very many components in cell culture media that you can vary. It’s many parameters at the same time and then most of the time even on a continuous scale. It’s not even choosing zero or one, but choosing any number between zero and one basically. The search space is quite big and the mix of continuous and discrete variables, that’s usually not a very easy problem to solve. Then if you have your search space, you’re looking at your objective again. Oftentimes we start with, ‘Yeah, we want to create the best cell culture media.’ At some point you need to formalize that and think about, okay, what does it mean? What means best in this place? Really defining those objectives is actually not as easy because you have often times many different objectives that are competing with each other. In many cases you end up with a multi-objective optimization where you have to really then define, okay, what’s the most important objective for you and how do we want to design this in the end.
Ross Katz: On that objective front, think about that as quality and quantity for example? Making more of something might mean you make a lower quality of something and so you need to weight the two objectives against each other and help the model understand how to balance quality versus quantity in the objective. Am I thinking about that right?
Wolfgang Halter: Yeah, it could be. For example, cost versus quality, that could be two objectives competing with each other. For example if you choose raw materials that are of lesser quality but they are cheaper and you may achieve the same outcome in the end at a lower price. But for us I think we didn’t encounter that too often to be honest. Most of the time we’re looking at things like ensuring pH in a certain range, but at the same time you want to have a certain activity of your proteins. pH and activity also influence each other. Those are the different objectives and it could be a quite big list actually in the end that we’re facing. We have the search space, we have the objectives and then basically we translate this using bayes. It’s a super simple interface. You translate this into an optimization problem and then you just get started. You get your first recommendation for your first experiment, people conduct that first experiment or first set of experiments, they play back the results into bayes and they get the next round of recommendation.
Ross Katz: Is there an endpoint? How do you know when you’re done with this kind of process or is it just that you reach some threshold where you feel like this is good enough?
Wolfgang Halter: That’s a very good question because that’s also one of the advantages of Bayesian optimization. You can actually get this measure of how much information gain can you expect with additional experiments. That’s also something that you can’t get from classical DOE. In classical DOE you can say, ‘Okay, maybe I just need to make my search grid finer and I conduct even more experiments’ or ‘I need to extend my search space somewhere.’ With Bayesian optimization you usually get a good measure for how well you covered your search space already and how likely is it that you find something that is even better. That’s actually a feature that we’re working on right now to also implement these dynamic stopping criteria.
Ross Katz: Interesting. I’m assuming that you developed this open-source library because you were already doing these sorts of things internally, you thought that having an ecosystem around this would be valuable. Could you share a little bit about why you decided to open source this?
Wolfgang Halter: It was a process certainly. We started as I said with one design of experiments use case, we tackled it with a Bayesian approach. After that we encountered one or two more of those use cases and then we reached out in the company and realized, okay, there’s huge potential for these kinds of applications and use cases. But each of them needs to be tailored in one way or the other. So initially it started out of just getting more efficient internally to be able to scale and deliver use cases in a faster fashion. Then at some point we said, ‘Okay, this is not just for us. We can also partner up with academic institutions and we can also get some input from other partners that are working on these kinds of problems.’ For us that was basically the idea then to go open source and to share this with the community to also get input from the community.
Ross Katz: Has that input been valuable? Has it been formative in the tool to date or are you mostly hoping that over the course of time it will continue to help the tool develop?
Wolfgang Halter: We open sourced it in December of 2023. So it’s really relatively fresh I would say. But we have been working together with Acceleration Consortium in the past already even before open sourcing it. We already got some really valuable input through them. We are hoping that we can expand that even further.
Ross Katz: Turning back to bayes and how it’s used internally, what has the adoption been like from data scientists on your team, from biologists or bench scientists across Merck Life Sciences? How is it being used today?
Wolfgang Halter: There are different levels of answers to that. When we look at core, bayes is a software development kit. It’s really made for data scientists to develop things. That means we have this layer of developers of applications that is still there. But every application that we roll out using bayes is adopted very fast and really well. We have a long list of projects that are waiting to be implemented basically even though bayes is speeding things up so tremendously; we collected quite a few projects in terms of demand. What we also see is that since we open sourced, we also allowed OEMs to work with it for things that they develop with us for example. There we also get really good feedback in terms of how easy it is to use and how easy it is to implement. That has been very valuable for us and showed us that we are on the right track.
Ross Katz: Are you able to share any examples of the applications that you’re developing? Obviously you call it Bayesian backend for a reason, so I’m imagining it’s a toolkit for developing the backend of an application but then there’s a variety of different frontends or use cases that people might see. Any examples would be really interesting.
Wolfgang Halter: One example is a product that is based on finding viscosity-reducing excipients. It’s excipients that you put into your formulation that reduce viscosity. You need that to be able to administer drugs in the right way. We basically developed a tool that allows you to test just a few excipients with your protein and then it suggests ideal combinations of those excipients to find the best excipient combination for your protein that you have and that you’re trying out. This is something that we offer our customers to use because we are also distributing some of those excipients and if you would have to test all combinations of those excipients, it will take a lot of time and effort and a lot of resources. We try to minimize that to allow our customers to speed up and go faster to market basically.
Ross Katz: Interesting. It’s a way in which you’re enabling your customers so that they can scale faster and that has benefits for you as well. It sounds like time savings both internally and externally are a big value. Do you have any estimates of how much time you’re able to save in a given domain or is it just difficult to estimate?
Wolfgang Halter: We have certainly some experiences. You can actually find numbers in the literature. Some numbers in literature claim that this approach would lead to 95% time saving. What we experience is more in the area of 50%, which is even then huge if you think about it. I would say the minimum is really 30%, so 30 to 50% is something that we consistently see in the projects that we conduct.
Ross Katz: Interesting. As you have different customers and business partners coming to you and asking for applications of Bayesian optimization to their particular problem space, how do you think about prioritizing which of these use cases you’re going to go after? Is there a way to bundle multiple use cases together in a way that’s smart for your team?
Wolfgang Halter: The latter is always what we think about first. How can we maybe serve someone’s need with an application that we already have developed for another group? Oftentimes this is already the case and that’s the best-case scenario. The other needs are something where we really look at what’s the potential impact of that solution. We try to prioritize based on how we can create the most value for our customers and the company in the end.
Ross Katz: Interesting. I’m imagining when you talk about all these different campaigns, the data that these campaigns collect is really valuable intellectual property that you’re developing, and I’m wondering, do you somehow store for later or leverage or combine or analyze the data that these experiments are throwing off in a way that yields more insight, and if so how do you approach that?
Wolfgang Halter: In principle, any experiments that are conducted in the lab are gold in that sense. It gives you some snapshot about a little bit of truth of the world and you really want to conserve that in a good way. I think that’s one of the biggest pain points in large corporations and with many different labs, to store that data in a meaningful and also inter-compatible way so that you can make sense of it as a whole. I would say we’re not there yet. That would be my vision, that you really have this uniform data model in terms of research and scientific insights where you can just feed all your experimental data in and in the end you can just throw bayes on it and you will get the perfect recommendation for any problem you’re looking for. That would be great. We’re not close to that vision at all. But what we see is for the applications that we have developed like design of experiments, we see these individual campaigns where bayes is running; one of the key features that we’ve just rolled out is really transfer learning. That means learning from past campaigns. When you do this viscosity experiment, when you look at proteins that are similar to that, you can learn a lot about how proteins behave with those excipients and you can basically get a warm start for your DOE. Maybe that’s not the best example because proteins really behave very crazy sometimes, but in other domains, cell culture media for example or for model fitting, these warm start features are where we get from the 50% to the 95% time saving.
Ross Katz: Are the conditions for that kind of transfer learning that the parameter space and the objective have to be the same or at least similar enough? How do you think about when transfer learning can be applied to a given campaign?
Wolfgang Halter: For us right now the status quo is that the parameter space at least in terms of what kind of parameters you have — maybe you expand the parameter space, but at least the physical properties are the same — that’s basically the condition right now I would say. You can extend it to unseen parameter spaces, that’s also possible, but you need to have some overlap. Without the overlap then you don’t gain anything. That’s certainly the main limitation right now. But that’s not something that you can fix.
Ross Katz: Right. So there’s exploration and exploitation happening, so even if you’re in the wrong place, the exploration will lead you eventually to the right route to use your metaphor. This method of optimizing or of learning from experiments just seems like the clear best path of the methods that we have available today, which makes me wonder what has stood in the way of adoption of Bayesian optimization in this context previously? Is it just that the open-source toolkit hasn’t existed before and now it does or is it something else that I’m missing?
Wolfgang Halter: I think there are several factors to that. I would have to say even classical DOE is not utilized to its full potential when you look at how people in the labs do scientific experiments. User interface and the gap between the developers of DOE methodologies and the ones who are using it in the end, that’s probably the biggest challenge for DOE in general. I believe tools like Synthace that you had earlier on this podcast can help significantly to lower that entry hurdle. That was certainly one of the biggest bottlenecks in the past. For Bayesian optimization in particular, it’s also something that comes with the computational power we have available these days. When we look particularly at transfer learning, you need computational power. When you have many parameter dimensions and you take into account all the data points that you have collected in the past, the standard Gaussian process approach is just eating up memory and computational resources. That’s something that we are facing still; even now with the seemingly unlimited cloud resources that you have, we see it is computationally hungry. We’re working on solutions for that to make it more efficient particularly for larger domains, with many data points and still being able to do transfer learning. I think that is certainly also one reason why it didn’t get that much attention maybe 10 years ago. But then again there’s so many things where you in hindsight think, yeah why didn’t they do it from the beginning.
Ross Katz: Right. Well, yes, and I also know that there’s a cultural aversion between practitioners of frequentist statistics and Bayesian statistics and so there’s probably an element of that in there as well, the extent to which classical methods had a hand on people’s viewpoint of the world. With the time that we have left, I’m interested in zooming out a little bit and getting your perspective on what you think are the biggest challenges that are currently facing the biotech industry at large regarding data handling, data analysis, data science in organizations like yours.
Wolfgang Halter: I think we touched on that earlier a little bit in terms of this universal data model. That’s for me the vision that I would like to get to. Right now we’re seeing still way too much Excel, Microsoft Excel. There’s a lot of manual data collection being done still in the lab. I think this is really hindering some of the advances. Particularly when you look at cross-lab insight generation. That’s something that I would see as one of the biggest challenges. In general looking at digitalization in the lab, the electronic lab notebook for example is something that many people treat as an electronic form of paper. But it can be so much more. It’s not just a different way of writing things down. It needs to be what I was telling earlier, it needs to be this connecting layer between your information. For that you need to treat it also in a different way than just jotting down information. You need to bring a little bit more structure and standardization into those processes, I believe. That’s certainly one problem, and there are data standards for experimental data but I don’t see that we have reached a good level of adopting those data standards. There are maybe also a few competing ones. All the vendors that are out there creating lab equipment, each is basically following a different standard or different interfaces and APIs. That’s just complicating this vision of a connected data layer for the R&D world.
Ross Katz: I’ve heard a lot about the idea that automation in many respects can resolve this problem of a consistent data model because at least the machine that you have is collecting the data points that you need, but that’s predicated on the idea that all of these different machines that you use for automation will come together on a data model that’s at least relatable or combinable together.
Wolfgang Halter: Exactly.
Ross Katz: When you talk about ELNs you’re talking about more structured metadata about the experiments that are happening versus the unstructured writing it down on paper thing. Is that correct?
Wolfgang Halter: Exactly. Yeah, absolutely.
Ross Katz: I know you come from an engineering background and you come into the field of biotech R&D. How do you think that colors your perspective on the problems that biotech R&D faces here?
Wolfgang Halter: That has been very surprising to me to be honest. At the beginning, during my PhD I pivoted a little bit away from just regular engineering into biological sciences. For me it was very different because in the natural sciences in general, researchers really think end-to-end. They think of their research problem from the very beginning, from the atoms to the therapeutic sometimes. It’s really very well-integrated thinking that the natural scientists have and conduct. I think this is really good because you really go deep on your topic. What it prohibits somehow is that you create interfaces. That’s something that coming from engineering, the key dogmas of engineering — why the industrial revolution worked out so well — is this concept of modularity and orthogonality. Designing things in complete isolation of the neighboring modules. Being able to do that allows you to narrow down on just your module, looking only at that and then becoming excellent in exactly this module and then just feeding everything you have outside to your neighboring module. This kind of modular thinking is something that I believe is missing in the natural sciences a little bit because of this end-to-end thinking. That’s something that I wish we can also move into more in the future and that’s probably coming back to the data standard. It’s exactly the same thing. You need to have standardized interfaces and then you can stay within your domain, do whatever you want as long as the interfaces don’t change.
Ross Katz: Wolfgang, thank you so much for joining us today. It was a really insightful conversation and look forward to connecting down the line.
Wolfgang Halter: Thank you Ross. It’s been a real pleasure.
Jason: And that’s it for this episode of Data in Biotech. If you enjoyed the episode, please subscribe, rate, or leave a review in your podcast platform of choice. See you next time!





