Listen on
Overview
The bioprocessing industry, while critical for drug development, faces a fundamental challenge: its progress is bottlenecked by manual data management. Scientists spend hundreds of hours per experiment on tasks like moving files, aligning timestamps, and overlaying data, rather than on scientific discovery. This manual burden delays insights, increases costs, and shifts highly educated talent away from their core work, hindering time-to-market and competitive advantage.
Host Ross Katz speaks with Martin Permin, Co-Founder of Invert, who brings a unique, software-first perspective from his background at companies like Airbnb and Hive. He discusses how applying software automation principles can resolve these bottlenecks, speeding up process development and scale-up, especially as biologics and AI-designed molecules accelerate drug discovery.
This episode explains how Invert automates data capture, cleaning, and analysis from diverse bioprocessing instruments. The conversation covers how such a platform enables faster experimental design, predictive modeling, and efficient tech transfer between CDMOs and sponsors. It also explores the critical role of data readiness for machine learning adoption and future applications of AI in bioprocessing workflows.
Key Takeaways
Manual Data Management, Not Complex Science, Bottlenecks Bioprocessing
Bioprocessing scientists often spend hundreds of hours per experiment manually moving, cleaning, and preparing data. This extensive manual effort directly delays crucial insights, inflates operational costs, and prevents scientists from focusing on advanced research. Automating these foundational data tasks is necessary to accelerate drug development cycles.
Data Preparation, Not Model Complexity, Limits Machine Learning in Biotech
Many biotech organizations struggle with machine learning adoption because of difficult data management and cleaning, not a lack of sophisticated algorithms. Effective data governance and automated preparation are essential. Resolving these foundational data issues allows for practical ML applications like time series forecasting and predictive modeling.
Shared Data Platforms Reduce CDMO Tech Transfer Risk and Cost
The transaction costs and risks associated with tech transfer between CDMOs and sponsors are high, involving significant time and potential indemnification issues. A common data platform standardizes process lineage and data sharing, reducing manual effort and providing near real-time visibility. This shifts data exchange from static packages to a seamless, collaborative experience.
Leveraging Historical Data Accelerates Expensive Experimental Design
Bioprocessing experiments are costly. Incorporating historical data and organizational expertise (referred to as ‘priors’) into experimental design helps companies begin optimization efforts from a more informed position. This approach reduces the number of expensive ‘design-build-test-learn’ cycles needed to identify optimal process parameters and achieve desired yields.
Related: CorrDyn helps biotech and life sciences companies build strong data engineering foundations. We also guide organizations in successful machine learning implementation and provide managed data teams to accelerate operational improvements.
Full Transcript
Jason: Hi everyone, this is Jason, producer of Data and Biotech. Before we get started, I wanted to let you know about our latest white paper. It’s a comprehensive guide to implementing machine learning models in biotech manufacturing. It’s a complete overview of all the potential problems of ML adoption, and more importantly how to solve them. To download it, simply visit connect.corrdyn.com/biotech-ml. We’ve also dropped the link in the show notes of this episode. Okay, let’s get into it. Welcome to Data and Biotech, a podcast from CorrDyn where we explore how companies leverage data to drive innovation in life sciences. Every two weeks we sit down with an expert from the world of biotechnology to understand how they’re using data science to solve technical challenges, streamline operations, and further innovation in their business. In this episode we speak to Martin Permin, co-founder of Invert Bio. During the conversation with Ross, Martin shared the reasons why Invert exists, to solve the biggest challenges faced in bioprocessing, and the integration of machine learning in the platform to enhance data analysis and decision-making. Martin also emphasized the importance of collaboration between CDMOs and sponsors, the future trends in biotech, and offered advice for tech founders looking to enter the biotech space. Here we go.
Ross Katz: Martin Permin, welcome to the Data and Biotech podcast. Thanks for joining.
Martin Tromp Permin: I’m always a little hesitant to give my background story on a biotech-specific podcast because it’ll be a non-biotech specific background. But I was very fortunate right out of high school to join a company called Airbnb, which worked out in a really big way. And after that was equally fortunate to join a company called Hive, a SAS company that sells enterprise software to marketing organizations. That takes us all the way up to somewhere around 2020 where early COVID I had a little bit of a — I hesitate to say come to Jesus moment because I’m deeply non-religious, but whatever the non-religious version of that is. And realized that I’d spent the last 10 years using bits and bytes to coordinate pixels on people’s screens and that in the face of an actual real world threat had actually almost nothing to contribute. So I ended up roping my best friend and many times co-conspirator, Holger, a childhood friend who’s a wonderful engineer, into building this COVID tracking app for Denmark and for sub-Saharan. And that rolled out and worked out really well and through that I got to meet all these folks in the wonderful world of biotech. And fell in love with their world and has basically spent the four almost five years since then as a full-time first investor in the space and now for the last almost three years a full-time operator and founder in the space.
Ross Katz: Yeah, so what made you fall in love with it? Why are you doing what you’re doing now?
Martin Tromp Permin: Well, the very obvious answer — which should go without saying — is that it is so much more tangible when I talk to our customers and team members that the thing that we’re doing is unambiguously good. That feels really nice. Not that I’ve ever worked in anything that I felt had any real chance of being wrong, but the magnitude of goodness here is just so tangible. Perhaps the non-obvious answer that I maybe should have expected but didn’t is that biotech is populated with people who tend to be really smart, meaning your high school bully probably didn’t fall upwards to become a bioprocess engineer or a molecular biologist or what have you. And secondly it’s usually populated by people who are well-intended. Nobody does it for money or prestige because there’s very little of either to be had outside of the big name founders of exited biotech companies. And what that leaves you with is this selection of smart, nice people, which just means that my day-to-day working with customers, working with the type of people that we hire, is just deeply gratifying and tends to be really nice and stimulating days.
Ross Katz: That’s awesome and definitely resonates with me. Can you give us an introduction to Invert — what’s the organization’s mission and how does it go about accomplishing it?
Martin Tromp Permin: Invert is a software company that we started about three years ago. Invert builds software for bioprocessing of various sorts. Tongue in cheek we say that we want to increase the GDP of biology, which is a wink to the SAS company Stripe who talks about increasing the GDP of the internet. That’s how we frame ourselves. In very practical terms what we do is we help groups who do bioprocessing of various sorts — literally everything from yeast for brewing to cell and gene therapies, with an emphasis on the slightly more profitable and impactful right hand side of that technical distribution. We help those folks grab data off their various instruments, map out their processes, get data in real time from online and offline instrumentation, all sorts of different sources of data. We then help them save an ungodly amount of time cleaning that data and making it square and aligning timestamps and adding context to it. All so that we can get to the point where we really start getting excited, which is helping people do analysis on that data that helps them design better experiments, train machine learning models without writing any code that helps them either design better experiments or predict what’s going to happen in the future for a given process. That’s what we day-to-day work on. There’s a grander mission in play here that we can go into as well, but that’s the day-to-day.
Ross Katz: Yeah, that’s really interesting. Coming at this from the outside, why did you view biomanufacturing and bioprocessing as fertile ground for developing a software company?
Martin Tromp Permin: There’s an inverted — no pun intended — answer of I felt I wasn’t smart enough to go into any kind of discovery work, so that was DQ’ed out initially. But to give you a more serious answer, it’s my general view that what AI does to drug discovery and drug development in general moves the fat part of that funnel into the later bits of drug development, which to me right now very much is process development and scale up of manufacturing processes. In many ways what we’re trying to do here is skate where that puck is going — if we’re beginning to see the first INDs now for these AI-generated molecules, at what point will we just start seeing a clip of them that is fully unheard of, and we won’t remotely have enough human resources to design bioprocesses? Can we build software that helps with that work?
Ross Katz: From the time that companies start going into the more development side of research and development, can you talk to me about how Invert comes into the picture for them — when they start thinking about using Invert and how do they onboard you into their processes?
Martin Tromp Permin: There are a few different typical types of customers we can talk about here. There’s what you’re alluding to, which is your traditional biotech where somebody spins out some technology, runs through the funnel to the point that they’re beginning to manufacture enough drug product — or hopefully about to begin manufacturing enough drug product for a trial — and now they’re beginning to have to think about how that process works. That’s typically our sweet spot for that cohort. There’s a different cohort of larger companies — either biopharma companies, CDMOs, or established bioindustrial companies — where they have multiple programs in parallel, they’ve ideally already scaled n number of processes where n is a pretty impressive number. For those it becomes more a question of what have we already done, how do we avoid redundant work, how do we get faster in time to market or time to milestone as we move those molecules through our process.
Ross Katz: For each of those cohorts, if it’s different, can you walk me through what is the value proposition that Invert brings to them that makes them think this software system is exactly what I need to optimize my bioprocesses?
Martin Tromp Permin: It’s perhaps best explained in a narrative, which is most of our users before they meet us will live a life where they wake up every morning, go in, run an experiment that generates a ton of data, and for the most part there’ll be next to no automated handling of that data. Meaning this instrument just generated a file, somebody has to go and manually move that file into wherever it’s going to be analyzed. Somebody has to manually overlay that on whatever the online time series from, say, a bioreactor or some downstream processing unit op. And all that has to happen manually. I just described that in 20 seconds, but that’s more like 20 hours — actually that’s a gross understatement, it’s more like 200 hours of work per experiment very often. So what tends to be the situation the moment after Invert shows up is that all that is just automated. The time from data to insight gets really shortened. That’s the main value prop for most of our users and customers. When you take a lot of those moments and aggregate them, what you have is just reduced time to milestone or time to market, depending on where you’re at in the stages. It really is just a question of taking the schlep hours that highly educated, deeply smart, passionate scientists spend doing work that is below them and making computers do that for them, so they can go spend their time doing the actual science that they’re passionate about and that computers can’t yet do.
Ross Katz: Yeah, you touched on this a little bit earlier, but what are the different instruments and data sources that Invert integrates, and how does implementation work on your end once they decide to go with you?
Martin Tromp Permin: We work with a number of process development labs, and we work with pilot plants, and we work with full scale manufacturing. If you ask somebody in a pilot plant how that’s different from manufacturing, they’ll have a whole range of answers. If you ask somebody in a PD lab how that’s different from manufacturing, they’ll have a whole array of answers, and there’s a lot of nuance there, and those two are not wrong. At the end of the day those are all trying to emulate each other, and we feel pretty strongly that we should have software that works across those scales such that we’re not reinventing the wheel at every different amount of liquid in a vat somewhere, to be crude about it. Different answers for different stages there, but at a very high level our bread and butter is a lot of bioreactors, various shapes and sizes and various degrees of automation. We feel very strongly that we have to be OEM agnostic — meaning we’ll engage with your Thermo Fisher hardware, your Sartorius hardware, your Eppendorf, or your own home-brewed system of some sort. For a lot of hardware we have standard integrations — the hardware market is Pareto distributed, so there are a few instruments that almost everybody has and those we just have standard integrations to. Every once in a while we’ll encounter some exotic thing or some very old thing and we’ll build a custom integration for that. We also integrate to a bunch of offline analytical instrumentation of various sorts, and obviously for the different types of modalities different types of instrumentation is relevant, all the way down to different downstream processing unit ops.
Ross Katz: Interesting. After you have that data, what are the types of data science problems that bioprocessing companies face that you think Invert positions them to answer more efficiently?
Martin Tromp Permin: I never quite know when to talk about data science and when to talk about just statistical problems — that seems like a bit of a sliding scale. But first off we start by just solving a lot of statistical problems. By the time the scientist we talked about before, who has now overcome the hours and hours of manual work they used to do, has their data automatically in a square format — what else were they going to do with it? Are they going to do regressions on that? Are they just going to plot it out to see what actually happened during a run in a bioreactor? Are they trying to use that data to feed a DOE for the next round of experiments? All that stuff we automate and try to make it a click. We try to build very goal-oriented software in which, rather than just give you a suite of tools, we give you a range of options — what are you trying to achieve here — and then walk backwards from that. Once we talk data science, there are a few things that we’re very excited about, which are first of all these very modern, state-of-the-art time series forecasting type models. We do a lot of hybrid modeling on bioprocesses in which we bring in mechanistic data on cell growth and tie that into a machine learning model on top of it. Right now we’re working a lot with transfer learning and model-based DOE — can we use whatever data a customer has historically generated to inform a much wider DOE than would just statistically be generated if we don’t take that context into account? Those are the types of solutions that we try to bring to the table.
Ross Katz: For the time series forecasting, is that anomaly detection and understanding whether the process is going the way it needs to go, or is that predicting what we expect yields to be based on what you’re seeing inside of the reactor — what are the types of insights that people are after?
Martin Tromp Permin: All of the above. You hit on, non-surprisingly, some of the most common or at least attractive use cases. But there’s a wider question there, which is how early can we call a scrap batch and not sink the cost of the marginal next n number of days of running what eventually was going to be a scrap batch anyway. Foaming is a very attractive thing if you can just solve that — it solves a lot of headaches for a lot of people. But that’s exactly right.
Ross Katz: Are there any of your customers doing unique assays or image-based detection, or gathering unique data sources and integrating them with what you’ve done in order to answer questions in ways you haven’t seen before?
Martin Tromp Permin: Let me tell you an internal story. I recently posted in our internal Slack, in our data science channel, a link to the new Meta image recognition model that came out — call it a month ago.
Ross Katz: Segment Anything?
Martin Tromp Permin: Bingo. About 15 minutes later somebody had taken some old footage of a full bioreactor run in a transparent glass vessel and trained a foaming model that just really works out of the box. That’s magical. I really dislike the word soft sensors, but things like that are going to be more and more prevalent. And then because we work with a few of these synthetic biology companies who manufacture — or at least are hoping to manufacture — various food products, we see a fun amount of endpoint taste testing, which is always a fun thing to work with.
Ross Katz: Yeah, that’s super interesting. Talk to me about what things look like for these companies in the absence of Invert. What are the challenges that companies typically face trying to get the answers to questions like this that Invert makes go away?
Martin Tromp Permin: My personal goal is to make it such that most bioprocess engineers or scientists don’t have to spend a lot of time in Excel — and maybe prior to that goal, don’t have to spend a lot of time looking for files. That’s really where there’s just a ton of value to be collected back. I would say 80% of the time, by company, when we meet a new company what we walk into is some file management system — think OneDrive, think Google Drive if it’s a cool startup — on top of a data management suite, and then Excel or JMP, or Jupyter Notebook if we’re really exotic, for the actual analyses on top of that data. The remaining 20% there’ll be some internal infrastructure in place — a data lake in a warehouse, often a bunch of data pipes that have been built out painstakingly over years and that tend to need repairs every once in a while when a vendor ships a breaking change to that API or OPC UA connection. That’s really the main picture in the before and after — how much time are we spending there. And from there it is so easy for us to automate all the Excel and JMP work that otherwise would happen, which is every time I run my Amber again I’ll have to redo all my formulas and calculations in a new spreadsheet because it’s a new run. Because we have all the context of what this data is and where it’s coming from, and we have a semantic understanding of what each metric or variable is, we auto-calculate all the oxygen uptake or carbon evolution rates or what have you. All those things are just auto-calculated for you so you don’t have to do all that work again.
Ross Katz: Yeah, that makes sense. You mentioned DOE. Can you talk about the design, build, test, learn loop and how Invert facilitates that process?
Martin Tromp Permin: On a meta point, the thing that has been the most surprising to me coming into this industry is that I come from software where the design-build-test-learn loop is something that you want to get as many repetitions through as you can per unit of time, because your only cost is more or less the headcount of the engineers who are building stuff, and the velocity of your loops matters a hell of a lot more than the accuracy of them. That is not the case for most of our customers. The cost of doing business here is just incredibly high. Dealing in biology is just not cheap. We’re hoping to make it somewhat cheaper, but we can only do so much. What we’re really trying to help people do — as we talk about internally — is the following. A process optimization problem, or getting to a process that works in the first case, are two separate problems but they’re related. What you’re trying to do there is walk around a landscape and climb the highest hill — to use a computer science analogy — or move through a maze of possible experiments that’ll finally get you to the highest hill. The design-build-test-learn loop is the engine that runs us through that, that is what that algorithm is. What we ultimately care about is reducing the number of loops that you have to go around. The few levers that we have to pull on are: a) use historical data to inform where you’re starting from — can we just start you from a higher place on the map and not have you climb quite as many hills before you find the global maximum; b) can we help you learn more from each loop and not have to do redundant work where you misunderstood the learning from the last one and have to repeat; and c) can we make it such that whatever the gap is between those four steps in a design-build-test-learn loop is as tight as possible — no real time from data to analysis, from analysis to insight, from insight to next experiment design. Can we just make that as automated as possible? That’s a lot of what we’re working on right now.
Ross Katz: It’s a very Bayesian setup to the problem — you’ve got selecting the best prior you can based on the information you have, moving in the direction of that highest hill as quickly as you can using the information that you have. And what I’m also hearing is that by setting up the data in the right way you can create that automation so that the compute cycles that are happening are not necessarily people’s time cycles — it’s all facilitating the movement in the direction that these bioprocessing organizations want to go in a way that seems really efficient even if it is computationally intensive. Am I thinking about that right?
Martin Tromp Permin: I think you hit that right on the nail, and I would even double down on a point you made there, which is I think a lot of the unexplored territory here is what are your priors walking into these designs. There’s a combination of experimental results as priors if you’re not starting from scratch on a project. There’s also the concept of priors in your head — you’ve designed processes before, you know how this microbe has behaved elsewhere, or CHO or E. coli. Fundamentally what you pay for if you go to a CDMO is some mix of drug product material, process expertise, and process development, and a lot of that is just the prior of this organization having done this work a lot more than you have. It’s akin to how the management consulting or professional services industry works. A lot of what we think about internally is how do we make sure that you’re operating on the widest prior we can, or at least the best-selected prior, as we’re going into experimental design.
Ross Katz: Yeah, that’s really interesting. You’ve mentioned CDMOs a couple times, and I imagine you’re working with companies on the CDMO side but also companies that are contracting with CDMOs. Can you talk a little bit about how you work with each party and how you work in the in-between space?
Martin Tromp Permin: Why do CDMOs exist? CDMOs exist because organizations grow until the internal transaction cost matches the external transaction cost — that’s the Coasian theorem. At some point it becomes cheaper for an organization to outsource some work than to do it in-house. That’s true on CapEx, it’s true in all sorts of other ways. What you then end up often not factoring in is just how high those transaction costs are in the non-dollar sense. I’m talking about things like tech transfer in both directions, identification on failed runs, things like this. One of the things that we spend a lot of time with both CDMOs and their sponsors on is how do we reduce the risk in tech transfer in both directions. By risk I mean a) just what’s the actual labor involved here, and b) how do we make sure that whatever we perceived of what you’re trying to explain to us about your process actually is accurate. A lot of our colleagues, especially in our product division, spent a lot of years at various CDMOs, and that’s no coincidence — those people have just seen a lot more reps going through a zero-to-one process development cycle and know what it takes to reduce that time. Really, to me, what a beautiful end state would be — and I’ll say we have met this for maybe one or two smaller CDMOs right now — is that if you have a customer who’s already on Invert and the CDMO is on Invert, sharing data about a given process and that process lineage should be completely the same experience as sharing a Google Doc with a friend or colleague. There really should be no other steps involved there, except maybe adding some context where we don’t have data completeness, and obviously a bunch of security steps. And in the other direction, when my CDMO is doing work for me I should be able to see, if not real-time or live data, something pretty close to it and in as raw a format as possible. I don’t think it’s a good state of affairs to receive a PDF package of data after paying hundreds of thousands or millions of dollars for some CDMO service.
Ross Katz: I want to talk a little bit about machine learning and AI and how they’re used in Invert today. You mentioned models that are doing anomaly detection, models that are predicting yield, models that know when to scrap a lot before it’s done manufacturing so you’re saving manufacturing time and money. How consistent are the models that you’re developing across the organizations that you work with, and how do you manage the problem of training, deploying, and validating these models for different organizations?
Martin Tromp Permin: On the question of how consistent they are — there’s a chunk of problems that everybody has and everybody wants to try to solve using machine learning. And then almost every one of our customers also has some pesky little problem that nobody’s been able to solve, or many such problems, where they really want something bespoke on top. We’re not a consulting shop. We build tooling. Invert at this point comes with some broad, thin models that’ll work for an array of universal problems, and then we provide the tooling for you to go do more bespoke work for yourself. We try to not peek into what that work is, just to respect the privacy of our customers. That said, I think part of the reason we haven’t seen a ton of machine learning adoption in bioprocessing is not that the data science is all that hard. I don’t mean to trivialize the work of the many genius data scientists who work in this space — but much more so, and I think all those data scientists would parrot this, the data management and cleaning problem is so nasty that even before you get to something like the ML operations question, way before we begin talking about model selection and fine-tuning, the questions are just so large and chunky that almost no one is at the stage where they’re doing a ton of optimization on the margin of their ML models. It’s not to say that nobody is — a lot of the big pharma companies are there in many ways — but they’ve gotten there painstakingly through a lot of offshoring and a lot of in-house development spend. What we really focus on is can we just make “machine learning operations” a phrase you don’t even need to know. The end state that is best here is there really should be no stack. There should just be a tool or many tools, but there really should be no need for your average bioprocess engineer who wants to train an XGBoost model on a series of bioreactor runs to know what Kubernetes is, or model management, or model cards — all of that should just be abstracted away, as most of the rest of the industry has done. That’s how we think about it internally: build tooling.
Ross Katz: That makes sense. I heard two things: one is the data preparation and cleaning, making the data machine-learning ready for people who want to build the long tail of bespoke models needed to make the business run faster. But I also heard you say that one of the things you’re planning to bite off, or have already started to bite off, is the machine learning operations side of the equation — where models that people want to train or have trained can be layered on top of the data where the data lives. Am I thinking about that right?
Martin Tromp Permin: Right now in Invert you log in, you select a number of runs, choose a type of model or what your end goal is with that model, and it’s one-click training — we’ll separate the data out into test and training, give you all the predictions, give you the confidence intervals, and give you a contour plot of whatever you’re targeting. We’ll have a few different versions of these types of models, but that all just comes in the box at this point.
Ross Katz: Interesting. For the organizations that want to do something that Invert doesn’t support yet, do you support training it outside of Invert and then bringing it in? What does that look like?
Martin Tromp Permin: Yes, we have a Notebooks product, which is basically just a Jupyter Notebook inside of Invert that has access to all the data you have in Invert. Any model that you have elsewhere you just bring right in, and now it lives there. Your colleagues who might be less data science or ML ops savvy can then go and apply that model to their next set of data without having to bother you for your time.
Ross Katz: Yeah, that’s really interesting. Are you seeing a lot of uptake of machine learning in organizations that weren’t using machine learning previously, or is this just something people are excited about looking toward in the future as they grow?
Martin Tromp Permin: No, we definitely see uptake. What I will say is that a lot of our customers, both existing and future, are surprised once they learn that the bottleneck sits elsewhere than where they thought — meaning they go out and hire a data scientist or two and then discover a year and a half in that those very expensive individuals have spent their time just trying to figure out where data lives. I still think adoption is lacking quite intensely. At this point it is clear that the delta between where we’re at scientifically and where we are as an industry is really large, and a lot of our opportunity lies in bridging that and bringing industry to where the scientific frontier is right now. And then to push it. Even the scientific frontier of machine learning in bioprocessing is quite dramatically behind where machine learning in a lot of other areas of the world is at.
Ross Katz: Yeah, that resonates. I read in one of your blog posts that you all also have a gen AI co-pilot integrated with Invert. Could you talk a little bit about what it does, how it’s used, and how you see it being used in the wild?
Martin Tromp Permin: What we don’t want to build is Clippy for bioprocess development. However cool that would be — I think I might want that for myself as a non-bioprocess engineer — I don’t think any of our users want that, so that’s the first disqualification. We have Gen AI built into the product in a couple of places. In our Notebooks product we have a little chatbot that’ll help you write code from natural language — hey, in Python do a two-way ANOVA, and so on — that just comes out naturally and is helpful to people who are mildly technical, just enough to be dangerous like myself, but also as I understand it from my colleagues who are actually technical and have PhDs in this space, it’s just a hell of a lot faster than typing everything out yourself. And then what we’re really interested in getting a lot better at right now is how do we help you surface data that you don’t know to be looking for as you’re analyzing something, or how do we help you surface a certain analysis on some data that you also weren’t going to naturally gravitate towards. Can we get there for you before you even know what you want to do with your data, and just present it to you? I think that’s really the sweet spot.
Ross Katz: Yeah, interesting. The code generation use case is the bread and butter use case across Gen AI, but it sounds like the surfacing data you didn’t know you needed is an offline agent that’s reading over your shoulder and seeing what else is in the system. Am I thinking about that right, or how do you think about it?
Martin Tromp Permin: There’s a RAG component to this — you’re looking at a certain set of runs, and you should have something that surfaces: hey, actually there was this other set of runs three years ago at a site across the world from you, same company obviously, that did something similar, and here’s how that turned out. That’s the thing that should just exist. I’ve often been tempted to try to figure out how much redundant work is actually done within some of these large organizations, and I’m scared to even start that work because it sounds horrible. The other bit that we really want to get good at is having enough context on where you’re trying to get that we can get you there. If we know that every time Ross produces a set of runs in his Amber 250 with a cluster of process parameters that looks like all these others, Ross usually goes and does these 10 different things to that data and generates a report with it — that should just be automated. Right now we have a wonderful reporting tool that helps you create interactive reports for your colleagues, managers, and external partners with data that lives on Invert, but that should just be auto-generated based on whatever data comes in. I think that’s really the next step for us.
Ross Katz: Yeah, that’s interesting. Being where you are, with the industry purview where people are doing these consistent workflows over and over again and asking the same types of questions across organizations, it’s very much in alignment with what you were talking about earlier in terms of starting from that prior that’s further along, that knows where you’re trying to get to and helps you get there a little bit faster than you would otherwise. As we head toward the end of our conversation, could you give us some insight into what are the metrics that you use to measure the impact of Invert on the organizations that you work with?
Martin Tromp Permin: Sadly, a lot of the metrics we care the most about we can’t look at ourselves — we have to ask our users. But it’s a lot of: how much time do you spend in Excel every week, how hard is it for you to retrieve files, things like that. We often have the debate of whether we should use time-in-app as a goal. The easy way to get time-in-app really high is to make it really hard to do things, and for that reason, once a metric becomes a goal it ceases to be a good metric. The things we really care about are: what is data completeness for a given set of experiments — meaning are all our automated pipelines just running? How many times do people have to go look for, say, the Cedex data that didn’t get uploaded for reasons XYZ? Can we just make sure that always runs? We spend a lot of time looking at subtle measures around the time from when somebody starts trying to learn something from their data to when they finally close that tab having learned it and moved on — can we make that as short as possible? There’s no really easy measure to look at here, but it’s a lot of looking at click rates and how many different things did they try to get to where they want to be, and how intuitive can we make it, how can we build intuition for where you should be going given this data into the product? That’s a lot of how we think about it.
Ross Katz: Yeah, that makes sense. Are there any segments of the market in particular where you’re seeing companies that are particularly hungry for the kind of thing that Invert does?
Martin Tromp Permin: We love our CDMO customers especially, because they have the data management problem n degrees more than anyone else — they have to deal with everyone else’s data. We’re also seeing a lot of traction with cell and gene therapy companies. Those folks are designing a lot of processes that haven’t been designed much before, and they’re very eager to have all their data ducks in a row, especially as they talk to the FDA about their processes. And doubly so, their experiments are just so expensive that they really need to make sure they’re not doing redundant work. Those are two specific groups we’re seeing a lot of traction with right now, but quite frankly we work with anything that’s biologically manufactured and we’re definitely mostly pharma-focused.
Ross Katz: Yeah, that makes sense. Are there any trends in the biotech industry that you think are going to really shape the work that you do over the next three to five years?
Martin Tromp Permin: This is very much why we started the company, but the rise of biologics is just undeniable. I’m still fairly convinced — and I mentioned this at the beginning — that a few years from now the bottleneck on bringing drugs to market will be process development. I’m not saying trials will be solved or anything else, there are a lot of pesky problems there. But I’m fairly convinced that we’re making it easy enough to design new molecules that the actual number of drugs making it through the funnel is going to be multiple times higher than it is today in fairly short order. And we just don’t have the workforce to even begin to design those processes manually. That’s really what we’re doubling down on. And then another trend — contrarian these days, but I’m still incredibly bullish on all the synthetic biology work that’s happening. We’ve had a cohort of companies that have been ill-fated financially. That was a function of their timing in the cycle, and maybe some of those projects are too scientific to be a startup and too businessy to be an academic project, and they just don’t have a home right now. That said, I think all the companies going out of business in that space right now will reform into new entities, and the science doesn’t actually die. You can fairly linearly extrapolate the results from those companies, and at some point we cross a lot of price parity thresholds on a lot of commodity products, and I think that’s very interesting.
Ross Katz: So where can people go to find more about you and more about Invert?
Martin Tromp Permin: I’d much prefer that they go find something out about Invert than about myself, but the company is on invertbio.com — Invert as in turning something on its head, bio.com. I myself am on Twitter under permin, P U R R M I N, as a cat. That’s where we’re at. Go talk to us.
Ross Katz: It’s been a pleasure having you on the podcast, Martin. Really appreciate it and look forward to connecting down the line.
Martin Tromp Permin: Thanks for having me.
Jason: And that’s it for this episode of Data in Biotech. If you enjoyed the episode, please subscribe, rate, or leave a review in your podcast platform of choice. See you next time.






