Listen on
Overview
Indoor air pollution by volatile organic compounds (VOCs) makes the air inside homes and offices up to five times dirtier than outdoor air, a problem traditional air purifiers fail to address. This insidious issue affects health, productivity, and represents a significant, unmet market challenge. CorrDyn frequently sees executives grappling with how to apply advanced science to real-world problems with clear business outcomes, especially in biotech where development cycles are long and data is hard-won.
In this episode, host Ross Katz speaks with Patrick Torbey, CEO and co-founder of Neoplants, a synthetic biology startup engineering indoor plants and their microbiomes to naturally metabolize these harmful VOCs. Patrick offers a unique perspective on translating deep scientific research into a tangible consumer product. He details Neoplants’ iterative R&D process, their challenges in genetically modifying plants with previously unsequenced genomes, and their multi-stage data collection and testing methods crucial for validating novel bioengineered solutions.
Patrick also discusses how Neoplants balances ambitious scientific exploration with market strategy, providing insights into operationalizing complex biotech development. Data leaders and executives will find value in understanding how a company navigates the empiricism of biology, uses structured decision-making to manage R&D resources, and builds sophisticated measurement systems where no off-the-shelf solutions exist.
Key Takeaways
Biotech R&D demands rigorous, structured iteration to manage resource allocation and accelerate productization.
Neoplants operates a ‘great filter’ system, where projects are defined by specific KPIs and timelines. Projects that fail to meet these pre-determined metrics are deliberately terminated, freeing resources for more promising avenues. This systematic approach allows deep tech companies to iterate at a speed closer to software development, despite the inherent biological lead times.
Measuring real-world efficacy for novel bioengineered products requires purpose-built, multi-tiered testing infrastructure.
Standard VOC sensors are unreliable for specific pollutants like formaldehyde or benzene. Neoplants developed a three-stage testing process, scaling from 2-liter jars for high-throughput screening to 35-liter flow-through chambers, and finally to 160-square-foot, full-scale rooms equipped with €300K PTRMS machines for real-time, parts-per-trillion sensitivity. This approach prioritizes precision as product development matures.
The ‘last mile’ challenge for AI in biology means contextual, real-world interactions remain largely empirical, making strategic data collection paramount.
While AI can predict protein activity in a vacuum, integrating real-world factors like pH, temperature, expression levels, and complex physiological interactions within a living organism is far more difficult. Neoplants’ approach involves building initial simulations based on gathered data, like air exchange models and plant metabolism, which are refined by rigorous testing. This highlights that deep understanding of biology, not just prediction, still drives fundamental advancements.
Designing a consumer deep tech product requires starting with common user behaviors and reverse-engineering the science.
Instead of inventing a plant, Neoplants started with the popular pothos plant, which lacked thorough genomic data. Their product strategy required sequencing its entire genome, understanding its metabolome, and then engineering it to perform new functions. This focus on an existing, desired form factor guides scientific exploration toward a viable, adoptable product.
Related: CorrDyn provides specialized data engineering services for complex scientific applications and helps biotech manufacturers realize data value. We also offer data assessment to optimize R&D processes and build data reliability systems for critical operations.
Full Transcript
Jason: Welcome to Data in Biotech, a podcast from CorrDyn where we explore how companies leverage data to drive innovation in life sciences. Every two weeks, we sit down with an expert from the world of biotechnology to understand how they’re using data science to solve technical challenges, streamline operations, and further innovation in their business. Today we’re joined by Patrick Torbey, CEO and co-founder of Neoplants, a synthetic biology startup based in Paris. Patrick and his team are pioneering a bold new vision, engineering indoor plants and their microbiomes to purify the air we breathe. In this conversation, we dive into the science behind volatile organic compound degradation, the challenges of genetically modifying ornamental plants, and the role of simulations and testing in scaling bioengineered products. Patrick also shares how Neoplants balances deep tech R&D with go-to-market strategy and why he believes that synthetic biology could hold the key to solving global environmental challenges. If you’re interested in the future of plant design, indoor air quality, or applying data science to biotech R&D, you won’t want to miss this episode. Here we go.
Ross Katz: Patrick Torbey, welcome to the Data in Biotech podcast.
Patrick Torbey: Thank you very much.
Ross Katz: Well awesome. Just to kick us off, would you mind giving us an introduction to your background and what brought you here today?
Patrick Torbey: Sure. I’m the co-founder and CEO of a biotech company called Neoplants, and we’re specializing in the engineering of plants through their genetics and their microbiome to have a positive impact in people’s lives.
Ross Katz: Awesome. Why don’t you just introduce us, what are you genetically engineering plants to do currently and what are the plants that you have on the market right now?
Patrick Torbey: The main problem we’re tackling is the problem of indoor air pollution. The air that we’re currently breathing, you and I, indoors is five times more polluted than the one we can breathe outside on the streets, even if you live in a big city. And that’s due to a specific type of pollutants called volatile organic compounds, things like formaldehyde, benzene, toluene, xylene, nasty compounds emitted by everything we have indoors that makes the air quality really bad. And the problem is that there’s no good mechanical solution for that. Normal air purifiers filter out particulates matter but can’t really handle VOCs in a real sense. And so when we talk about volatile organic compound, organic means carbon and plants love carbon. So what if plants could absorb these toxic chemicals and metabolize them? And how can we create a product out of this technology? It’s not just about the science but how do we democratize this to people to put it in people’s homes. This is what we’re tackling.
Ross Katz: What are some of the biggest scientific or engineering constraints that you had to face when you were thinking through the problem of how to engineer a plant to improve indoor air quality?
Patrick Torbey: The main difference between academic research and a startup is that you are creating a product. You’re not just trying to have results. So you need to start from a plant that people already buy. And so you start with a plant that science hasn’t researched before because nobody researches the physiology of indoor plants in a deep way. We started with a selection of the plant and we decided to work on Pothos, which is a Monocox vine kind of plants that is one of the most, if not the most, sold house plants in the US. But to be able to genetically modify it, first you need to sequence the entire genome, which hadn’t been done before. So we did a de novo whole genome sequencing of Pothos. I think it was the first whole genome sequencing, at least with publicly available information, of indoor plants ever, so that was very fun. And then we start to study it. So what’s the transcriptome? What’s the proteome? What’s the metabolome? How does it actually work? And what type of enzyme do you need to introduce to bridge the gap to go from formaldehyde to sugar or from benzene to an amino acid? And from there, we say, okay, what are the available enzymes you have within nature? The second challenge here is that most of these enzymes exist in prokaryotes, so bacteria. So you need to translate these genes and make sure they work on a higher plant. Make sure they’re expressed at the right level. Make sure that they actually don’t have a byproduct and they actually do things efficiently. We also worked in parallel on the absorption of these pollutants. What if the plants had more stomata, more pores that allow for gaseous exchange so they could absorb these VOCs a bit faster? So these were the two pathways we saw how to engineer the plant, and we’re five years down this process. It took a lot of time to understand and then engineer it. But now the plants are now growing in the lab, which is great. Same thing with the aesthetic feature of the plant. We haven’t announced it yet but we also had to look within nature — how can we change either the color or the shape or the fragrance, what could we change on that plant and how to make it radically different. So that was also something not explored by the literature. You find hints here and there, experiments that weren’t supposed to do something that did something interesting with the physiology of the plant, and just explore. So there’s a lot of creativity involved, a lot of luck, because this is empirical science and there’s very little data to go on. It was a big chunk of our time gathering this data, this knowledge, understanding that plant and then engineering it. And so these were the challenges, or I would call them the opportunities, the pathways that are open for us to create this type of plant. And the same pathways are also open — and actually the more we delve into it, the more we thought they were even more interesting — which is the microbiome. Interestingly, we thought that the microbiome would have a lesser impact than the genetic engineering of the plant itself. And what we saw is that it could be so powerful, so much more powerful than what we thought, just the microbiome alone. So that was also a journey of discovery because nobody tried to do this type of stuff before.
Ross Katz: That’s awesome. What I’d love to do is dig into the specific degradation pathways that you’re targeting with the plant and the microbiome, and then talk a little bit about your optimization process — how you gather data and how you got down the path of understanding better what was going on. My understanding is that you’ve got two degradation pathways: formaldehyde and BTEX. And your goal is to basically take as much of that out of the air as possible — I think you refer to it as flux optimization, getting as much of that out of the air and into the plant itself or into the plant microbiome as possible. What I’m interested in is how did you come to the conclusions about how you were going to do the degradation? Were you looking at the enzymes that were out there and using those to guide what was possible to connect these pieces together in order to end up with something consumable by the plant or the plant microbiome? Or did you have in mind that this was the path that had to be taken from a chemical perspective and then go out looking for the enzymes that were going to enable the processes inside of the plant or the microbiome? Or was it some combination of the two?
Patrick Torbey: We started with: we need to go from benzene, let’s say, to something that the plant can degrade. So we looked at all the available datasets online, enzymology datasets. What are the enzymes that can turn benzene into anything? So you get to this first circle. And then do we have anything that the plant can use there? Not really. So we had to go to: okay, from all these molecules that can be produced by benzene to something else by anything in nature, whatever the organism. With one enzyme we can’t get there, but with two enzymes we have a few different candidates, and catechol for us was the most interesting. So you get from: okay, we could go from benzene to catechol and it’s a two-enzyme step. But there’s a lot of different enzymes on very different organisms that can get you to phenol, get you to catechol. Some actually have one single enzyme that can do the two steps. So you test those. You say, okay, let’s have all these different types of enzymes from a lot of different organisms — and the more the protein sequence differs, the better because you’ll have a diversity of results. Test all of those inside plants and then see what works, what doesn’t, through analytical chemistry. So you can trace where the benzene goes and whether there’s a bottleneck at step one and step two. And then you get to making sure this is as smooth and the flow is as easy as possible. So you can then take, okay, these are the best enzymes. Is there a bottleneck there? Do I have to do protein engineering? In our case we also did that. We also found different enzymes where they didn’t even take benzene, but they took something that looks a lot like benzene, and we did protein engineering. Because either you start with basically bacterial enzymes that do exactly what you want to do, or you start with a plant enzyme that doesn’t take benzene as a substrate but takes something that’s very, very similar. This is still confidential information so I can’t go into the detail. But you change the specificity of that enzyme to allow it to take benzene as a substrate and do the chemical reaction it was supposed to do with its original substrate. And you can do that through protein evolution or site-directed mutagenesis. You can either go through directed evolution — which we’ve done with some — but the most promising approach and the ones that showed some interesting results is actually designing the protein itself. So you go into AlphaFold and you check, okay, what if I change this and that? How big will the pocket, the active site of the enzyme, how big will it be? What is the substitution I need to make to have the types of interactions that I want so that it doesn’t take the first substrate but it takes benzene. And this was actually quite successful.
Ross Katz: You’re going down all of these different experimental paths simultaneously. Can you give us some insight into what data you’re collecting as part of the design-build-test-learn loop to stay informed about all of the paths that you’re going down and be really smart about the resources that you’re deploying and the experiments that you’re running to get yourself to the end goal as efficiently as you can?
Patrick Torbey: Putting in place an R&D process is something that I love to do. We went through a lot of different ways to do parallel work, to iterate fast, trying to go at the speed of software with biotechnology, with plant biotechnology, which is not easy to do. The best system that we found that we currently use is working with projects that are defined in time and scope and objectives before we start the project. We have within the team something like 20 projects running at different stages of maturity. For the early-stage projects, we say, okay, we’re going to spend four months to get to this first milestone with this team and this percentage of time of each of the individual contributors, these resources. After the four months, we have what we call a great filter. This is where we try to kill the project. After four months, we’ve already decided these are the KPIs, this is the advancement that we needed to have, the exact results that we needed. If we don’t pass the results, we kill the project, which is something that we actually celebrate because it opens up resources to do another project. You keep iterating this way, trying to make sure that the resources are as well distributed as possible so you keep a flexibility but you know exactly at which stage you take a decision. Since it’s very much an empirical journey, you let things go their way and see what’s the most promising path through these great filters. So it goes through great filter one, then great filter two, and then the project is done. For example, a protein engineering project where you need to have a specific Km, Kcat, a specific activity. If you hit these numbers, you move forward and you get to a higher stage of development. You created this protein, now it’s about: can you make this protein work inside your plant? And then there’s another filter afterwards — now can you create a stably genetically engineered plant that produces this protein that actually absorbs these pollutants faster? And then the last stage: is it fast enough to have a real impact in real-life conditions? And so you have a lot of projects that die at the first great filter, which is what we want. We really try to kill the projects. We incentivize people to kill projects, actually. At the end, the only projects that go through and end up being ready for commercialization are the ones that survive. A selection of the fittest in terms of projects.
Ross Katz: Interesting. You gave the example of a protein engineering problem. Can you give some examples of some of the other classes of projects and the milestones that you had inside them to understand how you were progressing across all of these different avenues?
Patrick Torbey: Inside the Power Drops, there are two strains of bacteria. One strain is a Methylobacterium that can absorb formaldehyde, the other is a Pseudomonas that can absorb BTEX — benzene, toluene, ethylbenzene, xylene. These are naturally occurring bacteria that like to live close to plants, either in the rhizosphere or in the phyllosphere, forming a natural microbiome. The entire microbiome project: first step is let’s find organisms that aren’t harmful for humans or animals, that can live in symbiosis with the plant, and that can use BTEX as a sole carbon source, meaning they can grow with only these pollutants. That was the first great filter. If we couldn’t find these organisms, the project would die. And we found some of these organisms — one of them, Pseudomonas, is what we use currently. The second stage: now if you want to have an impact at scale, we run some calculations with data that we generate in-house. If you want it to have any impact in real life, you need to have this efficiency, this remediation speed. So there’s another great filter, great filter two: now let’s try different things, how to make it efficient. Let’s try directed evolution — manually. If you put the bacteria into media that are more and more concentrated in benzene, the ones that survive all these different steps should be the most efficient at using benzene as a carbon source. That’s one path. The other path was also directed evolution but in another way. You take the same organism and you try to evolve it not by putting it to more and more pollution levels and killing most of the bacteria, but you make it constantly grow on a low amount of benzene so you select for efficiency. How fast can it consume a small amount of benzene, instead of how do you resist to high amounts of benzene? That’s more about efficiency of use of benzene. Then you have: let’s use transgenesis. Let’s insert transgenes from other organisms to add to this capacity. And then there’s another angle — maybe it’s not about having the pathway be more efficient but something else within the bacteria. What if you could make sure that once the benzene is inside the bacteria, it never goes out? So it has to metabolize it. That’s not directed evolution, that’s a specifically targeted mutation inside the genome of the bacteria to knock out the genes that code for the benzene to exit the bacteria. It has to deal with it, it has to deglutinate. So this is another great filter we had to go through — which one of these projects is best to move forward. We found the best way, moved forward, and ended up with a strain that is orders of magnitude more efficient at deep-polluting this BTEX — high enough efficiency to have a real impact in your bedroom, for example, with one single plant. This is another example of different stages of projects that had to die for one of them to succeed. And then obviously it doesn’t create a product — you need to add these other bacteria, Methylobacterium, that went through the same exact process, put them together, make sure that they work well together, then make sure they can form a stable microbiome, at least stable enough that it doesn’t die within one or two days. And then you have something that can become a product.
Ross Katz: Throughout this entire process of experimentation, you’re keeping your eye on the end goal of having a plant with a microbiome that can metabolize these VOCs. What are some of the measurement methods or North Star metrics that you’re using to understand whether you’re moving in the direction of having the product that you’re targeting, even as you’re going down all of these avenues and hitting these different filters?
Patrick Torbey: Testing the efficiency of the product was the most difficult thing we had to unlock at Neoplants, because nobody has done that before. All the VOC sensors that you have in your home don’t work. They react to ethanol and not formaldehyde. If you put perfume on, you’ll see the total VOC levels of your sensor shoot up, but if you actually put pure formaldehyde at super high concentration, it doesn’t even react. Same thing with benzene — they don’t detect benzene, they detect other things. Basically the total VOC sensors right now can only tell you the air seems stuffy. That doesn’t mean there are harmful pollutants. It can be ethanol, which is not really harmful at the levels it detects. We needed to develop our own testing methods and we needed to be high-throughput and extremely sensitive, which is very, very difficult to do. We don’t have one testing method, we have many. On the microbiome side, the first testing setup we developed: obviously there’s the molecular testing making sure that the pathway is working — this is level zero testing. But once you have the strain, you want to make sure that it actually absorbs these pollutants. Outside of analytical chemistry and enzymology, you have the bacteria — we put it inside the soil of a small plant that you put inside a 2-liter jar. We’ve found sensors that are not very specific but very sensitive to benzene, for example. You put these sensors on the lid of the jar. You can close the lid and insert specifically benzene into the jar. Even if the sensor is not specific to benzene, since you know you only introduced benzene into this jar, everything it detects is benzene. And we have something like 20 of these jars in parallel. So you can test 20 different concentrations or strains or whatever you want in parallel with all the controls that you want. This is what we call Level 1 test — you have these small plants inside 2-liter jars and see what works, what doesn’t. It gives you a good sense and you can compare them between each other. But the concentration levels are way too high — much higher than what you can find typically indoors. It’s also a static chamber, completely sealed, only 2 liters. So then you have Level 2 testing. You take the best results from Level 1 and go to Level 2. Level 2 is a 35-liter chamber where you flow through polluted air. We have two of them running in parallel, control and experiment. They’re linked to GC-MS — gas chromatography linked to mass spectrometry — to be much more precise. But you can only do these measurements every 30 minutes or one hour, otherwise technicians go crazy. It’s active sampling, not passive data collection. You can’t do this at scale, but you’re very precise. You can actually now say, okay, this solution — if you flow through this amount of air, you have this efficiency. You start to have a good idea and you can plug in some numbers in your simulation, because we’ve also created simulations on how this can work. And the last level of testing: two chambers that are 160 square feet that we created, built in stainless steel so there’s no interaction with the walls. In these chambers we can control humidity, temperature, pressure, light source, amount of pollution, and we can track everything happening in real time at parts per trillion sensitivity. So extremely sensitive with a machine called PTR-MS — I think it’s Proton-Transfer-Resonance Mass Spectrometry. This machine alone is like 300k Euro. But it’s extremely sensitive and works in real time. We link this machine to these two rooms and put inside them Ikea furniture or any other furniture that does or doesn’t emit VOCs, walls that we paint with certain paint. You actually create an actual bedroom with the same conditions, but you control everything and you can test a normal plant or a genetically modified plant or a plant with or without the Power Drops. You have extremely precise data with this. But one experiment takes a month to run there. So it’s not high throughput. We needed to go through these three levels of tests — from 2 liters to 35 to a full bedroom-size test — where every step along the way you lose throughput but you gain precision.
Ross Katz: Very interesting. You mentioned the simulation process that you use to understand the dynamics of what’s happening in these tests. Can you give us some insight into what that simulation looks like and how the experiments that you’re running help you to understand the simulation and vice versa?
Patrick Torbey: The simulation is very rough at first and then you can add more and more complexity as you go along. At first, the first thing we’ve done is: okay, what is the typical airflow within a bedroom, what’s the typical concentration within a bedroom, how much do we think a plant could absorb? It was very rough, done with relatively simple equations. But the more you gather data — specifically the more you set up your testing methodologies — and the more data you have on wood emission levels, how fast do they in the world, etc., you can create a more and more precise simulation. It’s something that is going to be more and more powerful the further along this process we go. We’re really at the beginning because we have to invent or adapt a lot of things from different areas — from the indoor air quality field to microbiology to plant genetic engineering. You build this simulation in different parts little by little and then put them together. You try to see: does the simulation correspond to what you actually see? And it never does. So you have to go back to the drawing board and see what’s wrong. But at some point you get to a point where it actually is useful to have the simulation. It’s precise enough to give you a rough idea of some paths that you can say, okay, for sure this is not going to work. Even if you can’t predict exactly what is going to work, you can say with high degree of confidence that it’s not worth running these tests because this is not going to be useful.
Ross Katz: Can you give us an idea of what level the simulation is at, what are some of the variables that go into it, and some of the trade-offs that the simulation might demonstrate for you?
Patrick Torbey: Different bricks that come together — this is how we think about it, because you can’t have one whole simulation. The first thing we’ve done is air exchange within a room. We’ve done that with an external partner and you can have a simulation of the airflow that happens within a typical room. This is one brick that doesn’t require any data from our side. Another brick: once you get some data on the efficiency at which the bacteria absorb these pollutants, you simulate a well that absorbs — like an attraction force for VOCs — that’s the plant. And you can plug in numbers for this well of VOCs according to the data you have from the 35-liter chamber or the actual full-size bedroom. You try to refine them as much as you can. Another big part of the simulation, but one not yet completely linked, is the plant itself. As we go along our genetic engineering of the plant, we created a virtual plant. We assembled the genome, but we also had the transcriptome, the metabolome, a bit of the proteome. We have a simulation of the metabolism of the plant — a virtual plant that has as much data as we can have, studied only inside our computer, that allows you to say: okay, this enzyme is probably missing, so you probably need to bridge the gap with this additional enzyme. We’re not at the quantitative level yet, but once we gather enough data we can plug in numbers and then plug that simulation into the indoor air quality simulation. Something else we’re trying to build and give to people is actually this simulation part in and of itself. We’re generating a lot of data on how furniture affects indoor air quality, and nobody is generating this data — we talk to the people at Dyson or at the EPA, and none of those people have actually studied at this level of sensitivity the dynamics of VOCs being emitted by different types of woods, plywoods, paints, etc. We try to gather as much data as possible and then create an indoor air quality simulator that people can actually use to predict the level of VOCs they have inside their bedroom, since there are no good sensors available today. These are the different pieces we have when we talk about simulations — how we try to predict, first, the level of pollution in people’s homes, and then how our product affects that level of pollution.
Ross Katz: That’s amazing. We’ve gone through the experimentation process, some of the assays you’re using to measure the outcomes you’re getting, and the simulation component as well. I’m wondering, from your experience starting from a very low data environment and slowly building up the data you need to characterize the system you’re creating or trying to influence inside people’s rooms, inside people’s homes — what have you learned about the limitations of data science or computational approaches in helping you along the path of solving the problem you’re trying to solve?
Patrick Torbey: When I started this journey I was very much a wide-eyed, young, naive little scientist thinking we could generate fantastic levels of data and then build our models and then predict things and have the computer tell us what to do. It doesn’t work like that at all. It’s very much a step-by-step process, especially if you’re the only people researching exactly what you’re doing. You can have a bit of resources here and there, but you need to generate these data. The first thing that I learned is: you need to have a very, very precise test with very precise outcomes. Setting up this test and the right KPIs is the most difficult part, by far. Setting up the test and making sure the test is the right one is the most difficult thing. Running the test is actually the easy part — it’s just about throughput and making sure we’re creative enough to find the solution that actually works. If I had advice to give to my younger self, it’s: focus the first part of your R&D on just setting up the test. Don’t even try to create the product, just set up the test. Then once you know the tests are working, you start creating the product. You’ll go much faster. And the results of these tests need to be in a format that’s computer-friendly, even when you start. One example: since nobody has really tried to do genetic engineering of that plant before, we need to understand — once you find the genes you want — what can drive their expression? There’s a huge amount of work on which promoters work on that plant, which terminators work, and the combination of the two. It didn’t express high enough, so how can we increase the expression level? Once you explore all these possibilities and create even synthetic promoters and all these things, how do you test the activity? What’s the actual test? Do you need to wait for the entire stable plants to grow to test the expression level? That would take a year. And you can’t test it on another plant because it has completely different metabolism and genetics. So what test can you run that will last a week or two and give you good enough results with high enough throughput to predict, okay, once the plant is stably transformed with these promoters and terminators and this gene, it’s this level of expression? We need to establish: what’s the fluorophore? What’s the protocol to do what we call agro-infiltration — to transform, in a transient way, the genetics inside the leaf? How can you normalize these things? There’s a lot to control for, because if you transform more cells and you just look at the fluorescence level, you’ll have a high fluorescence level but each cell has very, very low expression. If you have one cell that expresses a lot, then your problem is not the expression level — it’s the fact that you can’t transform the other cells around it. So you have your fluorescence outcome, but even then you need to say: I also need to count the number of cells that produce this fluorescence, which gives you an idea of the strength of the promoter and terminators. And then you need to say, okay, now I’m going to replace my fluorophore with my gene of interest. Am I sure that the promoters and terminators are actually working when I replace the gene? Does the gene have any effect on the activity of these regulators? Literature was very ambivalent about it. We found out that yes. So what kind of test can you run if there’s a difference — if you change the actual enzyme, your reporter fluorophore doesn’t work as well? This is the type of challenge where biology needs data science: how can we have the best test possible that’s reproducible enough to give you high-quality data to really inform what’s the next step in your R&D journey? The only way to craft it is to go through it and see what doesn’t work, what does work. At the end, once you’ve generated a lot of data, you can extract a few learnings from it, a few general rules, a few general principles, and be more refined in your approach. But you can only deduce these rules of engagement of the plant’s genetics if you have the standardized test and enough high-quality data — even if you don’t think about simulation level, just enough data to lead you in the right direction or to remove a path you know will not work. It was really a step-by-step process.
Ross Katz: That makes a lot of sense and I think what you’re saying gets lost in a lot of the conversation about foundation models for biology as compared to foundation models for language — which is that, obviously, if you have data of a certain scale or there’s zero marginal cost to produce new data, then noisy data can teach you a lot about the world. But when you’re in an environment where there is marginal cost to creating new data points in both time and money, the design of that particular measurement system, and ensuring that you’re measuring the thing you need to measure on the time frame that you need to measure it, becomes so much more critical. I really enjoyed hearing you describe that.
Patrick Torbey: The most difficult challenge that AI will face is putting the results into their right context. Let’s say you want AI that can predict protein activity. We can predict protein activity with AlphaFold — you can do some predictions. But that’s in a vacuum, that’s in vitro. When you talk about a full organism, what’s the effect of pH and all the other things on the enzymes? What’s the effect of the expression level? How does this enzyme react to the transcription machinery of the cell to produce this enzyme? Is this enzyme stable in this specific context? Where do you want to put the enzyme — in the cytosol, in the mitochondria, in the chloroplast, in the nucleus? And then: you modeled the effect of your enzyme in one cell. Does every single cell express this enzyme, and even if it does, does it react with other stuff from other cells? What are the different organs of the plant and how will they react with this stable transformation? And then, now that you’ve predicted the expression level and even the phenotype of the plant, what’s the effect on the actual function you want to add? It’s not a plant in a vacuum — it’s a plant at this temperature, with this amount of sunlight, with this effect on its physiology, that is watered this amount of time in people’s homes. If we want to talk about the dream of having an AI that can predict: you want a plant that purifies the air from formaldehyde, here are the pathways, these are the enzymes — you need to integrate all this different context to predict, yes, this is the way to go. And this is just the context that we know about as scientists, which is I don’t know, 1% of how the organism actually works. At first you need to understand biology, then you need to model it, and then you can predict things from it. And understanding is basic research and we’re very far from completely understanding what’s going on. I believe that AI models in general will be very, very efficient, and are already very efficient, at very specific tasks that are very far from solving biology and predicting what type of changes you can introduce inside organisms to have this type of impact. This is still a very empirical-led research. But I do believe that we’ll get there. I do believe that AI is going to change everything when it comes to biology, but the challenges are insanely high.
Ross Katz: I’m interested in how you think about avoiding becoming an insular R&D-focused organization — balancing the engineering performance and biological stability, the research-oriented outcomes of what you’re trying to accomplish, with the product-focused outcomes, the manufacturability, the scalability of what you’re trying to do.
Patrick Torbey: It’s about creating a certain culture within the company. And it’s little things. For example, we have a monthly company review where we talk to the entire team, business and R&D, about what’s happening, updating everyone. It can also go bottom up. We involved almost all our scientists in customer support, so they can actually talk to the customer using the product. It’s super motivating for the R&D team to know that our product is out there, people are asking about it, and I can help them use the product. It’s not something we’re doing at scale, but it’s something we intentionally did at least at first to bridge the gap from the bench to the bedroom. Everybody in the team thinks in a product way. What’s going to be the impact? Most people in R&D don’t even think about the products and even less about the customer experience. So we try to do both. It’s really cool you found your fantastic little organism that does what you want. But can this organism live with indoor plants? If not, what are we even talking about? You need to have this culture within the R&D team of always thinking about the end product and even more the user experience. This is how you bridge the gap. Once you have your first product on the market, you want to have feedback and run it back directly to the R&D team, so you actually get the R&D team to ask for feedback directly. We do this at scale in a more data-oriented way with our marketing team, but this direct link is important. It’s these types of little things that you can inject into the culture of your company, the more it faces the market. It’s very useful because then R&D people can say: I’ve heard that consumers want to do this, and I actually can do it. Sometimes the marketing team hears the customer saying: I would love for the Power Drops to smell good. And the marketing team says, but this is not what we do — we do indoor air quality, air depollution, we’re not interested in fragrances. But if the R&D person hears it and says: that’s easy, I can just introduce one gene that produces a smell, why not do that? So we can then have a discussion between R&D and marketing saying, actually this is not very difficult to do in the next iteration of the product. R&D needs to be in constant contact with the market. This is how you build a product creation machine and not just an R&D institution.
Ross Katz: Amazing. What do you say to people who have some discomfort with the idea that you’re genetically modifying these plants, even if they’re in support of the reason you’re doing it?
Patrick Torbey: There are a lot of jokes I could do — like, if they have dogs, do they know that their dogs are genetically modified wolves? But in a more serious way, I think the most important thing is to try to educate the public about what genetic modification is. For me it’s very simple: we’ve been genetically modifying organisms since civilization existed. The first genetic modification technology we developed is called domestication. We domesticated crops and dogs and sheep and cows and a lot of different organisms that are not supposed to be the way they are. Cows are not supposed to be these big lumbering sacks of meat that are super docile. Maize corn didn’t used to look like this. Tomatoes are monsters compared to their original form. It took us 10,000 years to create these organisms, but they are definitely heavily genetically modified. People don’t think about it as genetic modification because it’s not artificial. But 100 years ago we introduced artificial genetic modification through what we call mutagenesis — random mutagenesis. You put UV on your tomato seeds, let’s say, because you want to create a tomato that’s a bit more red. You generate millions of different mutations and then you select the seeds that produce the most red tomato. You have millions of mutations inside your genome, but it’s not considered GMO by any country in the world. Why? Because if you outlaw this, you outlaw agriculture. Every single seed in the world is produced using this method of random mutation and then selection of what you want. That’s the second level, 100 years ago. Now, 40 to 50 years ago, we developed a method called transgenesis — instead of doing these random mutations, you can insert one single gene. Now the big question is: what is the function of that gene? This is where you realize whether a GMO is good or bad. GMO is like a technology. Is metallurgy good or bad? Well, it depends. If you’re producing swords to kill your neighbors, it’s bad. If you’re producing casseroles to feed people, that’s good. GMO is not good or bad in and of itself. But the first applications of GMOs were to put herbicide resistance into maize and put so many herbicides in the fields that the only plant that could survive was your genetically modified plant — which is now full of herbicides and pesticides, destroying the ecosystem, the soil health, and it’s bad for your health because you’re eating herbicides and pesticides. But the fact it’s GMO in and of itself is not the problem. It’s the application of these herbicides and pesticides. And it’s too bad that the first applications were those that are in service to big agriculture companies selling more herbicides and pesticides. That’s their revenue model. That’s called transgenesis. But if you insert a gene that is not a resistance gene to these herbicides but a gene that purifies the air, why would it be bad? These are all naturally occurring genes that I introduced from one organism to another. You can call it Frankenstein, but when you create a light bulb, it’s not supposed to be there either. At some point you need to accept: we have two paths in front of us. Either we work with nature or we’re destroying it. And the way we’ve been doing it is destroying nature instead of leveraging its power to find sustainable solutions. Nature is sustainable by nature. Why not use it? And then the last type of genetic modification is CRISPR-Cas9, genome editing — cutting bits and pieces of the genome, putting them back together, changing small things. People are much friendlier to that approach because you’re not introducing something exogenous to the organism. But even there, the current European regulation says this is a GMO. Even if you introduce one single deletion, that’s a GMO, whereas the millions of mutations you’ve done through UV — that’s not GMO. To answer your question, what would I say to people who are iffy about GMOs: learn about genetics. It’s beautiful, it’s elegant, it has so many applications that you don’t even think about. All the insulin that people have right now is produced by a GMO bacteria. That’s the only way we have to produce insulin — to introduce the insulin gene inside a bacteria. That saves dozens of millions of people’s lives every year. It’s a type of technology with bad and good applications. It can either save the world or possibly destroy it. We need to have the right discussion. It’s not about whether GMO is good or bad, it’s about what applications we want to promote and what applications we want to ban.
Ross Katz: Not what is the tool, but how is the tool being used. You’ve been very generous with your time today, Patrick, I really appreciate it. Before we let you go, where can people go to find out more about your work or connect with you online?
Patrick Torbey: You can go on Neoplants.com. This is where we sell our products. If you want to support what we do, go there, try it out for yourself — I hope you’ll feel much healthier, especially if you have asthma or allergies. This could be very beneficial and will help us continue to pour these resources into the deep tech R&D projects. You can also check out the science we have. We created a white paper that explains everything about the problem, the technology that we use, the solution, the raw data — we try to be as transparent as we can about the science behind what we do and how we test it.
Ross Katz: You’ve got one happy customer over here. I’ve got my Neoplant sitting here right next to me. We renovated our house and had to paint the interior and so getting VOCs out of the air is very close to my heart. Love the work that you’re doing at Neoplants. Thank you so much for joining us today, Patrick. Really appreciate it and look forward to connecting down the line.
Patrick Torbey: Thank you very much for inviting me. It was a very fascinating discussion.
Jason: And that’s it for this episode of Data in Biotech. If you enjoyed the episode, please subscribe, rate, or leave a review in your podcast platform of choice. See you next time.






