Listen on
Overview
Bringing a novel biomaterial from concept to industrial scale in less than two years is an extraordinary feat in synthetic biology. Achieving this demands more than scientific breakthroughs; it requires a deeply integrated, cloud-native data platform from day zero. In this episode, host Ross Katz speaks with Pierre Salvy, CTO, and Lucile Bonnin, R&D Lead at Cambrium, who are pioneering the design of sustainable proteins, like their vegan collagen, Novacol, for diverse biomaterial applications. They discuss the profound challenges of integrating disparate lab equipment, scaling processes across four orders of magnitude (from 1-liter tanks to 10-ton fermenters), and de-risking manufacturing—all powered by a sophisticated data strategy.
Pierre and Lucile detail their journey, from initially simulating lab operations to building a protein programming language, and then incorporating advanced generative AI models. They explain how these tools, combined with a commitment to early data capture and team collaboration, enable rapid protein design and biomanufacturing scale-up. This conversation offers executives and data leaders a rare look into how advanced biotech companies are structuring their data infrastructure to accelerate product development, reduce costs, and maintain competitive advantage in a complex scientific domain.
They unpack how data guides every step—from identifying market needs and computationally designing proteins to optimizing microbial ‘microfactories’ and ensuring quality control during scale-up. Listeners will gain concrete insights into the data systems required to transition from R&D to commercial production with speed and reliability, highlighting the strategic choices that enable such rapid market entry.
Key Takeaways
Rapid biomanufacturing scale-up requires a ‘digital twin’ of lab operations, not just experimental simulation.
Cambrium achieved industrial scale in under two years for its first product by first simulating lab operations and data flows—a ‘digital twin’ of throughput and failing points—rather than solely predicting experimental outcomes. This foresight allowed them to acquire the right equipment and anticipate data generation volumes, directly impacting their ability to launch quickly without common delays.
Cloud-native data infrastructure from day one is non-negotiable for competitive speed in biotech R&D and manufacturing.
From the very beginning, Cambrium implemented a cloud-native approach, storing all data in the cloud without local files. This foundational decision enabled rapid data organization, analysis, and iteration, significantly reducing the future cost and complexity of retrofitting data pipelines that often plague older organizations with unstructured or siloed data.
Scaling across vast orders of magnitude (e.g., 1L to 10T) hinges on correlating small-scale lab data with industrial production data.
Cambrium successfully correlated fermentation data from 1-liter lab tanks to 10-ton industrial fermenters. This ability to maintain process consistency across such a massive scale difference, achieved in just six months for their first product, is critical for de-risking tech transfer and optimizing production parameters.
Effective protein design integrates both precise ‘programming languages’ and advanced generative AI models.
Cambrium uses a unique ‘protein programming language’ for combinatorial design, allowing specific functional elements to be precisely arranged. This structured approach is complemented by generative AI models (LLMs, diffusion models) that explore broader structural possibilities. This combination balances controlled design with novel exploration, mitigating the risk of AI ‘hallucinations’ while still benefiting from wide-ranging data insights.
Related: CorrDyn provides data engineering and machine learning expertise, particularly in biotech and life sciences. Learn more about how biotech manufacturers gain data value.
Full Transcript
Jason: Hi everyone, this is Jason, producer of Data in Biotech. Before we get started, I wanted to let you know about our latest white paper. It’s a comprehensive guide to implementing machine learning models in biotech manufacturing. It’s a complete overview of all the potential problems of ML adoption and, more importantly, how to solve them. To download it, simply visit connect.corrdyn.com/biotech-ml. We’ve also dropped the link in the show notes of this episode. Okay, let’s get into it. Welcome to Data in Biotech, a podcast from CorrDyn where we explore how companies leverage data to drive innovation in life sciences. Every two weeks, we sit down with an expert from the world of biotechnology to understand how they are using data science to solve technical challenges, streamline operations, and further innovation in their business. This week, we sat down with Pierre Salvy and Lucile Bonnin from Cambrium, developers of NovaColl, the first micro-molecular and skin-identical collagen, designed specifically for highly efficacious skin care formulations. During the interview, we discussed the process of protein design, specifically touching on the role of ML models across this process, the challenges of integrating data in the lab, and the development of a protein programming language. Lucile and Pierre also unpack data digitalization in the lab, protein design and vector databases, designing the manufacturing process, and de-risking quality control in scale-ups. Here we go.
Ross Katz: Pierre Salvy and Lucile Bonnin, welcome to the Data in Biotech podcast.
Pierre Salvy: Thanks, Ross. Nice to meet you.
Ross Katz: Nice to meet you too. And just to kick us off, could you each give us an introduction to your career to date? Lucile, why don’t you go first?
Lucile Bonnin: Yeah sure, hi Ross. Thank you so much for the invitation. I’m a chemist originally coming from France, from the West Coast, now based in Berlin. I worked in big companies before in personal care and food and beverage industry before doing my PhD in physical and analytical chemistry. I developed a chemical platform for stable hair coloration — really fun with a lot of scales and big tanks. I then jumped into predictive analytics for chemistry, working in lab for Henkel, which brought me to think about data in a different way, which I brought to my R&D background, research and development background. From this point on, I decided never again to do R&D without using data. I think you have tools in your lab and you just should use them. I brought them to consulting for a while and now I’m working for Cambrium where the synergy of R&D and data is perfect and I’m leading the R&D team here for three years now.
Ross Katz: Awesome. Pierre?
Pierre Salvy: Thanks, Ross. Thanks again for the invite. Super happy to be talking about the exciting stuff we do here at Cambrium today. Historically, I’m a big fat nerd. I studied engineering, applied math to a bunch of different domains, from material science to microelectronics to nuclear physics. I actually almost got a degree in nuclear physics, but then I went to the US West Coast and decided to make jet fuel by engineering genetically beer yeast. That was really fun and I realized I really like biotech. I got a PhD in the field and then the founders at Cambrium, Lucile, and that’s basically the start-up, when we built the company that we’re going to tell you about.
Ross Katz: Very cool. And so, tell me about Cambrium. What’s the company’s mission and how does it accomplish it?
Pierre Salvy: What we do is we use generative AI to design proteins that will be the building blocks of the sustainable biomaterials of tomorrow. The general idea is most of the stuff that surrounds us, like the table on which this computer is, the computer in itself, is based on a lot of unsustainable material, right? Coming from petroleum sources, from animal exploitation and so on. There are actually a lot of material problems that had been solved by nature for the last two billion years of evolution, and it would be a missed opportunity not to take from these robust materials that make up your hair, your skin, your bones and try to find applications to that. That was the original idea.
Ross Katz: Yeah, so it’s a biomaterials company. You’re focused on developing sustainable biomaterials. My understanding is that you have a biomaterial that you’ve already begun to bring to market. Is that right? Can you talk a little bit about what that is?
Lucile Bonnin: Yeah. So NovaColl is our first product. It’s a micro-molecular and skin-identical collagen. So what does that mean? It’s basically a vegan collagen. So a lot of applications in the personal care and beauty care industry come from animals, unfortunately. So it’s not cool for the animals and also absolutely not sustainable. So our first product is basically tackling this issue. Pipeline is developing in other directions, but we launched last year and that was super exciting. So it means that for a three years, back in the day, three-years company, we already were at scale, at industrial scale to be able to commercialize our first product, which I think for a biotech company is pretty impressive. But you have to be that fast nowadays, yeah, to keep market traction.
Pierre Salvy: From idea to selling it took less than three years, which is absolutely unheard of in synthetic biology. But what was the status quo before? How do you get collagen before? Collagen was a byproduct of the meat industry. Your cow goes to the slaughterhouse, gets slaughtered, and then you have a lot of remainders. What people were doing is boiling this, extracting the collagen back, and then selling that to ingredient manufacturers who put them in creams. People were literally putting dead animals on their skin. The moment you tell them, they’re like, oh well, that’s gross, I’d rather buy the vegan collagen, which is good for us. That’s also how we chose this initial product — the market traction was already there so strong that for us it was a no-brainer. And I’d like to talk also a little bit about why cosmetics, because we do material science and we start with cosmetic. First, your skin is a very complicated material, but second, cosmetics is also a great market to enter when you’re a techy company, because it’s a low volume, high price elasticity, high drive for innovation and a lot of permeability of the market for new players. Not the case in healthcare, not the case in commodities. In that respect, it’s really a place to start for many companies.
Ross Katz: As I heard the story of NovaColl, one of the things I was wondering was did you start out with the idea that Cambrium was going to enter the cosmetics industry and that collagen was a ripe target to replace with biomaterials? Or did you develop the data and software platform to explore the protein space and then select collagen and the cosmetics space and go from there? How did the decision to use this as your first target evolve?
Lucile Bonnin: I think every target we’ve been addressing was pushed very hard and also in a very logical framework by the business development first. But the pipeline that we used in terms of data was already very abundant from what we could do, especially for the personal care industry. What Pierre is developing with all this data platform is very orientated to create the product and design it in its specificity, but when it comes to business, this is actually something where we got in touch with customers very, very early. This was early days, because now we are building a data set which is proprietary to us with our own lab, so it’s becoming richer and we have more freedom into actually guiding the business development into the right use cases. It becomes more and more interesting.
Pierre Salvy: Yeah, it was concurrent. We screened around 300 use cases — protein-market fit, if you will — which involved incredible conversations with people that were actually welding a rocket in the garage somewhere, really cool. But at some point you know you’re going to build a protein, so you know what are going to be the building blocks you need in your data infrastructure, but also your general software infrastructure. We’re going to talk about the toolkit that we built. Whatever protein you’re going to make, there are some common elements. These common elements we were already building at the time, and then once we decided to go for collagen, there were some more specific elements that we just fine-tuned for that.
Ross Katz: Since many of our listeners are new to protein design, it would be useful if we could talk about the problem at a high level. You have an idea of a target market that you’d like to enter with a protein that you could potentially design. What are the components of the process or the different problems that need to be solved between knowing what market you want to get into and designing that first protein that can actually solve the problem you’re setting out to solve?
Pierre Salvy: I’m going to use the space glue example because we’re not going to make space glue, so I’m not infringing on any IP. You’re talking to people and they’re telling you, I have these two materials I want them to glue together, but I also want it to be sustainable because in space, if something burns, you are in a can, you can’t have air, it’s toxic, fire-blown, not happy. What’s nice about proteins is if you burn, it just smells like burnt hair, smells bad, but it’s not going to kill you. So there was a real use case for it. People tell us, okay, we want these material properties, for example being able to stick metal to plastic, curing in so many minutes and at that temperature. We have the databases that allow us to check what are similar use cases in nature. Turns out the mussel, the seafood, is really good at making a water resistant glue to glue on the rocks, and this is a protein glue. Basically what we would do is look at this protein sequence, identify the part that is responsible for gluing and adapt the rest of the protein to our particular use case — melting temperatures, curing time and so on. This adaptation is a small word for many, many methods that span mechanistic engineering to generative AI. That’s how we start from a functional element that we know is well-characterized and then push it around. The design space there is often gigantic. For NovaColl the design space was 10 to the 18, which is big. We do really smart sampling and screening selection with computers, down-select to a bunch of high-likelihood candidates, select the top n percent, turn that into a DNA sequence, order the DNA inside the lab and then pass it down to Lucile.
Lucile Bonnin: And then it goes to my lab, to the wet lab, which is a beautiful space. It’s incredible that you can screen candidates there because it’s already de-risking the work I’m doing together with my team. Once the DNA arrives, you clone it into an organism — we call them our micro factories — which is what generates the molecule of interest. We screen them according to different factors. You want to check if your organism is actually secreting what you think it’s secreting. That’s the first step. Then you want it to be very productive. It’s great if it’s producing the molecule that you want, but if it’s a very low productivity, you’re not going to be able to commercialize that. Then you want to check its material properties. Everything is de-risked by Pierre’s team first. But you still want to check that the safety and efficacy is actually aligned with your commercial usage. In terms of space glue, is it actually sticking under which shear rate, etc. We check everything in the lab in high throughputs. We can also talk about automation later because this is something we onboarded very early on, because even though we de-risked the use case by having very few candidates coming to the lab, it’s still a lot of candidates. You can’t run that many things manually. Right now we’re at a rate where we can screen more than a thousand candidates a week with our platform, which I think is pretty cool and outstanding.
Ross Katz: Yeah, that’s fantastic. So it sounds like the business development team, when they’re first talking to potential customers about potential applications, needs a menu of all of the different properties of proteins that they might select from, so that, Pierre, you can then take that menu and sub-select the candidates. Am I thinking about that right? It seems like a very complex ecosystem where the communication would be difficult.
Pierre Salvy: Communication is key, of course. The truth is you have information exchanges in both directions. Sometimes we have use cases that are just downright impossible, like people that wanted something that’s water resistant but also biodegradable in a salt bath, and you’re just like, well, if it’s going to rain on it, then the whole thing is going to be degrading. Sometimes you need to align this, because often the customer doesn’t really know what they want. They’re like, I want something sustainable that doesn’t know how it behaves.
Lucile Bonnin: On NovaColl there is a really cool example. We try to bridge the gap between sustainability and efficacy. For an anti-aging product like NovaColl, we bring anti-aging properties — no surprise. But you want to be as good as what is on the market. On top of that, you want to offer new things which were not there before. We have a peptide — we identified a specific peptide which is also antioxidant, which is the number one reason your skin is actually aging: oxidative damage. The molecule is doing both at the same time. It’s really interesting to see that on top of leveraging what exists already, you’re also bringing new things to the market through this new technology.
Ross Katz: It is very exciting. Starting from the beginning of developing the data systems you needed to explore the protein landscape and understand the different features of the proteins that were out there — so you could identify the candidates, sub-set them, leading up to handing it off to Lucile’s team to do the experimentation — Pierre, can you walk us through what the data journey was like? How did you build up those systems? What were the different components of the systems you were developing during that early phase?
Pierre Salvy: To be very honest, it wasn’t something that just built from the get-go and worked. It started as a bunch of scripts and Java implementations of some algorithms. There were maybe three different important milestones to this journey. One was: if you’re going to screen, you’re going to get data. If you get data, what can we infer from it? And how do you include this inference loop inside your design? One of the first things we did was actually simulate the whole lab in a computer to understand what the data input-outputs were, what the possible throughputs were with the money we had. From this, we knew: this is the data that we can use to improve our algorithms at a later stage. At the same time, we developed the protein design toolkit, which took the form of a protein programming language. The way people talk about proteins is very functional — this bit is responsible for that. But at that point there was no tool to do that well. People were literally copy-pasting protein sequences in Word files, and I heard this horror story of somebody getting their Word file auto-corrected for spelling, which changed the protein sequence without them knowing and it took them six months to debug why the thing wasn’t working. That was a real thing. So we thought, if you have a functional description of something in a database, then maybe we can just write a programming language. That was before the LLM boom. We wrote our protein programming language, which allowed us to actually design a lot of stuff in a very smart way. It’s a proper language that’s compiled and so on. This would be ingested by our AI to do diversification and optimization of the sequences. That’s how we designed NovaColl. Then the LLM boom happened and the way we looked at protein sequences data changed a bit. Now we have our own RAG for gathering information from experiments and informing the next design, the next experiment. We have diffusion models for proteins. Everything was added ad hoc to our little garden of models and today I think we have one of the nicest gardens.
Ross Katz: I want to unpack some of the components of what you said. When you say building a model that acts as your lab, is that the model that does the experimental screening for you to determine what the characteristics of a given protein sequence is? Or what does that mean to you?
Pierre Salvy: It’s closer to what people market as the lab digital twin now. Not necessarily simulating what’s happening in the experiments, but simulating the lab operations — what are the hubs of data, what are the failing points, what are the throughputs — so that you get an idea of whether you’re going to get a gigabyte of data per week, per day, per month. What type of data and which machines you need to buy and what are the marginal costs to increase your throughput. That was really basic. But if you don’t do that, then you might buy wrong machines and you lose six months and you can’t launch in two years.
Ross Katz: Right. So it’s a lab optimization algorithm to determine what are the different components of the lab that you need in order to execute end-to-end the entire manufacturing screening and manufacturing process that you need.
Lucile Bonnin: Exactly. When you work in a lab, you have a lot of machines which are not supposed to work together. You’re going to have a shaker, oven, a machine which is going to screen your proteins — they don’t communicate together. One of the foundational pieces of work was really understanding your lab end-to-end and putting all those pieces together to have the right database afterwards. That was a huge effort that we started with Pierre in the beginning — designing the lab right, buying the right machines. R&D is not always the most innovative place because you work with very expensive machines and sometimes some of them are not able to communicate with a standard computer. There were a lot of challenges that we faced in the beginning we were not even expecting, but now it’s all connected.
Pierre Salvy: You have to imagine the vendor calls with the business that sold us the robots. Lucile was asking all these super sharp questions and the guys were like, okay, I think I can deal with that. Then I was hanging in the back like, yeah but does it have an API? They’re like, what do you mean? And then I would start pitching them: I have an AI in the cloud, I want it to control the machine, is that possible? Literally some of them told us, why would you do this? And other people said, oh but you have an iPad and there’s a button on the iPad and you can tell your technical assistant via the AI how to turn the button? And I’m like, I don’t think you understand the vision, but that’s cool.
Ross Katz: Yeah, that’s very funny, and that’s been a running topic of conversation on this podcast — the way that these different automation tools don’t interface with each other, and the challenges of creating the data you need to understand what’s happening in your lab by integrating all of the different data sets together. It’s unnecessarily challenging. It’s difficult to do ML and AI if you don’t even have the basic data infrastructure and data integration in place to run the models on.
Pierre Salvy: Absolutely.
Ross Katz: Pierre, back to the protein programming language and the way that LLMs and diffusion models and the recent advances in ML and AI have changed things for you. What were your intuitions that led you to the approach of the protein programming language, and what are these new models doing for you that’s replacing the need for that?
Pierre Salvy: When I was working on the West Coast at Amyris with Total New Energies, we were making biofuel. They had this genotypes specification language that was really sick. They had a bunch of DNA banks and you could just type in ligation between the DNAs and build your own DNA construct, hit enter, it would compile into robotic instructions and build the DNA — basically DNA as a service, which is absolutely insane. We are not trying to repeat exactly the same thing but the use case was the same. You have sequences, you have functional pieces within the sequence, and there is a way to arrange it in a functional way that is not relying on fuzzy natural language but on computer instructions. So you would say, here I want a sticky bit, you go to the database, take one of the sticky bits, optimize for whatever I want, then add a linker, then here I want a blocky bit, and so on. Once we were able to write that down on a piece of paper, it made a lot of sense and we got the implementation there. It was super cool, but then LLM came in and it kind of overpowered it a little bit. Now we have tools that allow us to specify natural language and basically influence diffusion generation so that we’d have the protein that we want. It’s not a one-to-one replacement but it allows you to have more structure-based elements taken into account rather than sequence-based, which is important for some types of proteins. The answer is we use a bit of everything.
Ross Katz: Yes, it’s interesting. The protein programming language was a combinatorial approach where you have these different elements that have functions you’ve seen in the wild associated with these portions of DNA, and then you combine them to select a set of the most likely candidates — those sequences that then go out to be ordered and go through Lucile’s lab to determine whether they’re going to work. But these more powerful models, because they’ve been trained on such a large corpus of data, have a deeper understanding of the structure of the proteins and the functions of the proteins, and I’m assuming they can get you closer to the right protein faster because they have that understanding already. Am I thinking about that right or are there pieces I’m missing?
Pierre Salvy: It’s just yet another way to explore the forest of possibilities. We had sonar and now you have lidar on top. There are pros and cons for both. With the protein programming language, you get what you type, what you optimize for. At the end of the day it’s basically a combinatorial multi-objective optimization problem — which marketing doesn’t like me saying because it sounds too nerdy, but that’s what it is, and I think it’s exciting. It doesn’t hallucinate. A very deep network might hallucinate and you have very little control on that. So we use a bit of both.
Ross Katz: Yeah, that makes a lot of sense. As you’re moving toward maturity, what has the development of the data platform looked like? Has it been moving to the cloud, doing more parallelism, using Kubernetes? How has that journey looked for you?
Lucile Bonnin: We analyzed the R&D pipeline and where we can leverage data and identified three fields which were very, very useful. The first we’ve addressed deeply already, which is protein design. Then we looked into strain engineering — there are plenty of applications Pierre is working on to optimize how our micro-manufactories are working. Pierre wrote a PhD on that so he was pretty qualified for the job. Then we looked into manufacturing and scale-up. As a chemist, scaling up processes is where you’re going to meet the most problems. That’s where you put the most money and you really don’t want to miss on that, so strategically it made a lot of sense to look into scale-up. Really early on we integrated all the bioreactors we have in the lab with Pierre to make sure that all the data through the whole process is actually uploaded to the cloud and that we are able to correlate that with our partner who is producing for us at large scale. Now we have this correlation where a one liter in the lab is correlated to how exactly a one ton works at our partner, and we are able to tweak the process with that. We put a lot of money there, but you also have to be really fast there for the sake of speed and competitiveness, so that was a very cool angle to go for. Still continuing working on those three axes.
Pierre Salvy: To put it in the context of your question, Ross: from the beginning we were cloud native. I like to say digital native, but from the beginning there were no notebooks, no local files, everything went to the cloud from day zero. That’s what allowed us to move so fast, because first you store your data, then you try to organize it. It’s always cheaper than 10 years later trying to organize data that is in an unstructured format. Ask the BASF or Bayer of this world how much money they’re spending trying to retrofit an infrastructure on their existing data pipelines or notebooks. Terrible. From day zero we were cloud-native. Then what has changed is the technologies we’ve been using. Now that we’re doing more and more protein embeddings, we’re starting to use vector databases, benchmarking several solutions. We have a lot of structured storage. We’re starting to move toward more unstructured data storage, testing different methods and solutions there as well. The AI research assistant we built — the RAG — was a good way to better test several features.
Ross Katz: In the beginning you were designing the manufacturing process from scratch. Could you talk about how you set a vision for what the manufacturing process should look like? Any insights from your PhD are welcome as well.
Lucile Bonnin: Manufacturing. The same way we mapped the lab, we mapped our supply chain. You have all the R&D which is happening in the lab and you scale it up to a certain scale. We don’t have massive fermenters in the lab yet but we have a lot and it does the job. Then you look into partners — we produce with a different partner who covers those steps at large scale. You obviously don’t get from the lab to manufacturing at commercial scale in one day. You have a scale-up phase which is pretty impressive and also very data-heavy. The engineering team from Pierre helped us a lot to understand how to de-risk the tech transfer. When you run a precision fermentation you really have to understand what’s input, what’s output, what’s the risk, and how to use the data to make the right decisions during production. This was a collaboration between both teams. Then you ferment, you filter everything, you clarify everything, you stabilize it and make it commercializable. We actually have two partners in place for that, which is also very rich in data that we collect on our ELN to make sure that this is transferable as we scale.
Pierre Salvy: I want to point out two things that are really not obvious but really fantastic in what Lucile just said. One is, if you ferment 10 cubic meters, that’s 10,000 liters, that’s 10 tons of material — a gigantic amount of things to process. The second thing is, having correlation between the fermentation in a one liter tank, that’s the size of a water bottle, and the 10 tons tank is insane. That’s what scale up is — you go across four orders of magnitude and you get exactly the same cake you’re trying to bake. If you don’t have a background in bio-manufacturing, this is absolutely amazing because it takes years for other processes to get there.
Lucile Bonnin: That’s what I did during my PhD when I worked on the chemical platform. You realize that the physical laws in action in the big tank are actually completely different than in a small glass tank in the lab, and you have to transfer those laws from a mini scale to a very large one and try to understand how to de-risk that. Every time you work in a different setup you have to recalibrate your model. That’s what we did and that’s why it’s also important to have a very good partner in the first place and not to work with everyone, because you want this database to be consistent and qualitative. A scale up from lab to very high scale in six months, which brought us to commercial scale — two years ago I would have been laughing at that. But we made it and the team was fantastic, involving a lot of people: bioprocess, engineering, all the protein people.
Ross Katz: Lucile, can you give us some examples of the steps you took to de-risk moving from the lab to scaling up in manufacturing, and how you use the tools at your disposal to monitor and control that risk?
Lucile Bonnin: De-risking is probably the right keyword. There is a lot of that in what I’m doing. When you run R&D in biotech, a lot of experiments fail and you don’t want this to happen at large scale, so you’d rather de-risk properly. The first thing to do is play with parameters. When you look at precision fermentation you have a couple of feeds — you look at oxygen rate, aeration, mixing, etc. You stress test your system. If the feed of glycerol is not working on a specific day, what’s the lower boundary and what’s the higher boundary so that your cells stay happy and still produce something? There is all this stress testing happening, and throwing it on the cloud is also very useful to guide the manufacturer. Then you also have a stage called pilot — you don’t jump from a one liter to an actual 10 ton, but you go with hundred liters and maybe thousand liters if necessary and you test the robustness of your process, to make sure that the transferability is actually true but also that the process is robust, that you can repeat it twice in the same condition and nothing changes. That’s the de-risking strategy you put in place in the beginning.
Pierre Salvy: There was this use case where we had an anomaly feeding from the tanks and we put the data inside one of the models of the cell we have, and you could see that there was something in the metabolism that was not acting as we wanted, and we could act on it at the genetic level and improve that. A perfect proof of concept.
Ross Katz: Right, and that’s the feedback loop you create by having the data infrastructure in place — it takes the stress testing happening in the lab and brings the data up into the cloud where it can be quickly integrated into the models you have available to analyze what’s happening from a biological perspective. That’s really interesting. I’m imagining that the output of this exploration of the parameter space — the different variables in the environment or in the inputs that you expect to vary within your manufacturing process — is guidelines for your manufacturing partner: these are the controls we need in place for how the manufacturing process needs to go. Is that the way that it works? And do those parameters change over time, and how do you manage that process?
Lucile Bonnin: Parameters are going to change from an environmental reason — maybe this day is colder and you have to adapt the temperature differently — but you also have the human factor. When we are working in the lab we automate everything. I can write a script on my fermenters and I know it’s going to run the way I want because there is a closed-loop system. When you work at higher scale, not everybody is equipped with those technologies. There are people literally pressing buttons, making decisions — when they see these parameters reaching certain levels they need to press the button. But there is still this interpretation of: am I in the range, do I need to press the button? This guide is helping people make the right decisions and also de-risk it. You give very direct guidelines — feed glycerol at that time — but you know if there is 10 minutes difference, in the end your process is going to go well. It’s exactly a guidelines document, just very complex to read and not accessible to everyone.
Ross Katz: How do you approach quality control at that scale as you’re scaling up? When it’s in your lab and you have a liter it’s very easy to sample it and run your assays against it. But how does quality control work as you’re scaling up?
Lucile Bonnin: One thing to say is that Cambrium is very innovative in everything upstream — protein design and getting to a process which is very efficient very fast. When it comes to quality control there is not so much room to be innovative. You can be much more efficient than everyone else but there is no space for creation. You follow the guidelines, and the key for me is partnering. You just partner with the right people who did that for years, who know the industry, they know the assays and they can guide you there. We have a lot of support from consultants, really good partners. There is unfortunately less improvisation available, but you can be creative in driving that the proper way — you just have to surround yourself with the right people.
Pierre Salvy: Yeah here the marginal value of human expertise is so high you’d rather bet on the right people.
Ross Katz: That makes a lot of sense. Cambrium is a biomaterials company — you’re not just focused on NovaColl, you want NovaColl to be the first of many biomaterials you’re bringing to market. How do you think about the development of broader, more integrated software, data science, and manufacturing capabilities at Cambrium that allow for that expansion? How repeatable do you feel the process you’ve developed is for other biomaterials you might develop, versus additional work that needs to be done?
Pierre Salvy: I’ll give a two part answer — one Cambrium-specific and one about the general field. We built the platform so that for the next molecules, the price decreases because the platform already exists and is working. We’ve already seen results like that in the lab — increasing efficiencies in molecule classes that are totally different. That part of the bet is already looking good. There is always the question of can you guys do something else than cosmetics, which we need to answer. We have so much cool stuff we want to show. The product cycle has a given rate, so hopefully in the next months or years you’ll hear about some other cool stuff we are preparing. I’m super confident about it. We proved that it was possible with NovaColl and now we can bring it to industrial scale throughout the material world. For the market in itself, a lot of movement has been happening on the pharma and healthcare side of things. Bio-manufacturing is still the misunderstood child of life sciences in the industry, and there is an opportunity for anybody who is innovative, anybody who is an entrepreneur, to make things happen. There are things to get done. Bio-manufacturing is the solution to many problems of tomorrow. The AI boom that we are living right now is something we absolutely have to catch in order to make a difference, and I would really encourage anybody who’s super excited about this to dig in there because I think there are an infinite amount of possibilities.
Lucile Bonnin: There is definitely a lot of space to create really cool new bioproducts. Beside the pipeline that we are creating — which is super exciting and we can’t reveal yet — you can imagine protein-based everything which is vegan and can offer different properties. The space is massive. The data point that was amazing a year ago was when we tested the efficacy of our first product and it was mirroring what Pierre designed a couple of months before. If it worked on a substrate as complicated as skin for anti-aging, I’m really excited to see how this goes on different additional applications. Cool stuff cooking in the lab right now.
Pierre Salvy: I really want to make space glue, that’s my goal for the next 20 years, but I want to get this space glue out.
Ross Katz: It sounds like the way you structured the problem at the beginning, Pierre — the data platform you’ve described and the way you think through the design and scale up of proteins — is repeatable. What I’m also hearing is that the space for exploration of these different applications of protein design is wide open. There’s a lot of opportunity for innovation there, and you don’t seem particularly scared of competition because there are so many applications of what you’re doing. Am I hearing that correctly?
Pierre Salvy: Obviously investors don’t want to hear that. But one of my good friends said, it’s like the cake is so big you can’t eat it by yourself — bring friends along and build the bio-economy. I do want to echo this. There is a lot to be done and a lot to collaborate on. We’re friends with people who make molecular cheese, for instance. We’re not going to make cheese for them and they’re not going to make materials for us, so why not help each other? At the end of the day we’re both sequencing strains. There is an opportunity for exchanges, for building the community together. One thing I want to mention is the space is big but we have to be smart about how we work here. When we are trying to make biofuels, we’re trying to make commodities with deep tech, which is good when you have backing from oil and gas companies, but way harder when you’re a startup. You start with things that have high added value, specialty chemicals, and then you work your way down to commodities, because ultimately the way you make a material change in the world is by replacing plastic and concrete. But the price points and quantities for those are so high you can’t start there as a startup. You do specialty chemicals first and work your way down the value chain so that you can ultimately change the world. And for that we need many companies.
Ross Katz: Yeah that’s great advice for new startups.
Lucile Bonnin: We have a lot of really cool profiles in the company who are trained on both sides — as they come along they are trained on one side or the other. That’s really the base of creating the team which is going to give you really excellent results. My team and Pierre’s team are different but at the end of the day they work together, almost like a single team. That’s one very strong point. Also, when you welcome people in the lab it’s about onboarding them very fast on data and data capture, being very diligent about why you do that, why it’s important and where quality matters. That’s been one of my battles and it’s very close to my heart — I always talk about it because if you do R&D the old way you have no chance to be competitive in the future and there is no reason not to use those tools. I still don’t know how to code but I can sit beside Pierre and tell him look, I need that, do you have a solution for me, and then we build it together. That’s a very strong decision that a company can make in the beginning: embedding the data journey from very early on.
Ross Katz: It sounds like a commitment at the leadership level within Cambrium to collect that data and make sure it’s driving the decisions on a day-to-day basis is really important. Are there any challenges you’ve faced in trying to do that rapid onboarding in the lab or trying to make the teams collaborate more effectively?
Lucile Bonnin: Outside Cambrium there are always people who say they like to do it the old way, but after repeating and showing them that you’re actually very fast compared to the rest of the industry and competitive, usually they get convinced. In-house at Cambrium, people working in the lab have a lot of manual work. Everything takes a lot of time. It’s very complicated work. You have to be thinking but at the same time be very manual, and you add an additional constraint of them needing to enter some data in a certain system when they could just draft something on a whiteboard. Really talking to people — that the time they take for data capture is part of the job and part of what they need to deliver, and it’s going to save them a lot of time later — that’s the end point I always mention to motivate people to get behind it. At Cambrium the onboarding has always felt very smooth and natural, but it’s also something we screen for when hiring. Checking the enthusiasm about those tools is something we do.
Pierre Salvy: Once you have buy-in it snowballs really fast across the team because people see that you’re saving me so much time by doing this. It’s pretty obvious.
Ross Katz: As we draw to a close, do either of you have any resources that you would point data scientists and biologists toward if they’re thinking of exploring in silico design of proteins — anything out there that helps people get up the learning curve of this field?
Pierre Salvy: The first thing I want to say is it’s not because you never did biology that you shouldn’t try — it’s absolutely captivating and easy to understand at the high level. I would always recommend understanding AlphaFold, not necessarily reading the paper but understanding how it works, and understanding the Facebook ESM embedding model because it’s foundational for a lot of the stuff that’s been happening in the protein space in the last years. Look at diffusion models and understand the power with natural language. Proteins are letters, natural language is letters. Composing a protein is the same as composing a poem. If you just randomly choose your letters you’re going to get gibberish. But if you do it properly you’ll get Shakespeare, and that’s what this whole field is about. I think we’re really good at Shakespeare in proteins at Cambrium.
Ross Katz: Fantastic. This has been a really insightful conversation and I really appreciate both of your time. Pierre and Lucile, thanks so much for coming on and looking forward to connecting down the line.
Lucile Bonnin: Thanks Ross.
Pierre Salvy: Thanks Ross.
Jason: And that’s it for this episode of Data in Biotech. If you enjoyed the episode, please subscribe, rate, or leave a review in your podcast platform of choice. See you next time.





