Listen on
Overview
Transcription factors are critical regulators of disease, often linked to conditions like cancer. Yet, their complex, disordered structures have historically made them “undruggable” targets for small molecule therapies. This challenge means countless therapeutic avenues remain unexplored, limiting treatment options and competitive advantage for biotech firms.
In this episode, host Ross Katz speaks with William Fondrie, Head of Data Science and Engineering at Talus Bio, who details how his team confronts this problem head-on. Talus developed TFScan, a unique proteomics platform that monitors protein-DNA interactions in live cells. They pair this with a sophisticated, AI-driven recommender system to rapidly identify and prioritize potential small molecule compounds. Fondrie shares how this combination of novel data generation and machine learning accelerates drug discovery and provides a blueprint for leveraging data to solve persistent biological challenges. He also shares critical lessons from building scalable bioinformatics pipelines and repatriating outsourced data infrastructure.
This conversation offers vital insights for CEOs, VPs of Engineering, and Chief Data Officers grappling with data architecture, the build-versus-buy dilemma, and the strategic application of AI in data-intensive environments.
Key Takeaways
Previously ‘undruggable’ targets become accessible with targeted data platforms and machine learning.
Transcription factors, long considered too complex and disordered for traditional drug discovery, are now within reach. Talus Bio’s TFScan platform generates unique protein-DNA interaction data in live cells, enabling machine learning models to predict effective small molecules and accelerate drug development against these critical disease regulators. This demonstrates how novel data collection combined with AI can redefine what’s possible in a field.
Strategic data model design prevents costly rework and enables scalable operations.
Early on, Talus outsourced its metadata warehousing, leading to a “very rough” system with unreliable queries and poor data integrity (e.g., newline characters in unique keys). Bringing data modeling in-house, cleaning the data, and rebuilding the schema was painful but essential. This emphasizes that a well-conceived and well-managed internal data model is foundational for auditability, downstream analytics, and achieving operational scale.
Foundation models extend domain expertise without requiring massive internal datasets.
In data-limited biological domains, pre-trained foundation models like ESM2 for protein sequences allow companies to inject vast prior knowledge into their predictive systems. These models provide strong representations for biological entities, accelerating the development of recommender systems for drug discovery. This approach enables startups to build powerful AI capabilities without collecting petabytes of proprietary data from scratch.
Balancing build versus buy decisions impacts speed, cost, and long-term explainability.
While time is critical for startups, William Fondrie advocates for building core scientific components in-house, especially when open-source options exist and team expertise is high. This approach provides full explainability, easier troubleshooting, and avoids reliance on black-box proprietary tools for critical functions. However, for non-differentiating or commoditized tasks, off-the-shelf solutions can save valuable development time.
Related: CorrDyn helps companies in biotech and life sciences build reliable data foundations. We offer expertise in data engineering and AI strategy to accelerate drug discovery, ensuring data reliability even at massive scale.
Full Transcript
Jason: Welcome to Data in Biotech, a podcast from CorrDyn where we explore how companies leverage data to drive innovation in life sciences. Every two weeks, we sit down with an expert from the world of biotechnology to understand how they’re using data science to solve technical challenges, streamline operations, and further innovation in their business. Today we sit down with Will Fondrie, head of data science and engineering at Talus Bio, to explore how machine learning, mass spectrometry, and innovative computational models are transforming drug discovery. We discuss how Talus Bio is targeting transcription factors, once considered out of reach, with scalable, high-impact data science. If you’ve ever wondered why transcription factors are historically hard to drug, or how mass spectrometry offers high-throughput, unbiased views of protein-DNA interactions, this episode is for you. Here we go.
Ross Katz: Will Fondrie, welcome to the Data in Biotech podcast.
William Fondrie: Hi Ross, thank you for having me out. I appreciate it.
Ross Katz: Well, just to kick us off, can you give us a brief introduction to your background and what brought you here today?
William Fondrie: Yeah, I am the head of data science and engineering at Talus Bioscience, a Seattle-based startup where we’re developing drugs to target transcription factors in a variety of diseases. My background — to get into that role — I think is interesting. I got my undergraduate degree in chemistry at the University of North Carolina at Chapel Hill. There I was actually an undergrad researcher in a physical chemistry lab, which was very different than what I do now. But one of the things when I was looking at grad schools was looking into going into genomics and things like that, and my PI at the time was like, won’t you miss the big machines that we work with here when you do that? He ended up being right. I did my PhD in molecular medicine at the University of Maryland Baltimore, where I ended up in a proteomics lab where we did a lot of work with exosomes and characterizing biomarkers in exosomes. And then also a lot of other random proteomics projects and non-proteomics mass spectrometry projects as well. During my graduate school career, I started out in the wet lab doing a lot of wet lab work which I think has been beneficial, but it turns out I’m not very good at it. Most of my wet lab experiments failed in one way or another, but one of the things that all the labs I was a part of needed was somebody to do computational work and to develop the analysis strategies that we needed to make sense of all the data we were acquiring. That’s where I naturally fell in and I ended up loving that aspect of it. At that time I feel like machine learning was on its rise again and became something I got really interested in and started teaching myself on the side. When I was done with my PhD work, I ended up going to the University of Washington and working with Bill Noble in the Department of Genome Sciences there, really developing machine learning methods and AI methods for proteomics specifically. Both analyzing the raw data and then making sense of all the data that we get out of that. And actually that’s where I met the founders of Talus. Lindsay Pino was a graduate student in Bill Noble’s lab at the time. She’s one of the co-founders and then Alex Federation was the other co-founder. He was a postdoc in a lab just next door. I got to know them during my time there and when they were starting the company, they reached out to me and thought I’d be a good fit and I decided to join and haven’t looked back since. It’s been a lot of fun so far.
Ross Katz: That’s a wonderful backstory for you and also for Talus. Can you tell us a little bit about Talus, and what you’re working on and the problem that you’re trying to solve?
William Fondrie: Absolutely. At Talus, we’ve built a proteomics-based platform to let us develop drugs against transcription factors. We’re trying to develop small molecule drugs that target this class of proteins that binds DNA and causes a whole host of downstream effects on the cell. In disease, these are master regulators of all sorts of genes and processes. They tend to turn on processes that shouldn’t be turned on or turn off processes that shouldn’t be turned off. In cancer specifically, they tend to enhance cell growth and proliferation or turn that off in the normal cell. What we’re trying to do is develop these small molecule inhibitors for transcription factors. They’ve been known to be good drug targets for a long time, but they’re traditionally very difficult to develop drugs against, for a wide variety of reasons. One of them is that they tend to have these highly disordered structures that are not captured well. They’re hard to crystallize, except for their DNA binding domains, which tend to be not great places to develop drugs against because of their high homology across proteins. So you end up hitting a lot of targets when you try and target a DNA binding domain. Traditional structure-based drug discovery approaches have not really been effective for them. They’re also really hard to measure in a test tube in isolation. They’re difficult to purify, and then once you’ve purified them, they really need all the components of the nucleus, all the DNA and their cofactors, to be active in their normal state. These assays that take place in a test tube with a purified protein, while they can be effective, have traditionally not been very effective for transcription factors for that reason. Our platform leverages the strengths of proteomics to do this in a high-throughput way where we perturb a live cell with a drug or a small molecule compound. Then we go in and we isolate the nucleus and measure what proteins are bound to chromatin and how much of each of those proteins are bound to chromatin at any given time. We get a snapshot of that cellular state after it’s been perturbed. What that lets us do is infer — when we see certain proteins falling off of chromatin, we can infer the target of the drug or the small molecule that we’ve treated with, and work to develop those to be more selective and more potent for the targets we care about in disease.
Ross Katz: That’s really interesting. Can you dive into why mass spectrometry as an assay is the best way of taking that snapshot of the cellular state when it’s been perturbed, and understanding how these transcription factors are being upregulated or downregulated?
William Fondrie: One of the things I love about mass spectrometry is that we get a fairly unbiased view of the system that we’re trying to measure. In our case, we get a high-throughput, unbiased view of what proteins are bound to chromatin at any time. That means we can measure the abundance of thousands of proteins that are either directly bound to DNA or bound through a cofactor to DNA in a single assay. We’re very high-throughput on the range of targets we can measure and then medium-throughput on the compound perturbation side. What that lets us do is instead of just taking a single target and throwing a bunch of compounds against that single target and seeing what works there — and then having to consider selectivity down the road — we get to measure both the on-target specificity of a compound and any of the off-target effects that we can measure in the nucleus simultaneously. Mass spectrometry is the technology we’ve chosen to let us do this. There are a lot of exciting new proteomics technologies that are coming online, but they’re just really not to the depth and reliability that we need for these particular assays. The other strength that mass spectrometry also gives us is that when we’re looking at small molecules that have a covalent mechanism of action — meaning when they react, they physically react with a protein and will stick to them — one of the things we can do with the mass spectrometer that is very difficult or impossible with other technologies is we can actually detect where that small molecule is bound in some cases. We can both get confirmation that we’re engaging the target of interest, we can measure the functional effect of that binding event on the protein’s binding to DNA, and we can measure thousands of proteins at the same time. Together those make it a really powerful technology for characterizing what these small molecules are doing to proteins binding DNA.
Ross Katz: Thank you. I think that serves as a really great platform and background for talking about the computational approaches you’re using at Talus. One of the things I’ve heard you talk about previously was this recommender system approach to drug discovery. Can you explain what it means to have a recommender system and how that relates to the data that you’re collecting and the feedback loop you’re creating between the wet lab and your computational work?
William Fondrie: Yeah, absolutely. The space of pharmacologically active small molecules is just absolutely massive — so massive that there’s no way we could possibly measure every single molecule we would want to measure, or find that needle in the haystack just randomly searching that space. What we really need, especially with a technology that’s medium-throughput on the compound side, is some system to prioritize what compounds we should measure first — which ones are going to be most likely or most informative for us to measure to hit a particular target. The way we’ve decided to tackle that problem is: first, we’ve developed our own chemical space that we want to search within. Within that, we’ve developed our models to recommend which compound is most likely to perturb a transcription factor. This goes back to the Netflix challenge — I forget what year that happened in, but the idea was that Netflix has this huge problem where, given a list of users and a list of movies or TV shows, they need to figure out and predict how well a user will like a particular TV show or movie. You can think about this as a giant matrix where one axis has the users, one axis has the movies, and the elements of that matrix are some rating of how well a user will like it. The way folks have solved the Netflix challenge were a variety of these different matrix factorization methods to learn something about what represents a user and learn something about what represents a movie and then combine those in different ways to get that predicted rating for each user-movie pair. Netflix, when you go to browse movies, can sort that list by your personalized rankings and give you the things that you’ll binge watch for days. The way we’re thinking about this is pretty similar. We have a slightly different data structure where we have cell types as one axis of this data matrix, or tensor. We have the protein targets that we measure and potentially target as another axis. And then we have the small molecule compounds as another axis. This looks like a giant cube if you were to visualize it, and each element in that cube would be a measurement of whether that compound affects the DNA binding activity of that protein in a particular cell type. Our goal is to build models that can predict the elements of that tensor, help us rank those for particular targets so that we’re measuring the compounds that are most likely to work first. The way we’ve done that is what’s called a deep tensor factorization. One of the things that’s been very exciting over the last few years has been the development of these foundation models for representing biological sequences in a lot of ways. One of the ones that we use for protein sequences is the evolutionary scale model. We’re currently using ESM-2 to represent these protein targets. What that does is give us a starting point for what it means to be a given protein.
Ross Katz: So you’ve got the protein types that you’re going to target, and you’ve got the small molecule compounds that you’re targeting them with, and then you’ve got the different cell types in which the small molecule and protein interactions are going to occur. Can you help me understand what role the cell types are playing in your exploration of what the drug is that you’re trying to discover?
William Fondrie: When we’re thinking about cell types, one of the things we’re always thinking about is what diseases is a particular transcription factor relevant in? What indications is it going to be active in, and is it going to be effective to develop a small molecule against? In different cellular contexts, transcription factors will actually complex with different cofactors and bind different regions of DNA. We want to make sure that when we’re developing an inhibitor for a particular target, it’s inhibiting that target in the correct context and not in a context that would cause wider ranging effects than curing the disease that we’re targeting. That’s really what the cell type axis represents. We’re trying to make sure that as we’re building out programs and taking on new targets of interest, we’re measuring cell types and onboarding cell types that are representative of that underlying disease biology, to maximize our success in the downstream drug development activities as well.
Ross Katz: Yeah, interesting. So if you’re trying to inhibit a transcription factor that has implications for a disease of the lung, you want to make sure it’s not having effects on that transcription factor in other organs of the body, like the liver. Am I thinking about that right, or are these cell types different in some other way?
William Fondrie: No, exactly. That’s one aspect of it. The other aspect is that a lot of these cell types have specific subsets of transcription factors that they actually express. Especially when you start looking at rare diseases — for example, our lead program is in a disease called Chordoma and there’s a transcription factor, the Brachyury protein, which is the TBXT gene. That one is often turned on, and it’s something that is normally only expressed during development. The opportunity to measure that protein in an environment that matters means that we have to go to Chordoma cell lines — or there are some other cancer types that have that expressed as well. We really have to make sure that we’re targeting that specific cellular context to make sure that we’re measuring the things that we care about. Does that make sense?
Ross Katz: That makes sense. The example you gave from Netflix was users and movies, right? In order to do a really good recommendation, you need to understand a lot about the users and a lot about the movies. In your case, you need to understand a lot about the small molecule you’re using to do the inhibition or the stabilization, and then you also need to understand a lot about the transcription factor you’re trying to bind to. Am I thinking about that right? How does ESM-2 give you the information you need to make the recommendations coming out of the system as good as they can possibly be?
William Fondrie: Yeah, you hit the nail on the head there. The naive approach we could take to build this recommender system would be to try and learn representations for drugs, proteins, and cell types all de novo from the data that we’ve collected. The biggest downside to that is that, as in most biological domains, we are data limited ultimately. We can’t train massive models from solely our data. These foundation models like ESM-2 give us the opportunity to inject all of that prior knowledge without having to learn it directly from our data or to bring in those sources ourselves. What that looks like for us is we embed all of our protein target sequences into their ESM-2 representations. Those embeddings are what we’re using to represent each protein. We’re using Mol2vec for our compounds. Those give us a good starting point for what it means to be a particular molecule. Given those two starting points, we can then feed both of these through a neural network, learn about cell types de novo, and combine all of those three to get our prediction to fill in that element in that cube tensor that I described.
Ross Katz: And the element that’s being filled in is some measure of transcription factor binding — what’s the value that goes in that little box?
William Fondrie: The readout of our platform, which we call TF-Scan, is an abundance of protein on chromatin. We transform that into essentially a percent inhibition, or percent gain on chromatin if a protein is stabilized. That transformation lets us look at the relative change of a compound to whatever basal level of binding a transcription factor has in a cell type. That’s what we’re predicting. That’s what we’re imputing into that matrix. We can then rank compounds for a protein target by that prediction and look for the compounds that would perturb a given target of interest the most, and measure those first.
Ross Katz: So you’ve been given this list of compounds that are likely to perturb a transcription factor of interest the most. What do you do other than running another experiment — how do you use that information to understand what’s happening in the biological system more effectively?
William Fondrie: One of the things we think about is, when we start to get into the chemistry of these compounds, we look for similarities in the structures of those compounds. We want to look for patterns that emerge there. You can do some fairly basic clustering to get clusters of compounds that come out of these predictions that are predicted to be highly effective against the target. But when we’re actually trying to narrow down, say, this list of a thousand compounds to a smaller subset that we might want to test quickly, one of the things we try to do is make sure we’re selecting for some structural diversity in those predictions. We get our list of a thousand and then we use a submodular optimization process to select the best representative subset of these compounds to actually physically test in the lab. The idea is that hopefully by testing around that space, probing around that space, we’re getting much more information than if we took the top 10 and they were all very similar to one another in our chemical space.
Ross Katz: Selecting a representative subset of the structural diversity of the compounds that are predicted to be really good is your way of maximizing information gain from the new experiments that you’re running. That’s creating the best balance you can between explore and exploit.
William Fondrie: Exactly. One of the things I just want to mention is that the recommender systems we’re learning are not static — we are constantly integrating data that we’ve collected based on previous predictions to build better models in an active learning loop. That’s the real key to making progress on particular targets: making sure that you’re reintegrating data that you’re collecting to maximize what the model can learn.
Ross Katz: You’re integrating with more than one model and bringing in data from all of these different sources to add information to the system that can be used to make better recommendations. Can you give us an orientation to what your data stack looks like in order to manage all of those moving pieces?
William Fondrie: When we think about our data stack, starting at the cloud foundations, we’re built on top of AWS — pretty much everything we do is processed and built in AWS. Our raw data we automatically upload from our instruments, it’s dumped on S3 in an organized data lake. We then have to run our bioinformatics pipelines to even begin working with the mass spec data that we’ve collected. Our TF-Scan experiments come out as these raw mass spec files that you then have to interpret to get the protein quantities that are useful for biology. For those bioinformatic pipelines, we’ve traditionally leaned on Nextflow a lot to build those — we’ve been using Nextflow orchestrating AWS Batch to make that happen pretty effectively. We contribute to one of the NF-core pipelines. NF-core is one of the great parts of Nextflow, where there’s a lot of community support around these open source pipelines for a variety of different bioinformatics tasks. One of those is for proteomics — there are multiple for proteomics — and we contribute to one of those to make it scale better and work for our use case. We run our data through that as part of our bioinformatics processing. That’s the vast majority of the compute we actually have to do. We’ve also recently been using Redun from Insitro, which is their open source orchestration platform. It’s been a lot of fun to use that one and to be back to orchestrating using Python instead of Groovy for the programming language. That’s what we use to orchestrate between our bioinformatics pipelines and then going into model training and updating, which is what we need to do on a regular basis. That’s what connects all those pieces together and is the plumbing that makes it all happen smoothly.
Ross Katz: You’ve got the data that you’re generating from the mass spec machines, and you also mentioned bringing in external data sources. Can you just orient us toward what are some of the external data sources that you’re bringing in and how you’re using them in the context of that stack?
William Fondrie: We’re not actually bringing in that many external data sources, to be honest. Most of the external data really gets injected through the model embeddings that we’re using. We are using protein sequences from UniProt for our bioinformatics workflows and to know which proteins we need to embed using ESM-2. But the vast majority of the external data that we’re bringing in for these models comes from the injection from these foundation models. One of the things that’s interesting about that is managing the feature store — making sure that as you’re updating models and using different models, there are all kinds of foundation models you can pull from, and making sure you have systems in place to know which ones are A, the most effective and should be used in the production models, and B, making sure you’re consistent about how those are updated and maintained.
Ross Katz: Right. That’s also a challenge with experimenting with different foundation models and bringing different embeddings into your ecosystem — knowing which ones are the production-ready embeddings, which ones are the ones that you’re experimenting with, which are the old versions that you might need to roll back to.
William Fondrie: I feel like our young team has really just started to dive into all the fun and all the headaches of ML Ops and engineering there.
Ross Katz: For sure. Relatedly, what lessons have you learned about designing scalable bioinformatics and drug discovery pipelines based on what you’ve been able to build?
William Fondrie: One of the things that is perhaps obvious to folks who have been doing it for a long time, but was not necessarily obvious to me getting into it, was that things unexpectedly break at scale all the time. Even on AWS, if you’re running tens of thousands of jobs at the same time, you’ll have an EBS volume fail, and you have to have systems to account for that. We’ve gone through the painful iterations of going from the scale of analyzing hundreds of TF-Scan experiments to the scale we’re at now of analyzing tens of thousands of TF-Scan experiments simultaneously. Being able to handle that data and trade off the practical aspects of how those pipelines should be run with the theoretical best practices is tricky. The other thing has been the relationship of building software versus buying pre-made solutions. I tend to lean heavily toward the build side of things, but one of my own flaws has been to realize that and remember that time at a startup is your absolutely most valuable resource. When you can pick something off the shelf, it’s often worth doing as opposed to re-implementing something yourself.
Ross Katz: I want to ask you a couple follow-ups on that. Operating at the scale of tens of thousands of TF-Scan experiments simultaneously is the scale where things like hardware just start to break, so you need reliability and recoverability. Are those some of the contributions you’re making back to NF-core to make it handle those kinds of things, or where in your pipeline do you need to build in that reliability?
William Fondrie: Part of our contributions to NF-core were to help it scale to the thousands of mass spec runs — we’re specifically contributing to the QuantMS pipeline, in case anybody is interested in that. We’ve used that up to the thousands-of-runs threshold. But when we really hit, I don’t know, 7,000 or so TF-Scan experiments, that’s when that started to break down. There’s one step of the traditional proteomics pipeline — this might be going a little into the weeds — where essentially all the data has to be aggregated onto a single node, which becomes impractical eventually. That’s where things really start breaking. We designed our own internal pipelines to avoid that step using some heuristic approximations of what’s theoretically the right thing to do, and I think that has helped us to really break through that barrier and allow us to scale further. But a lot of the optimizations we made to make that happen are really directly tied to the infrastructure we’ve deployed. It’s one of those things that could be open source, but I don’t think it would be useful for it to be open source.
Ross Katz: Unless people are replicating exactly the workflow that you’re undergoing. That makes perfect sense. And I know that you historically had outsourced some of your data modeling and warehousing, but later brought it back in house. Maybe that experience is part of why you have this bias toward building. Can you give us an understanding of what that experience was like and what it’s taught you about the build versus buy tradeoff, even in the context of a startup where time is so critical?
William Fondrie: Just to introduce everyone to our metadata storage strategy: we wanted to track all of the metadata associated with all of these TF-Scan experiments. Given a TF-Scan experiment, we wanted to be able to know what batch of compound was used to treat those cells, what cells they were, where they arrived from — all of the provenance information for that. Because it’s important for auditability, and important for downstream tasks in modeling to be able to know exactly what an experiment was. Early on, we had built or theorized a data model for what this could look like in a relational database, and worked to try and get it implemented through some third-party vendors. When we thought we had it implemented, it ended up being very rough. We would try and query the data in the expected patterns and then we would either find that the API from the vendor wouldn’t exist, or we would need to have some pre-knowledge of UUIDs to specifically retrieve the things we were looking for — when that was the entire point of the query was not to use those. We made the tough decision to bring that data modeling and warehousing back in house, where we now have a relational database with all of that metadata stored there and then linked out to where all the raw data and experimental results live. That transition was very painful, in part because there were other data modeling decisions that had been outsourced in that process where there were just nightmares lurking in the database — like unique key constraints were enforced, but some of the unique keys were the same key with a newline character. Things that should never exist in your database to begin with. We went through the painful process of having to extract all that data, clean it, and then rebuild our schemas on top of that to actually manage it ourselves. I think the biggest lesson I learned from that is: you want to think about your data model early on and know exactly what you want to capture in it, what purpose you want it to serve. If you start too broad, you try and capture too many things and add a lot of complexity. If you start too narrow, it’s very painful to go back in time and add things on. Really think about what you want your data model to do and specify that correctly. And then be involved in that data modeling process and verify at each step of the way — if you’re using an external vendor — that it is actually functioning as you expect it to, rather than going through this painful transition and finding out in the end that things are just not working reliably. That’s our story of how we had to undergo this painful process of repatriating that data into our systems. Now it enables us to do things at scale that would have previously been impossible. All of our internal systems now reference it and can use that data in a variety of ways.
Ross Katz: That highlights the extent to which external vendors will not always have the same quality standards that you have for the way your data is architected, or the same sense of what the long-term strategic implications are of the decisions that you’re making. Even if you’re outsourcing the writing of the code, the rigor with which a process needs to be managed and the attention paid to it is critical to making sure that it establishes a solid foundation for you.
William Fondrie: Exactly. And just to be clear, I’m not saying people should never work with an external vendor for their data modeling. You absolutely should — just be involved in the process and be more involved than I was, and don’t make the mistakes I made.
Ross Katz: That makes perfect sense. I know that you’re passionate about open source scientific software. Can you help everyone understand from your perspective why open source software is so important for progress in science?
William Fondrie: For science in particular, open source is really the only way to make progress as a whole. When we look at science in general, when we make great claims or great discoveries, we have to present the evidence to back that up and present how something could be. When we think about a wet lab experiment, you have to have the evidence to back up the claims that we make about that experiment. I think the same thing is true for the computational tools and resources that we use. In modern times, a lot of the really big discoveries — protein folding and other types of challenges that AI has really come in to solve — are really cool scientific advances, but if we don’t have the code, which is the proof that it is doing what it says it’s doing and describes exactly how it’s done, then we’re really just advertising about a cool new technology at a company or a research group. Open source is what allows folks to both integrate that new advance into their own workflow, test it and verify that it actually does what it claims to do, and make sure the claims are verifiable. More importantly, it allows people to build off of it, create new things, make new discoveries, and expand science faster. Those are the big reasons why open source is so critical for science in general. There’s also the reproducibility side: without the code used to analyze data or for a model, it’s impossible to reproduce the results that are presented in a paper, and we already live in a time where it’s very difficult to reproduce experiments anyway. Not having the code there just makes it orders of magnitude more difficult.
Ross Katz: Right. If you’re able to speak to the exciting results that you’re producing but nobody else can reproduce those results, it’s an advertisement for your organization but not necessarily a validation of the capabilities that you’ve claimed to create.
William Fondrie: Exactly. That’s not to say there’s no place for closed source software — I absolutely think there are places for closed source software. I just think that if you’re conflating the closed source software with the scientific advance, then you have a huge conflict of interest there that needs to be resolved. Those critical pieces of open source scientific software can be the foundation of great companies that build on top of it and offer things outside of that. But the core science aspect of it I think is what really needs- like you really need to understand what’s going on there to contribute to the scientific body of knowledge.
Ross Katz: That makes a lot of sense. For people who are considering utilizing proprietary closed source tools for their biotech or bioinformatics workflows versus the array of open source tools that are out there, how do you think about the risks of adopting one or the other?
William Fondrie: The tradeoffs here in part depend on the expertise of your team. If you’re depending on closed source proprietary tools to perform some very core function of your company, that is a risk you’re assuming because you’re assuming that those tools are doing exactly what you expect them to do without actually knowing how they do it. You lose the explainability of what those tools are doing. That being said, it also means you can push some of that risk onto the company that’s licensing you the tool, depending on the licensing terms you have with them. In regulatory restricted labs, having that vendor responsibility makes a lot of sense. On the research and development side though, the reliability of those tools is not covered under those licenses in most use cases. You really are assuming the risk for the data that comes out of those tools and the availability of that tool into the future. If you have a lot of expertise on your team in that particular technology though, the open source side of things when it’s available is really a good option. If you have the expertise to implement those tools and to dive into them and understand what’s going on, that adds a lot of value to the data that you’re getting from the tools and a lot of explainability when things go wrong — and it can help you with that troubleshooting and understanding all the edge cases in your data that you need to model.
Ross Katz: That makes a lot of sense, and with adopting open source tools you have the ability to accelerate the work that you’re doing internally, which when you’re a startup and resource constrained is a good place to be — especially when you don’t have to pay for the open source alternative to the proprietary thing. Whereas when you’re later on and you want someone else bearing the liability and cost is less of a concern, it makes sense to adopt the proprietary tool at that point. As we head toward the end of our conversation, where do you see your work in computational biology, or the computational biology field more broadly, going over the next three to five years? And how do you see the work that Talus Bio is doing going to shape that future?
William Fondrie: There are so many exciting things going on in computational biology, and machine learning and AI in biology in particular are just really exploding with popularity. One of the things that has traditionally been a fundamental limit of what we can do with computation in biology has been the size of the data that we’ve been able to collect. There are a lot of new technologies that are making that easier and faster and cheaper. As more of those come online and we have these larger data collection efforts, we’re going to see more and more foundation models and large models specifically tailored for biology come out and be very useful in a lot of domains. I’m really excited for the synthetic biology angles that a lot of this unlocks. I think we’re going to see the protein design fields really explode as well, if they’re not already. When I think about Talus’s role in this, our platform is unique in the type of data that we collect and complementary to a lot of other approaches. It’s really valuable for us as an integration into that, but also when we have models that essentially speak the language of transcription factors, that can be very useful for modeling a cell and understanding when you perturb a particular transcription factor what the cascading effects are for the cell as a whole, instead of just the little bit of the cell that we’re looking at in isolation currently. There’s a lot of cool modeling that will be happening at those secondary, tertiary, and farther-down effects of what we can observe and model.
Ross Katz: It occurs to me as you’re talking that Talus, with its narrow focus on transcription factors and your vision for understanding the broader biology — there are platform companies being developed with all of these different angles on the cell and an opportunity for that knowledge to combine itself into a picture that creates a much more nuanced and complex view of the biological phenomena than would have been possible in the absence of any one of these angles. Am I thinking about that right?
William Fondrie: Exactly. There are all these very cool technologies that are being developed and have been deployed in various ways that, combined, could really unlock a lot of interesting biology and push forward human health in a really good way. I’m really excited for the future and what’s going to happen over the next few years.
Ross Katz: Will, it’s been a pleasure to have you on. For people who are interested in learning more about Talus Bio, where should they start?
William Fondrie: We have our website, talus.bio, and the company is pretty active on LinkedIn and we have a BlueSky account as well. You can find me on LinkedIn and BlueSky as well if you want to connect.
Ross Katz: Will, thanks so much for joining — I really appreciate everything you’ve shared today and look forward to connecting down the line.
William Fondrie: Yeah, thank you Ross. I appreciate it.
Jason: And that’s it for this episode of Data in Biotech. If you enjoyed the episode, please subscribe, rate or leave a review in your podcast platform of choice. See you next time.






