Listen on
Overview
Host Ross Katz speaks with Ming “Tommy” Tang, Director of Computational Biology at Immunitas Therapeutics. Small biotech companies face an inherent tension: how to build reliable, scalable data infrastructure while simultaneously hitting aggressive, time-sensitive deadlines in a survival-mode environment. Tommy provides a pragmatic perspective on this challenge. He emphasizes prioritizing rapid, “good enough” solutions that deliver immediate value, iterating towards perfection rather than risking missed milestones with an upfront, resource-intensive approach.
Tang’s journey from wet lab molecular biologist to leading computational biology teams at institutions like MD Anderson, Harvard, and Dana-Farber, and now spearheading Immunitas’s computational capabilities, offers a unique, dual-perspective on data science in biotech. His insights are crucial for data leaders and executives aiming to balance speed, scientific rigor, and strategic data investment for competitive advantage in drug discovery.
This conversation explores how Immunitas established its computational function, the critical role of data quality and harmonization (especially with public datasets), the pragmatic application of machine learning versus simpler statistics, cloud data strategy, and the necessity of deep, mutual learning between computational and wet lab scientists to ask the right questions and drive clinical impact.
Key Takeaways
Prioritize ‘Good Enough’ Data Infrastructure to Hit Biotech Deadlines
Small biotech companies operate in a fast-paced environment where missed deadlines carry significant risk. Instead of striving for perfect data infrastructure from the outset, a “good enough” approach that delivers immediate, functional solutions allows teams to move quickly. Gradual enhancement of reproducibility and making processes more efficient can follow, ensuring critical milestones are met without over-investing in upfront perfection.
Data Quality and Specificity Outweigh Sheer Volume for Research Impact
The true challenge in immuno-oncology isn’t simply generating more data, but acquiring and curating high-quality, relevant datasets. Asking the right biological questions enables the design of experiments that yield specific, enriched data, even if in smaller volumes. This focused data, often generated through close collaboration with wet lab teams, drives more meaningful insights than applying complex models to large, unrefined public datasets.
Deep Collaboration is Essential for Effective Computational Biology
Successful computational biology in drug discovery demands continuous, mutual learning between computational and wet lab scientists. Computational teams should provide formative input on experimental design and tool selection, while biologists offer crucial context to interpret findings, distinguish biological signals from technical artifacts, and refine research questions. This interdisciplinary feedback loop accelerates discovery.
Metadata Management is the Primary Hurdle for Machine Learning in Biotech
While advanced machine learning models receive significant attention, the most substantial practical challenge in applying them to genomic data, particularly public datasets, is messy and unharmonized metadata. Cleaning, standardizing, and tracking metadata is a labor-intensive but critical prerequisite for any effective data analysis or machine learning application. Adequate metadata infrastructure scales beyond simple spreadsheets as data volumes grow.
Related: CorrDyn helps biotech companies establish solid data engineering foundations and improve data quality. For a deeper dive into industry-specific challenges, explore our insights on biotech and life sciences data. We also assist with practical AI strategy that delivers measurable results.
Full Transcript
Jason: Welcome to Data in Biotech, a podcast from CorrDyn where we explore how companies leverage data to drive innovation in life sciences. Every two weeks, we sit down with an expert from the world of biotechnology to understand how they’re using data science to solve technical challenges, streamline operations, and further innovation in their business. This week, we’re excited to be joined by Tommy Tang, director of computational biology at Immunitas Therapeutics, a single-cell genomics-based therapeutics company focused on immunology. During this interview, our host, Ross Katz, speaks with Tommy on the challenges and importance of data analysis in the field of computational biology, how to select the appropriate machine learning models to solve specific problems, what it takes to build an internal culture that encourages shared learning, and why he is motivated to build a community around computational biology. Here we go.
Ross Katz: Tommy Tang, welcome to the Data in Biotech podcast.
Ming (Tommy) Tang: Thanks for having me.
Ross Katz: Just to get us started, in a minute or two, would you mind telling me about your career to date?
Ming (Tommy) Tang: Of course. I came to the states in 2008 to pursue my PhD at University of Florida, majoring in genetics and genomics. But I was trained in the wet molecular biology lab, cancer biology. I did a lot of experiments pipetting, cell culture, all this stuff. Four years into my PhD, I was forced to learn computational biology by myself because my advisor at that time asked me to analyze this publicly available ChIP-seq data and I couldn’t open it with Excel. It crashed. I realized that I had to actually learn some computational skills, how to use Unix, Python, and R, by watching online classes, reading books, Google a lot, and it’s a lot of work. Two years later, after my PhD, I joined MD Anderson Cancer Center to pursue a full computational biology postdoc. At that time, my advisor was Dr. Roel Verhaak and he was leading the TCGA, the Cancer Genome Atlas glioblastoma project. There, I had the opportunity to analyze real large-scale TB-sized genomic data sets and I gained skills to analyze different kinds of datasets, like whole-exome sequencing, whole-genome sequencing, bulk RNA sequencing, bisulfite sequencing, ChIP-seq, ATAC-seq, all sorts of different sequencing data. Later I moved to Harvard as a senior bioinformatician. There I started to work on single-cell RNA sequencing and single-cell ATAC sequencing. Two years later, I joined Dana-Farber Cancer Institute as a senior scientist first, then a lead scientist to lead a bioinformatics team of seven people to lead this big NIH-funded consortium project called CIDC, Cancer Immunologic Data Commons. There, four cancer centers in the states, MD Anderson Cancer Center, Dana-Farber, Mount Sinai, and Stanford, they carry out immunotherapy clinical trials in partnership with big pharma companies, and they profile those patients using NGS, and our team responsibility was to process all those NGS data securely on the cloud and then deliver them back to the clinical trial for downstream collaborative analysis. We are also responsible for some of the cross-trial data analysis. Stayed there for almost two years and I joined my current company, Immunitas Therapeutics, as the director of computational biology to establish the computational capability for the company. Immunitas is a single-cell genomics company using single-cell RNA sequencing and single-cell TCR sequencing to find new targets for immuno-oncology. Our leading program IMT-161 is in clinical trial phase one with several patients dosed already, and it looks like it’s still safe. It’s exciting and fingers crossed that it does well for phase two and phase three and eventually can benefit more patients.
Ross Katz: It’s quite a journey across the Pacific Ocean and all over the US to various academic institutions and now into the private sector with a therapeutic solution in phase one trial. Why did you enter this field in the first place? And why have you continued on this journey in computational biology?
Ming (Tommy) Tang: First of all, for me, there are two reasons. To learn computational biology, usually there are two reasons for any person to do something. First is stay away from the pain. Second is you really like it, that’s your passion. Initially that was just the pain because I had to beg other people to do analysis for me and definitely your project is not your priority. I said to myself, I had enough. Let me start to learn it myself. Initially, that’s just the motivation. But later after more experience with computational biology, I really like it because I like the Unix commands. Initially could be intimidating but after I mastered how to use Unix commands, I just couldn’t go back to use the mouse and drag and click. Of course for my current role in the biotech company, what we are doing now can really benefit patients and make a clinical impact. That’s something I’m always passionate about. If we can help others save the patient, save lives, that means a lot to me.
Ross Katz: What impact has computational biology had so far on immuno oncology as a field? And what kind of work are you doing at Immunitas Therapeutics to advance understanding of immune responses in cancer or the development of therapeutic treatments to treat cancer?
Ming (Tommy) Tang: That’s a great question. With this sequencing technology advancing so fast, the cost just dropping so dramatically. But then the data is exploding. I think last week I was in this conference, the festival of genomics and biodata, and the sequencing cost you see it drop dramatically along the years. But how to efficiently use those data, analyze data, and derive biological insights, that’s still a very big challenge. That’s where computational biologists come to play a role to crack the code from those data. For us, immuno oncology, for any field, data is still a major block for us because when you talk about data, you talk about the quality of the data or the quantity of the data. Because I did some wet lab stuff, I know generating high quality data actually is not easy. I really appreciate if the biology team for our company hand us really high quality data. I really appreciate that because it’s a lot of work. Quality of the data and the quantity of the data, that means whether you have high throughput experiment assay in your company to generate those high throughput data. Of course for us, we always use those publicly available data, but one of the major problems of those publicly available data is they’re messy. Inherently that’s just messy and you need to spend a lot of effort to clean those metadata and put them together, harmonize them together in order for you to analyze or even just run any machine learning models on them.
Ross Katz: So when you’re coming into Immunitas and building this function of computational biology, where do you start? How do you grow that capability inside of a company like that?
Ming (Tommy) Tang: That’s a great question. From my perspective, it really depends on the company that you are in. The big pharma companies are very different than the small biotech companies. As a small biotech company, you are always in this surviving mode, there are always deadlines that you have to hit. There should be a balance between perfect versus good enough. For us, when I joined, we had one computational biologist and he established the computational platform using Terra, which was developed by Broad, and we also use native Google Cloud GCP to do those large scale single cell data processing. I would say quick and dirty initially, as long as it works, that’s fine. Then gradually you can enhance different properties, make the process more streamlined and more reproducible, things like that. Of course, some other people think a little bit differently. They think you have to be perfect even from the very beginning, but then you have to invest a lot of resources on that to be perfect. But then you may miss the deadline because a small biotech is just moving so fast. That’s my strategy is move along and as long as it’s good enough to work and I don’t break it.
Ross Katz: That makes a lot of sense. So you come in, are there already a set of priorities that Immunitas had in place for you, like these are the questions that we want you to ask of our data, or were you given the creative freedom to understand what was happening at the company and identify the priorities?
Ming (Tommy) Tang: Initially, I was recruited to do new target identification, using machine learning approaches, still conventional machine learning approaches to find new targets. At Immunitas, we can generate single cell RNA sequencing data in house, and also single cell TCR sequencing data in house because we are an immuno oncology company and we’re really interested in the T cell phenotype. We have this 10x Genomics machine to do those single cell profiling, and then we send those libraries to Broad for sequencing. Essentially we grab all the data from the cloud and then process them in house. In terms of priority at that time, we just want to find new targets. The approach that we use is actually simple. For example with 10x Genomics, you can profile the gene expression level at single cell resolution, but you also get TCR sequencing data from the same cell so you can match them by the barcode. That actually gives you a really powerful data set because now you can correlate the gene expression pattern with the clonal expansion pattern. For example, the T cells when they recognize tumor antigens or neoantigens, they expand and then they replicate. You can correlate the gene expression with that phenotype and then try to find which gene is positively correlated or negatively correlated with that phenotype. We’re not using too fancy stuff here, and although we do use deep learning for other purposes, let’s just be honest, we don’t need deep learning and tell you the truth, we don’t even need machine learning. Sometimes simple statistics, like correlation, will just tell you which target to pursue.
Ross Katz: I love what you’re saying about just using simple statistics even to do something as complex as identifying drug targets from these correlated data sets. Are you giving formative input to the data that you want to generate from these next-generation sequencing instruments, or are you collaborating with biologists in the wet lab who are having conversations with you about where the analysis path should go from here?
Ming (Tommy) Tang: Exactly. That’s where my strengths actually play because I had the experience of doing wet lab stuff, and I know their language. I know how the experiment is carried out, and I had a clear question to ask. At Immunitas we sit together, the wet biologist and the computational biologist, we sit next to each other, and when we have questions, we’ll always ask them. Before they carry out any experiments, we’ll discuss the purpose of the experiment, what should we do. For example, we were really interested in target identification in the dendritic cells, which is a commander for T cells and NK cells, it activates both T cells and NK cells. The publicly available data, because dendritic cells probably comprise maybe 1% of the total immune cells, it’s just very few if you don’t enrich them. For public data sets, you maybe find ten cells per sample, that’s not enough for you to study. Then we can talk to the wet biologist in the company: we want to study this type of cells, if you do single cell RNA sequencing, can you please enrich it first. They will do flow cytometry to enrich that subpopulation and send for sequencing, then we can get a thousand cells per sample, which is a lot more than the publicly available data. Computational biology, in my perspective, is really to use computational skills to address biological questions and you need to work really closely with wet biologists. When the data actually arrive, when we’re doing those analyses, we just turn back and say, we found something interesting, can you tell us if this is making sense or not. We learn a lot of immunology from them, and we also explain those complicated methods, for example batch correction, what does that mean to them. We learn from each other, and I really like that.
Ross Katz: So there are situations where you’re picking up on something you think might be signal and you’re asking the wet lab biologist, what do you think of this? Could this be a signal of something important? And then maybe there’s a data gathering exercise where you’re figuring out how to amplify that signal so that you can validate what you’ve seen in the data. But then there’s other situations where there are existing questions that someone has in the organization and so you’re collaborating to chase down whether or not the answers are there. Am I understanding that correctly?
Ming (Tommy) Tang: That’s absolutely right. We learn from each other. Sometimes biologically that’s just a technical effect. If you just talk to them, we found this subset of cells that is interesting, they’ll tell you, that’s probably when you process the samples, it may be a heat shock, it’s upregulated purely technical. They can tell you that. On the other hand, we can also tell them, let me touch on this a little bit. For example, we are really interested in this gene called TLR9 and try to quantify this gene’s expression. From the RNA sequencing data we find why is the gene expression level so low, it should be high. When we dig into it, it really depends on what kind of computational tools you use for RNA sequencing quantification. Those more conventional methods, you use STAR to align them and quantify them using HTSeq or featureCounts to quantify the number of reads that fall into a certain gene. But for those methods, they cannot differentiate the reads that fall into exons that overlap from two different genes. For example, if you have overlapping exons, the reads that fall into those overlapping features will be discarded and then you get reduced number of counts. But we know the state-of-art, what are the tools that we can use, and then we use this other tool called Salmon or Kallisto. Those tools are alignment-free tools, they are not checking those reads to the transcript base-by-base, they just check whether it’s compatible with the kmers from the transcriptome model, and it’s much faster and it also tends to be much more accurate. They will assign those reads in the overlapping features probabilistically to different genes and then you can rescue those counts. Then they also learn that from us, so it’s a mutual learning process.
Jason: Are you a biotechnology company looking to unlock the potential of your business data? CorrDyn can help. We’re an enterprise data specialist that helps companies working in life sciences make smarter, strategic decisions. From developing the right data strategy that starts with our data maturity assessment to building and delivering bespoke technical solutions, we are equipped to tackle the most complex data challenges. We have partnered with dozens of high-growth organizations from manufacturers of custom oligonucleotides to molecular diagnostic companies to achieve data competence. Whether you need to supplement existing technology teams with specialist expertise or launch a data program that lays the groundwork for future internal hires, you can partner with CorrDyn to unlock the potential of your business data today. Simply visit connect.corrdyn.com/biotech to learn more. Now back to the show.
Ross Katz: Is there a typical day for you in your role as director of computational biology or is it just every day you’re coming in building on the work you did the previous day, like answering the new questions of the moment?
Ming (Tommy) Tang: That’s a good question. As a manager, first of all I had to set up a goal for the year, there’s a long term goal, usually we have a document to write down what we need to accomplish in this year. You need to have one or two big things that you want to accomplish in this year. Then you also set quarterly goals and then weekly goals, and you assign those tasks to team members. My typical day as a small biotech, even as I’m a manager I still do a lot of hands on analysis. At the same time I also need to supervise three members, and my job as a manager is also to develop them because I want them to be successful and their success is my success. I also learn those management skills and leadership skills along the way. Even when I was at Dana-Farber I took leadership classes and some coaching sessions to learn the responsibility of a manager is really to help your direct reports to be their best version and I fully agree with that and that’s what I practice.
Ross Katz: That makes a lot of sense. Can you give some examples of or a template for what an annual goal looks like in your computational biology role?
Ming (Tommy) Tang: For example, when I just joined the company, maybe one of the goals is, let’s take all the publicly available data and harmonize them and establish a pipeline that we can process them uniformly in house. That’s what we did. I actually wrote a Snakemake pipeline just to process all this data uniformly on the cloud and we also developed this tool to annotate those cell types uniformly. Then we can have the same cell type that is annotated for different data sets because that’s essentially what you want to do. Of course different data sets of different cancer types, let’s check the gene expression of a particular gene across different cancer types but in a specific cell type. But if you don’t harmonize the cell type label across different data sets, you can’t do that. That’s one of the major goals for a year.
Ross Katz: Right. So it’s setting up the scalable data infrastructure that’s going to enable you to do the kinds of analysis repeatedly and efficiently over time.
Ming (Tommy) Tang: Correct. Of course, if you ask me again today, I don’t know, I might just turn to some other vendors. It really depends on your strategy. If you want to be quick, then you can spend some money working with vendors. Vendors such as Elucidata, they actually harmonize all this single cell data for you, and then you can just request. Another platform would be BioTuring, they are really good at single cell RNA sequencing data curation; they also expanded to spatial transcriptomics as well. High quality data and harmonize all this metadata using ontologies so you can easily search. But we still need to have some in house capability because if you are relying on them, they process data, it takes time. If there’s a newly published data set you want to do and then request it, it will take some time for them to do it for you. That’s why we still have some in house capability, we can crunch the data ourselves when needed.
Ross Katz: Interesting. I know that you’re monitoring the research as it’s come out to anyone who doesn’t follow you on LinkedIn, I would recommend following you on LinkedIn because you share a lot of papers. So are you keeping an eye out for new data sets that get published that you can then pull into your pipeline and combine with the infrastructure that you’ve already got?
Ming (Tommy) Tang: That’s exactly why I’m on social media. One of the main reasons is that I follow the most recent papers, bioinformatics tools, and I’m aware what data sets could be available, what data sets are relevant, especially if it’s single cell immuno oncology, I’ll just write it down, we have an Excel spreadsheet. You always start with this Excel spreadsheet. I will put the title of the paper and also the link of the data there so in case we want to use it, we can process that by ourselves. I also share a lot of bioinformatics tools every day. Especially for single cell genomics, you probably see tools every day published. It’s just too much and they can be overwhelming. I read the abstracts a little bit and I’ll know this could be useful for other people, and I would retweet or post it on LinkedIn, but I also curate them in my GitHub repository. That’s my superpower: if I write it down when I need to use something, when I have a problem I can remember it and quickly find that link and try that tool. Before I try that tool, I will take a look at their documentation. No matter how good your tool is, if the documentation is not good, I’m not going to use it because I don’t know how to use it.
Ross Katz: Yeah. I’m interested in your ideas about the tool proliferation that we see. In what ways is that enabling you to do new things that you’ve never been able to do before? And in what ways is it confusing people in the field and making it harder to do the kind of work that you do?
Ming (Tommy) Tang: It’s a double edged sword. On the one hand, there are so many tools published for single cell specifically, and you really need to be able to pick a good one. But without testing it, it’s really hard to know. That’s one of the problems, you just don’t know which one is good. As you said, it’s probably too fragmented. In terms of data analysis even for single cell, from 10x Genomics, most people use their platform and they have their Cell Ranger software to do those pre-processing. Some other people use Kallisto or Salmon, they also have a single cell version to process those data. It turns out we were actually using a quite old WDL workflow on Terra and it was using a really old Cell Ranger version, 3.1 something. Then I tried Kallisto or Salmon, it actually can detect many more genes and many more transcripts per cell and also pick up more cells and can recover more cells from your experiment. Different people use different tools and the count table eventually people get from different methods, it’s really hard to do a meta analysis. How do you harmonize them if they are from different process pipelines? What you can do is start from the raw FASTQs and process them all at the same time but that also takes computing resource and that can be a lot of work.
Ross Katz: Those are some interesting challenges. You mentioned that you’re in Google Cloud. I’m interested, how do you use the cloud to enable the workflows that you’re talking about? Where does the cloud come into play?
Ming (Tommy) Tang: The good thing about cloud is that it’s so scalable because you can use as many CPUs as you want and as big a disk as you want. For single cell data the challenge is that data can be big, for example some data sets can have millions of cells, it’s just hard to do that on my local desktop. For pre-processing, as I mentioned we use this Snakemake pipeline, and it’s also Dockerized for the reproducibility purpose, and it’s easy to run. We wrap it in a Docker container and just one command and with the same input format we can crunch the data regularly and without too much human intervention, so that helps to improve our reproducibility. The other thing is also about data storage: on the cloud you have a lot of space, you just need to pay for the storage. Just be careful, it can accumulate so fast because the raw FASTQ files are so big. Make sure you put them in the cold storage, it’s much cheaper. You have to really organize how you store, for example, the raw input, pipeline output, and also the standardized downstream analysis output. You have to organize all this infrastructure.
Ross Katz: That makes a lot of sense. It’s like you have these raw files that you’re storing in a cloud storage environment. And then you have these pre-processing pipelines that are Dockerized that you’re deploying whatever in a virtual machine or in Cloud Run or something like that. Do you use a cloud data warehouse like BigQuery at any point in that process or are so many of the analyses that you’re doing using file systems and Docker containers that there’s no room for that sort of thing?
Ming (Tommy) Tang: We haven’t used any of those big data warehouse stuff yet. It really depends on what type of company you are in. If your company is generating tons of data in house, then it can be beneficial to have such a thing that you can easily query the data. For us, one of the major challenges is metadata tracking. How do you track those metadata? Even for public data you have to clean them, go to the supplementary information tables and go to GEO, they may be hidden in some places in the GEO record. It’s a lot of work. The spreadsheets are also not greatly formatted; if you read into Python or R, then you have to spend some time to clean it up. Different data sets may have different metadata columns, and even if they have the same column meaning the same thing, they may be named differently, and the content of the column may also be different. For example female could be F or capital Female or lowercase female. You need to harmonize all this stuff. That’s one of the major hurdles for machine learning models or any data science models to work on those data.
Ross Katz: So for the metadata, do you store the metadata in a database or something like that and then reference the metadata as you’re pulling the raw data through the files or do you just write the metadata back to files that you’re storing in cloud storage?
Ming (Tommy) Tang: That’s a great question. Now we just have a spreadsheet. When you have small enough data sets, if you still have less than a thousand, I would say a spreadsheet still works. Perfect is the enemy of good. But if you are generating high throughput data every day it can easily explode. I believe even Broad, when they first started the CCLE, the cancer cell line project, they also started with a spreadsheet and I heard that until one day they couldn’t handle it, they said let’s do something more complicated.
Ross Katz: That makes a lot of sense. So you mentioned machine learning and you’ve talked about situations where you need machine learning and the kinds of analysis that you’re doing in the situations where a statistical test or a data visualization is good enough to tell you exactly what you need to know. How do you determine where machine learning should best play in the kinds of computational biology approaches you’re using?
Ming (Tommy) Tang: It’s a great question. As far as I can see, I always try a simple machine learning model first. For example, in this TCR versus gene expression example, we just used conventional methods for that purpose. Random Forest usually performs well and even logistic regression is really interpretable; you can get the weights of the features and whether it’s positively or negatively correlated from the logistic regression. But for other more complicated data analysis, for example image analysis, that’s where deep learning shines because those convolutional networks really do well on image processing. You may also be aware of AlphaFold, they can predict protein structures really well. In those cases deep learning can benefit computational biology. We actually did try deep learning using AlphaFold 2 to predict a potential ligand for a target that we are interested in. We got some top hits and we asked the wet biology team to validate, and it didn’t work. But we tried. It’s still not there. Also recently there are several drugs on clinical trial which are designed by generative AI, small molecules and also large molecules like antibody sequences. They are still in early phase, phase one, phase two, and we still need time to see how they perform in the later trials and whether they can hit the market. But definitely when we have enough data, deep learning or machine learning approaches can have really great potential.
Ross Katz: So I haven’t heard you talked about a lot of images in your data pipeline today at Immunitas. Is that something that’s on the horizon for you or is that something that you’re already working with?
Ming (Tommy) Tang: We don’t work on a lot of image data, but recently we purchased this spatial transcriptomics platform called Vizgen. You can look at single cell gene expression level on the tissue slice, but you can only look at five hundred some genes. It’s an image-based platform. For that you have to do cell segmentation and then assign the transcripts to each cell. We’re doing a little bit on that but we are using an existing algorithm like Cellpose and it does a pretty decent job at this moment. We’re not actually working on image processing per se, but we touch a little bit on that.
Ross Katz: Interesting. I really appreciate the pragmatic approach that you use where you’re using the simplest, most understandable methods at any point in the process and you only move to the more complicated method when you feel like the question necessitates the application of that more complicated method.
Ming (Tommy) Tang: Exactly. That’s why with the AI models, people, there’s really hype there: we can address this using AI, we can address that using AI. I think we still have a long way to go. Really don’t overcomplicate the problems. If it can be answered using simple methods, then just use the simple methods. It’s more interpretable with the simple methods, and with the more complicated method you don’t really know, especially for deep learning models, it’s a black box: how many layers, what units you need to have for each layer, there are so many hyperparameters you have to tune and it’s just not worth it.
Ross Katz: Interesting. Given that AI’s not there yet, as you look out one year or two years from now, where do you see the biggest opportunity for computational biology to really contribute to immuno oncology? Is it just application of the same methods but with better data or is it are there methodological improvements that you’re going to see and if so just any insight you can share there.
Ming (Tommy) Tang: I still think the data is the key. If you have better data, high quality data, more data, then your machine learning models that can be trained on those data may have some sensible predictions. That’s the thing, just like Google or any of those big companies, you’ll see who will win in the AI world. I would say Google because Google has the most data. For example YouTube, all the data on YouTube they can transcribe into words and then they can train their large language model on that huge data set. ChatGPT is really popular and OpenAI has pioneered that, but if they don’t have data that’s their biggest disadvantage because if you don’t have data then you don’t have the moat. There are many open source large language models, they can even perform as well as the OpenAI model. The model itself will become a commodity. But the data really is the key, it’s your moat, and then you can really apply your machine learning approaches on those data. That’s why you really need to have a platform that can generate those high throughput data in house, and that’s what I’m thinking about. Platform first, AI on top, so you can generate data, then you can apply your models on those data and that can potentially give you, for example, new targets and even new molecules, drugs, protein sequences, de-novo protein engineering and antibody engineering that can potentially give you some sensible results and can speed up the whole process of drug discovery and development.
Ross Katz: So when I think about data as a competitive advantage at a company like Immunitas, there’s the quality of the questions that you’re asking and the collaboration between the computational biologist and the wet lab biologist, and then there’s the quality of the data that you can generate and the volume of data that you can generate that provides insight into those questions. And those are the sorts of places where you can really gain an advantage and then the methods that you use to analyze those data are relatively commoditized. You just need to have sort of the capacity in house to do that kind of analysis. Is that right?
Ming (Tommy) Tang: Exactly. I want to echo your first point, really you need to ask the right question. That’s even more important than your model, than your data. Because if you ask the right question, you may not need huge data sets. You just need the right specific data sets. If they’re not available, then we can design an experiment and generate that data set. I really appreciate that point.
Ross Katz: As a practice, when you’re growing a computational biology team, I imagine that one of your roles as a manager is to help your people to ask better questions. How do you help them to do that?
Ming (Tommy) Tang: I usually sit with them to draft out their career development plan. What they want to learn. For example, if you want to learn a little bit more about biology, immunology, I will encourage them to take courses on Coursera, edX, those open platforms. If they want to learn more machine learning or deep learning, I encourage them to take courses as well. I also bought those machine learning books. By the way, this one is really good: Josh Starmer, he has this really nice machine learning book that uses cartoons and plain language to explain those complicated models. I find it’s really useful. I bought several copies for them and also bought several more practical coding books for them. It’s a learning process and I actually learn together with them. You have to foster a really good learning environment together as a group and then we share different expertise with each other. We also have this Friday teach and learn session: each time one person will present one thing that they know well and maybe other people don’t know so we can learn from each other. I encourage them to think about questions proactively, how to think in the big picture. Not only the technical things, that I can write this code to analyze this data, but before you even analyze the data, what kind of question are you asking? Is this data set suitable for answering that question? That’s more important.
Ross Katz: I heard a couple of things in there. So at an individual level, you want people to have the full scope of capabilities that they need as a computational biologist. So you want people to have the biological grounding to be able to understand what’s happening in the data and in the wet lab. You want people to have the statistical and technical grounding to be able to understand how the data sets are structured and how to clean them and then also the methods that they apply when they’re analyzing that data.
Ming (Tommy) Tang: Yes. Correct.
Ross Katz: Interesting. And then the other thing I heard was the culture that you foster as a team where if all of you are going down this journey together, then you’re all going to learn much faster than if one of you is just consuming resources all day on arXiv, for example.
Ming (Tommy) Tang: I really appreciate this saying: if you want to go fast, go alone, but if you want to go far, go together. You can always learn much more if you learn together.
Ross Katz: I’ve noticed that you share a lot of resources in public about bioinformatics, about computational biology. What motivates you to share and build community around computational biology?
Ming (Tommy) Tang: Initially I didn’t think too much. I was just posting and sometimes I was posting just complaints about bioinformatics or how difficult this problem is. But gradually I started to post the solutions that I have or the new tools that I found that could be useful for other people. Because I went through the pain. I didn’t know where to start ten years ago. Although I was speaking about my transition from a wet biologist to a computational biologist for ten minutes or so, it really took me ten years of grinding, learning, relearning, frustration. I didn’t know how to start. One of the purposes is that I want to help others to avoid the same pain. At least when they start, if they decide to make the switch or they want to learn, they know where to start and they have guidelines on what course they can take, what books they can read. I have some posts about the books that you can use, the courses that you can take. Ten years after I started, there are much more resources there, and I also helped to curate them, and it’s much easier for you to start today. One of the purposes that I started to share is to help others to avoid the same pain.
Ross Katz: Yeah. Just at a high level, as you think about your journey and all the resources that you share and the sort of people that you manage, for people who are interested in getting deeper into computational biology, what would you recommend to them as a starting point?
Ming (Tommy) Tang: Starting point. I have a lot of resources on my GitHub repository and also my own website, blog posts. I write many blog posts to explain some technical stuff, and sometimes it’s more general talk about AI in general, has it revolutionized drug discovery or not. I write all sorts of things, so you can go there and take a look. In one of the GitHub repositories, I listed all the courses that you can take. Make sure you go there and take a look, so you can follow the guidelines there.
Ross Katz: That sounds good. We’ll make sure to include some of those links in the show notes and in the blog post that we write about this episode. As we bring the conversation to a close, are there any resources that you’ve discovered recently or ideas that have come about recently that are on your mind that you feel are at the leading edge?
Ming (Tommy) Tang: You were asking what kind of platforms we were using, and I told you Google Cloud. To be honest, establishing a Google Cloud platform can be painful as a small biotech, because you have to deal with all this: to set up a GCP you have to set up an organization, then set up the administration, the security, all sorts of things. Recently I had this problem: when I attached a disk to my compute engine, it turns out that disk was too small. How can I pause the engine and increase the disk? Then I have to Google around, find the manual and do that, some commands you can type. It can be challenging sometimes. Recently there are other platforms emerging. Before that there are old platforms like DNAnexus, Seven Bridges, those are old companies, and some new companies for example Watershed, Latch Bio, and even Deep Origin. Those types of platforms are trying to solve these pain points for bioinformaticians, so they have their platform ready. You don’t have to set up all this GCP stuff or Amazon cloud computing. You can just say, I want six CPU, I want this big machine, you can easily select them and they will handle all this security, all this stuff behind the scenes. It’s much more convenient for us and saves us time. I look forward to trying some of them in the future.
Ross Katz: So these are cloud computing platforms that are designed specifically for the workflow of the computational biologist. I’ve experienced the pain of being in different cloud environments and being challenged to figure out how to accomplish a particular task, and there’s a lot of documentation and a lot of tools on the table and so it’s very easy to get confused. What are the aspects of a platform like that that you feel would be most valuable to you in moving from one cloud to the other?
Ming (Tommy) Tang: One of the main things I care about is reproducible computing. I’m an advocate for reproducible research. How can you integrate that into the platform, it’s important. For example, now we’re using a Docker container, and also Git version control for the code. But for those files that are generated by the pipeline, you also need to version control the data. So if a platform can help to advance all those aspects, the computing environment, the code, and also the data version control, all three of them, that could be really helpful. For example, you generate the data, you really know there’s probably a unique ID that is attached for that data. You can directly trace back how you generated that data, what command did you use, when did you generate that data. That could be really helpful.
Ross Katz: Each time you rerun that pipeline, it can be considered an experiment and you want to track the different metadata surrounding that experiment including the methods that were applied, the compute that was used, the specific data resources that were input, the specific data resources that were output, and be able to return to them later on and look at differences between them over time.
Ming (Tommy) Tang: Exactly. For now, what we are doing at this moment practically is, let’s run our stuff on Docker container, Git version control them on GitHub, and when you finish this analysis make a deck and then put the notebook link into the last slide. Then you can trace back how you generated the data in a not perfect way but it’s good enough for this moment.
Ross Katz: But as your team grows, as you’re doing more work for a larger organization, the complexity and the number of experiments that you’re running as a computational biology team is growing and you want to be able to navigate that data through a single interface. That’s really interesting. Tommy, this has been really insightful. I really appreciate you coming on. Look forward to connecting down the line.
Ming (Tommy) Tang: Thank you very much Ross, I really enjoyed the chat.
Jason: And that’s it for this episode of Data in Biotech. If you enjoyed the episode, please subscribe, rate, or leave a review in your podcast platform of choice. See you next time.





