Listen on
Overview
Biotech companies face immense pressure to develop accurate diagnostic tools quickly and reliably. The core challenge often isn’t the statistical model itself, but the integrity of the data used to build and validate it. Developing a blood-based cancer diagnostic, for example, demands meticulous attention to data provenance, potential biases introduced at the lab bench, and the careful construction of training sets that reflect the target patient population. Overlooking these early steps leads to flawed models, delayed regulatory approval, and significant wasted resources.
Michelle Wiest, Director of IVD Biostatistics at Freenome, brings a unique perspective from both academic research and high-growth industry roles. She understands the gap between theoretical data science and its practical application in regulated environments. In this episode, host Ross Katz talks with Michelle about how biostatistics guides diagnostic model development, the critical data quality issues that arise, and strategies for managing strict regulatory requirements with the FDA.
The conversation covers the evolution of data sets from convenience samples to controlled clinical trials, identifying and mitigating biases like late-stage cancer overrepresentation, and the careful preparation required to present data to regulatory bodies. Michelle also discusses the limitations and potential of real-world data and agent-based models in accelerating diagnostic innovation while preserving patient privacy.
Key Takeaways
Diagnostic model accuracy hinges on mitigating early-stage bias and lab artifacts.
Models trained on convenience samples often overrepresent late-stage disease or embed lab-induced noise. Actively balance training data with early-stage cases and ensure reliable lab randomization to prevent models from learning artifacts. This collaboration between lab staff and machine learning teams is critical for building trustworthy diagnostics.
Proactive regulatory engagement shapes successful diagnostic approval.
FDA interactions are iterative, starting with pre-submission meetings to align on study designs and assay challenges. Seek early buy-in on methodology—especially for extrapolation or complex modeling—to avoid costly re-work and ensure clinical relevance. Regulatory bodies prioritize understanding how a diagnostic will perform in real-world clinical use.
Data quality extends beyond complete fields to sample provenance and metadata detail.
Issues like mishandled samples, missing age data, or unstaged cancer diagnoses degrade model reliability. While classic imputation helps, internal data generation demands strict data architecture and engineering controls. Understand the limitations of external “convenience” data sets; not all data can answer every question.
Meaningful data collection, not just machine learning, is the true bottleneck for diagnostic innovation.
While advanced analytics accelerate discovery from existing data, the greater need is for targeted, high-quality data collection tailored for specific scientific questions. Real-world data, though valuable for post-market monitoring, often lacks the specific fields or controlled conditions required for initial diagnostic development. Leadership must understand these data limitations to set realistic development goals.
Related: CorrDyn helps biotech and life sciences companies develop effective data strategies. We specialize in ensuring data quality for critical applications, from foundational engineering to advanced machine learning model development. Read more about how biotech manufacturers access data value.
Full Transcript
Jason: Welcome to Data in Biotech, a podcast from CorrDyn where we explore how companies leverage data to drive innovation in life sciences. Every two weeks we sit down with an expert from the world of biotechnology to understand how they’re using data science to solve technical challenges, streamline operations, and further innovation in their business. This week we’re excited to be joined by Michelle Wiest, director of IVD biostatistics at Freenome, a high-growth biotech company that helps create tools to prevent, detect, and treat disease. During this interview, our host, Ross Katz, speaks with Michelle on the use of biostatistics in the field of diagnostics, the different types of data sets that are being used to develop diagnostic models, what biases can corrupt diagnostic tests and how to catch them early, and how to prepare to present data to regulatory bodies such as the FDA. Here we go.
Ross Katz: Michelle Wiest, welcome to the Data in Biotech podcast. Excited to talk with you today.
Michelle Wiest: Yeah, thanks Ross. I’m happy to be here.
Ross Katz: So just to get us started, in a minute or so, can you tell us about your career to date?
Michelle Wiest: Sure, I’ll try to cram it into a minute. It’s been a winding path. I am an epidemiologist and a statistician by training. I graduated from UC Davis and started at a small biotech company in the Sacramento area as my first job. I actually started there after I finished my master’s in statistics and was working full time while trying to finish my PhD. It took a little longer to finish. That was a company called Lapomex. Then wanted to head up to Idaho where I was born, like a salmon returning to my native spawning grounds, I really felt drawn to North Idaho. Headed up there and worked at the University of Idaho for 12 years. I ran our statistical consulting center there, trained up our students in how to synthesize the information they were learning and talk to people and provide recommendations. I also did a bunch of different research projects, collaborated with a bunch of folks from across the university. I did take some time and spent a year in Australia and worked at Murdoch Children’s Research Institute as a senior biostatistician. That was wonderful. Had to come back because my mom was ill. Returned to the university and she passed away which got me reevaluating what I wanted to do and I always enjoyed the fast pace of industry. I wanted to get closer to patients. I ended up coming back to biotech in the diagnostics field. I’ve worked for a couple companies now as the director of biostatistics, building teams to support the development of blood-based cancer diagnostics.
Ross Katz: It sounds like the kind of work that you were doing in academia was somewhat different than the sorts of things that you’ve done in industry. Are there any differences you would highlight?
Michelle Wiest: It was very different focus. I was working with clinicians directly, small rural hospitals, bigger rural hospitals, and supporting small clinical trials there. I was also leading research using public databases, secondary research on data that was collected, so very epidemiological type of research, trying to tease out things that were associations and things that were driving disease, and more on the methods development side. Can we do a better job of estimating the rates of suicide in rural counties, for instance? It’s a challenge because there’s just not a lot of people, so there’s not a lot of data, solving challenges like that. That was academia, and also trying to open opportunities for students, for your graduate students to get exposure to these things, thinking about projects and how can I involve students in this and what are they going to learn and take into their career? In industry, everything is very team oriented. There’s set goals that everyone is working toward, which is really nice, and I enjoy that back and forth with different functional groups and figuring out how we’re going to tackle challenges that are thrown our way during the development. It’s also a bit more regimented in that you’re following guidelines, you’re following recommendations from regulating bodies. You have this lens of how does this fit into what the FDA is going to want to see, and if the assay doesn’t lend itself to this particular approach, how do we defend that? It’s a bit more of a regimented approach from the statistical side of things and the study design side of things.
Ross Katz: That’s interesting. I think later in this conversation I’d like to return to that idea of the way that the work that you do needs to be tailored to the expectations of the regulatory bodies. But as we’re talking about your career and how you reevaluated what you wanted to do and decided to come back into biotech, can you talk a little bit about why you do this type of work and what motivates you?
Michelle Wiest: Like I said, I wanted to get closer to patients and actually having a direct impact on quality of life of patients. Academia is wonderful and you do have opportunities to advance the science, but there’s still that ivory tower thing where it may not necessarily be totally integrated and grounded in what can be applied in real life. Depending on the friendliness of your tech transfer office at the university, it may be pretty difficult to get things that you’ve developed out of the university and into the market. You’ve got to get it onto the market if you want it to reach patients. I was able to get around that by working directly with hospitals and directly with public health departments. Their needs are much more logistical and maybe not as complicated approaches that we were developing at the university and that were interesting from an academic perspective. It was a little bit hard to marry those two. I did it, I got some students involved in understanding what that is. But in order to really make an impact on patient health, I think industry is where you go unless maybe you’re at Harvard and you have a great tech transfer and people spin out businesses all the time. Coming back to industry allowed me to help move forward products that were going to help identify cancer earlier, to tailor identification of cancer. The thing that got me so excited and I was like, yes, this is the right decision, is being able to work on minimal residual disease identification. That’s where you are taking the signature of someone’s tumor and characterizing it and then developing a test that is going to specifically look for those mutations, personalized medicine, to see if their cancer has come back. That just got me super excited and I was like, I want to work on this, this is game changing.
Ross Katz: That’s absolutely fascinating. Can you give a little bit of insight into the way that data science, biostatistics, the use of data is applied in that diagnostic approach?
Michelle Wiest: It is a pretty multi-disciplinary development for these genomic and multi-platform assays. Ideally, you have a really tight collaboration between domain experts, folks that really understand genomics and the types of mutations that are associated with which cancer. That may be a computational biologist, it may be a nurse oncologist. It depends on the individual. People these days seem to be much more multi-disciplinary in general. It could be computational biologist who has both mathematical modeling training and biology. It could be somebody that is well versed in proteomics or metabolomics. These folks would be hopefully trained in both the laboratory science and in the mathematical or statistical modeling. Or they may have expertise that’s more on data management side, depending on which program they come from, it could be more comp-sci or it could be more statistical. Then you have statisticians and epidemiologists who are going to think about what kind of performance do we need to see from these tests? And what is the market like from an epidemiological perspective? How many people are out there, how many people do we think we could detect with a new test? Working with the statistician, what is the clinical efficacy likely to be? Then you put the specs together for your ideal product. Meanwhile, you have computational biologists, maybe machine learning folks, that are using your -omic database and samples from that population to train a predictive model. It could be a simple logistic regression, it could be a black box. The FDA accepts both types of algorithms that would say, this person is positive or this person is negative. That’s the general approach that is taken to developing those.
Ross Katz: Interesting. It sounds like there’s a lot of different models that are under development simultaneously for different types of outcomes that are trying to be driven by the organization. At the end there you were talking about the computational biologists or the machine learning experts, that’s the model that’s being deployed in a productized fashion as part of the diagnostic tool itself. But then there’s also the statistical models that are being used to understand what the impact on the population of this diagnostic tool is going to be, and then also some understanding of the economics of the deployment of the diagnostic tool for the organization itself. Am I thinking about that right?
Michelle Wiest: Exactly, exactly. You got it. There’s the economic impact and then there’s also will the company make money? Is it going to be reimbursed? How many people are going to be actually able to have access to this? Is the company going to make money?
Ross Katz: Right, and there’s another set of domain experts that need to be brought in for that portion as well, your legal or insurance experts or your financial people who are understanding the cost side of the equation. When you talk about working with stakeholder groups across the organization, I’m imagining that those are the kinds of partnerships that you need to have. Is that right?
Michelle Wiest: Absolutely. And physicians as well because you need to understand how the test would be used in the trenches. What challenges they might face in communicating results. It’s definitely a lot of moving parts.
Ross Katz: For sure.
Jason: Are you a biotechnology company looking to unlock the potential of your business data? CorrDyn can help. We’re an enterprise data specialist that helps companies working in life sciences make smarter, strategic decisions. From developing the right data strategy, that starts with our data maturity assessment, to building and delivering bespoke technical solutions, we are equipped to tackle the most complex data challenges. We have partnered with dozens of high-growth organizations, from manufacturers of custom oligonucleotides to molecular diagnostic companies, to achieve data competency. Whether you need to supplement existing technology teams with specialist expertise or launch a data program that lays the groundwork for future internal hires, you can partner with CorrDyn to unlock the potential of your business data today. Simply visit connect.corrdyn.com/biotech to learn more. Now back to the show.
Ross Katz: Let’s take an example model for the actual diagnostic model that’s going to be put out in the field. I just want to walk through these different types of models and understand what are the data sets that are being used for the modeling, what are the methods and approaches, what are the outcomes that are being modeled, just to do a little bit of a deeper dive into the way that these models tend to be structured and who tends to build them. For the diagnostic model that’s doing in vitro diagnostics that’s trying to detect cancer early, for example, what are the sorts of data sets that go into a diagnostic model like that?
Michelle Wiest: It may be that you take advantage of a study that has already been conducted and proteomics was run on it, and you use that data because there were also a few people that had this outcome of interest. You might start with a convenience data set. It’s already there, the data are already there, and you start getting an idea of what kind of associations you’re seeing between your outcome of interest and these hundreds of measurements. Sure. What you would do, probably simultaneously as you’re working on that, is you start building up a sample bank of your own. You might purchase samples from a supplier who collects blood or if you’re doing solid tumors, they’re collecting biopsy slides. You go into a more controlled designed phase where you’re starting to verify what you found in that first data set. In that first data set, you might be using things like holdout sets and stuff to help with your model training. There’s any number of different approaches to that first discovery step. The next step, you’d want to make sure because you’re probably running this through your own laboratory and you want to make sure that you’re getting as much randomization into that process as possible because, say, you purchased all your cases first and then all your controls second and those just went through the laboratory in that order, then you could very well be training your model on artifacts because the noisiest, most unreliable biomarkers are the ones that will be super flaming hot because of some systematic problem with the laboratory runs. Maybe they’re super sensitive to temperature. One was a hot day and one was a cold day or it was humid and then it was dry. Laboratories are climate controlled, but not to perfection, for instance. Those are things that you really want lab staff and your machine learning team working together and the lab staff understand what the machine learning group and the discovery group are trying to do. You want to make sure that’s in alignment, they understand why the randomization throughout the process is important. Then you might do a completely new classification to see if you get the same results. You might try to just verify what you’ve done in the other one. That goes on. You’ve got maybe a convenience data set, you’ve got your semi-designed data set with convenience samples maybe that you’ve purchased. You keep building that up, maybe you collaborate with a physician’s office and you do a final feasibility study on what you think is your diagnostic test. You iterate and you’re like, nope, I think this is it, and then you do a feasibility study and evaluate is it doing what I think. If it looks good in that feasibility study, then you start moving on to development.
Ross Katz: Got it. Once you start moving on to development, that’s when the regulatory and compliance concerns start to come into play.
Michelle Wiest: Now you’ve got to keep track of your design history, and you probably started that in feasibility because you would also start doing some guardbanding studies for your laboratory work. Then you move into development once the feasibility is all there. It’s pretty prescriptive in terms of what the FDA is going to want to see in terms of analytical performance. Once you’re looking good on your analytical performance, then you move on to looking at a big data set for your clinical validation. That’s the last step. You’re looking at that sensitivity and specificity and getting really good estimates of those in your intended use population.
Ross Katz: That clinical verification, that’s the clinical trial period? That’s the big data set that we’re talking about? Understanding how you build up that data set from the very beginning is really interesting. I appreciate you sharing that. It sounds like there are a lot of artifacts that can pop up from all of the variables that can enter the equation. There’s a big concern about the bias of these diagnostic tests that might lead to a lot of wasted time and resources if you’re not catching it early. Would you mind sharing what are the types of bias you look for and how you go about ensuring that bias doesn’t enter the equation?
Michelle Wiest: There’s a couple of different things. I touched on the laboratory and artifacts and you could end up building a diagnostic test on artifacts if your machine learning folks don’t understand what was going on in the laboratory and accounting for that. The other thing that can happen is that you can have a data set where you have mostly really sick people. You’re catching really late-stage cancer in those and you’re training on late-stage cancer and then assuming that it’s going to translate into earlier-stage cancers as well. What can happen is you start off with a really big effect and then you regress to the mean. You may train on these more severe cases and the way you combat that is to make sure that as you’re iterating through those early studies, you are actively trying to get more balance in that training set. If you can’t get earlier stages, you can do weighting of stages, but again, that’s assuming that you’ve got representation and maybe just the handful of samples that you have in those early-stage cases. You can do some statistical adjustment there, which will help, and you should do that early on. But the only way to ensure that your product is going to perform well in your intended use population, which is actually probably going to have more early-stage cancer, hopefully, that’s the purpose, we want to catch it earlier, is to train on early stage and test on the correct distribution of your intended use population.
Ross Katz: That makes a lot of sense. Throughout that early part of the process, you’re acting as the curator of this data set and trying to bring balance to the data set or identify areas of imbalance and account for those in the way that you can, but also make it a priority to balance out the data set and remove the bias.
Michelle Wiest: In that case, your dev team is working really close with your clinical team, who are the ones that are sourcing those samples for you or identifying collaborators to help build that model.
Ross Katz: Right. You have this wish list as you’re going for the sorts of samples that you would want if you had the opportunity to purchase them and you’re also, I’m assuming, having to work with stakeholders who are controlling the resources that are available to go out and purchase the samples that you need to accomplish the clinical goals that you’re trying to accomplish.
Michelle Wiest: Exactly. That’s it.
Ross Katz: Are there data quality issues that creep in as well? Do you end up with samples that don’t meet the quality standards? Diagnostic tests that are providing inputs to your models that are not quite up to the standards that you would want them to be? Just curious if that occurs and if so how you deal with it.
Michelle Wiest: Never happens. All the data is always perfect. Yes, you mentioned sample quality. That’s definitely an issue, especially if things have been stored for a while or mishandled, maybe sat in 90-degree weather for a few hours and everything thawed. You want to have a good understanding of the history of your samples and it doesn’t mean that they’re all ruined always, but if they’ve been through some stress, if things look weird for that sample, you probably don’t want to use it. Other data quality issues that plague data sets are what folks would call the metadata surrounding those samples. Incorrect sex labels, which if you’re doing genomics, you can figure out. But ages are missing, maybe if this is a sample from a cancer patient, maybe it has the gross diagnosis, but it wasn’t staged so you don’t know what stage it was or how big the lesion was, so you have varying levels of data quality in terms of that metadata around those samples. Hopefully the company is in control of their own internal data production, so when they are creating their data sets, they’re working with data architects and database architects and data engineers to make sure that in-house they’re producing quality data and that should be on top of mind when creating data in-house. But you definitely need to be aware that the information that you’re getting from either the place that you’re purchasing samples or the clinical site may not be perfect, especially when it’s not a controlled trial where you’re coming in and saying, here’s your CRF and we need these fields filled out exactly like this and this detail. Mistakes happen still, but that improves data quality quite a bit when you’re actually conducting and leading the data collection.
Ross Katz: That’s interesting what you’re saying about the data sets that you have control over versus the data sets that you don’t have control over, and I’m assuming there’s your classic data science methods of dealing with missing data that you apply to the data sets you don’t have control over, but there has to be collaboration internally to make sure that the data that’s being delivered to the modelers, the ML and statisticians, is of a certain quality. Do you find yourself having conversations, or have you found yourself having conversations in the past about the data that we’re getting, we need additional data, we need different data, we need higher quality data, or how does that sort of thing go?
Michelle Wiest: Sometimes you’re able to go back and get it. If you’re working directly with a clinic, that data exists and you can go retrieve it with their help. In other cases, it’s just not there and you’re not going to get it because it was an anonymous blood donor. They didn’t collect it and they’re gone. You’re just not going to get it. When you mentioned missing data approaches, those are applicable when you have another part of your data set that is complete and you essentially can use that information to get an idea of the distribution and the marginal distributions of those data and extrapolate that to the missing data. Of course anybody doing that should be well versed in Rubin’s papers and understand missing at random versus missing conditionally on random versus you shouldn’t do imputation because it’s not a random process. There’s lots of different approaches to extrapolating that information and some are better than others. I wrote a paper with my colleague on the case where you have some data that have been grouped, maybe to preserve identities, and if you have another data set from say a larger population where you didn’t have to group those, you may be able to use that to inform what the distribution of the subgroups would be within that. It doesn’t mean that you’re able to assign a value to every single person, it’s a distribution, it’s a conditional distribution, it’s just helping you get a better estimate of whatever you’re trying to estimate from the population.
Ross Katz: Right. There is a whole other field of generating data and agent-based modeling that is more grounded in mathematics, you’re actually modeling multiple people or agents and what would happen under different conditions they were put under. This seems to be an emerging trend in the field, these sort of agent-based models. I’m curious what your view is on the appropriate use of those models and what if any limits there are or should be to that approach in evaluating the validity of a diagnostic test like the ones that you’re creating.
Michelle Wiest: Gosh, I think that if I were to use or apply agent-based modeling in the diagnostics field, I would probably do it on the economic side because they’re very good at modeling decisions that people are making and the conditions that those decisions are made in. Those have been really helpful for epidemiologists as well in modeling disease intervention. I remember during the very large Ebola epidemic in West Africa, there was a lot of concern that it was going to come over to the US, and the CDC and others were relying on those agent-based models to get an idea of what would happen with the interactions of these folks in cities and if we imposed this condition, which is an intervention to help prevent the spread, how would that change it, how would it slow it down or prevent it? Those folks also made synthetic data sets. That I can see would translate to the economic side of diagnostic tests in terms of uptake, in terms of the decisions that clinicians might make from the results of the tests. On the training side or discovery side, I’m not quite sure how it would be applied.
Ross Katz: You have a diagnostic test that’s in front of the regulatory bodies and these regulatory bodies have specific expectations for how you present data to them, and it seems from your background that you’ve had a variety of opportunities to represent statistics to these regulatory bodies. Can you just give us a sense of what are some of the considerations that you have to bring in when you start the process and when you’re preparing to present these diagnostic tools to regulatory bodies?
Michelle Wiest: First of all, you’re going to be working really closely with your regulatory group within your company. They are typically trained as lawyers, actually. You’ll be working with them about the strategy in terms of how things are going to be presented and managed with your regulating body. The first thing that you do, you come up with a plan, a design, saying I read the guidelines and this is how we are going to apply them to our assay. We’re going to challenge our assay the way that you want to see it because you don’t want to do these analytical tests in perfect conditions because that’s not real life and things are going to go wrong. The idea in those analytical validation studies is challenging the assay. With that in mind, you come up with these designs and you organize a meeting with the FDA and you work with them to make sure that your study designs are in line with what they’re going to see. They may have other companies that are working on very similar tests and they’re trying to hold those companies to the same standard. It is really important that you speak up front with your FDA reviewer and get their buy-in and feedback on your study designs because this is a really cutting-edge field and we’re learning more about what works, what doesn’t work, and they may have specific questions that they want to see and they’re like, you really should do this because they saw it fail under these conditions with somebody else. Yes, so one, you need multiple pre-sub meetings. Once you’ve got this agreement of these are studies including your clinical validation studies, this is going to be the whole suite and the totality of data that we’re going to send to you about our assay, you get to work. You’ve got this agreement, you get to work and you execute. Then you end up doing things in chunks and sending things in chunks and you’ll get feedback. They’ll come back and be like, this is interesting, but can you also make a table of X, Y, and Z? Then you scramble because you’ve got 24 hours to respond to them. All the statisticians and coders are going back into that study and making this table real quick and then writing up an explanation to answer their question and then sending it back and then you don’t hear back from them for a while, and then they send another question. There’s that process, and then you get a final decision once all of your modules are submitted and all of their easy questions are answered.
Ross Katz: During the formative input stage, are there any examples that you can share of the types of questions that you might be asked or the areas where regulatory bodies might be more sensitive to the way that you’re designing your clinical approaches?
Michelle Wiest: There’s two maybe that I’ll touch on. One is anytime that you are extrapolating data beyond, so if you tested under this range of temperatures, can you extrapolate in between these two temperatures? Does it make sense to do that? Definitely they’re not going to be happy if you try to go beyond the limits of your study. It may be any other kind of difficulty if you’ve got multiple systems that are coming together and there’s any number of combinations, you could end up with thousands of potential combinations of conditions, any kind of modeling that you might do to smooth out that design space, they are going to want to discuss with you and they have some great statisticians there that will give you guidance on that. If you put an idea out to them, they’ll say, why don’t you look at this paper? Or why don’t you look at this presentation I did on this type of thing? That is one area, when you’re doing extrapolation. The other area is when you’re using clinical samples in particularly analytical studies, they want your performance to be grounded in the clinical relevance. There’s just so many different ways that this could go, but always having in the back of your mind that they are thinking about how is this going to be used, what are the implications on the clinical performance from this study? They’ll dig into that as well, which has implications for the statistician because you are doing those power calculations and trying to figure out, if we need to have this sensitivity, what does that mean for if we measure the same sample again and again if we’re getting the same result in this condition? That’s the other area that they will lean into.
Ross Katz: That makes a lot of sense. The first point that you made about extrapolation, does that apply to some of the imputation stuff that we described earlier? Yes. Oh, do you have experience in that? Yes. I think that methodologically you always have to be sensitive to the way that you make assumptions in your models and imputation is a form of extrapolation. I was just assuming that’s another area where they might press you in terms of the models that you develop.
Michelle Wiest: Yes. That’s right. Imputation, Bayesian models—some of the best approaches use Bayesian approaches for imputation, what’s written down from FDA is the assumptions that go into Bayesian modeling can be dubious.
Ross Katz: I don’t want to get into an argument about Bayesian versus Frequentist at this time, but I think that it makes sense that the Frequentist approach is better established and more theoretically grounded inside of a regulatory body like the FDA. That makes a lot of sense. As we come toward the end, just zooming out a little bit, obviously there is a lot of innovation going on in this diagnostic space and it sounds like one of the big challenges is the collection and availability of data to develop these approaches. Given that we as a society have a collective interest in the improvement of these approaches, but also in regimes that protect medical privacy, I’m wondering if you see a path forward for how we enable this innovation at greater scale, how we bring more of these diagnostic approaches to bear.
Michelle Wiest: There is a big movement in not just diagnostics but also in drug development to use real world data and make use of all the data that’s being collected during your visits to the doctor. You’re spot on that we need to preserve folks’ privacy while leveraging the common benefit we get from using those data in aggregate. There are companies that specifically specialize in creating data sets that are anonymized and usable for mostly post-market monitoring of things, and that’s very important because that information can lead to the improvement of a diagnostic or some sort of treatment that’s already out there. The approaches that marry both the data collection in the EHR and the outcomes that you see in those patients, but then also get rid of that potentially identifying information are super valuable. They’re never able to get rid of all identifying information, and so anybody using those data sets does need to be well versed in data privacy and adhere to I’m using these data for this purpose and that’s it and then it goes away. Anytime you’re using those data sets, folks need to be trained, there need to be internal controls for any company that’s going to use those. On the discovery side, it’s a little bit more difficult because data are not collected to help with that. They’re collected primarily for reimbursement and for quality control for the clinic or hospital. If your outcome is quality of life, you’re going to be hard pressed to get data from real world data sets on people’s quality of life that are already existing. You’re going to have to go in and hand somebody a questionnaire on their phone, of course, because you don’t use paper anymore, about their quality of life. You need to understand the limitations of these real world data sets. You’re not going to be able to look at, oh, this person is on their way to developing cancer because they went to the doctor three times. That’s not a great outcome. You need something more definitive than that and making those kind of assumptions is just going to get you into trouble and waste time. It’s so important that leadership not be asking folks to be using data sets for things that the data sets can’t provide. Just because this data set exists does not mean that it’s going to answer every single question that the company may want, and so I think there’s this two-way education that needs to happen to leadership so they can set realistic expectations and goals and go get the resources and data they need to build whatever they’re trying to build.
Ross Katz: That makes a lot of sense, and I think across disciplines there are a lot of requests for analyses that presuppose the answer, presuppose the availability of the answer in the data. Trying to avoid that situation is always a good idea. As we bring the conversation to a close, I’m just interested in what emerging technologies or trends you foresee having the most significant impact on your world.
Michelle Wiest: I think it’s the real world data, but where the innovation is and needs to happen is at the data collection point. We’ve advanced so much in terms of machine learning and the speed at which we can do discovery from data sets. What is lagging behind is the meaningful data collection and targeted data collection and an understanding of the alignment of different ways of collecting. You want to triangulate those. I think that is really what has the potential to change the field and accelerate discovery.
Ross Katz: Michelle, thank you so much for joining us today. It’s been a really enlightening conversation and look forward to continuing it at a later date.
Michelle Wiest: Wonderful. Thank you, it’s been my pleasure.
Jason: And that’s it for this episode of Data in Biotech. If you enjoyed the episode, please subscribe, rate, or leave a review in your podcast player of choice. If you’re a biotech company struggling to unlock a data challenge, CorrDyn can help. Whether you need to supplement existing technology teams with specialist expertise or launch a data program that lays the groundwork for future internal hires, you can partner with CorrDyn to unlock the potential of your business data today. Simply visit www.corrdyn.com to learn more. See you next time.





