Listen on
Overview
97% of healthcare data remains unused, a staggering figure for an industry where timely information can directly impact patient outcomes and research breakthroughs. This vast, dormant dataset exists primarily because of severe fragmentation and the persistent challenge of securely linking sensitive patient information without compromising privacy. For data leaders and executives, this represents not just a compliance headache but a significant barrier to accelerating drug development, improving care pathways, and realizing competitive advantages.
In this episode, host Ross Katz talks with Vera Mucaj, Chief Scientific Officer at Datavant – a company focused on secure healthcare data logistics – who offers her perspective. A molecular biologist turned technologist, Vera uniquely understands the chasm between scientific discovery and real-world health implementation. She explains how Datavant’s technology bridges this gap.
The conversation explores how Datavant tackles data fragmentation and availability using privacy-preserving record linkage (tokenization). Vera details how connecting clinical trial data with broader real-world data enables more reliable, longitudinal studies, thereby reducing patient burden and accelerating regulatory approvals. She also addresses the evolving role of AI in de-identifying unstructured data, while underscoring the critical, ongoing research needed to ensure privacy in AI-driven data analysis. Finally, Vera points to semantic interoperability as the next crucial frontier in maximizing the value of connected healthcare data.
Key Takeaways
Healthcare’s data problem is primarily one of connectivity and trust, not scarcity.
An estimated 97% of healthcare data goes unutilized. This isn’t due to a lack of data generation but rather its fragmentation across disparate systems and the significant barriers around secure access and re-identification risk. Organizations hoard data out of legitimate privacy concerns, limiting its potential to inform research and improve patient care.
Privacy-preserving record linkage enables secure, longitudinal research with real-world data.
Datavant’s ‘tokenization’ technology cryptographically links individual patient records across different datasets (medical claims, EHRs, lab results, wearables) in a de-identified manner. This allows researchers to aggregate complete, longitudinal views of patient journeys—crucial for understanding long-term treatment efficacy and safety—without revealing personal identities, adhering to strict regulations like HIPAA.
Connecting clinical trial data with real-world insights accelerates drug development and benefits patients.
By linking data from clinical trial participants to their broader real-world medical history, researchers can conduct longer-term follow-up studies more cost-effectively and with less patient burden. This supports accelerated regulatory approvals, provides earlier safety signals, and helps secure full market authorization for novel interventions, particularly for complex therapies like cell and gene therapies requiring decades of monitoring.
AI is making unstructured healthcare data available, but introduces new privacy challenges.
Advanced AI models are now capable of de-identifying and structuring complex unstructured data, such as doctor’s notes and medical imaging, making vast new data sources available for analysis. However, training large language models on de-identified healthcare data and interpreting their outputs creates novel re-identification risks that the industry, including privacy experts and regulators, is actively working to address.
Related: CorrDyn specializes in building reliable, compliant data systems for the biotech and life sciences sector, including healthcare technology. Our data engineering services focus on secure data integration and processing, while our data quality expertise ensures information is fit for critical analytical and regulatory purposes.
Full Transcript
Jason: Welcome to Data in Biotech, a podcast from CorrDyn where we explore how companies leverage data to drive innovation in life sciences. Every two weeks we sit down with an expert from the world of biotechnology to understand how they are using data science to solve technical challenges, streamline operations, and further innovation in their business. This week we’re excited to be joined by Vera Mucaj, Chief Scientific Officer at Datavant, a data logistics company for healthcare whose products and solutions enable organizations to move and connect data securely. During the interview, Vera shares how Datavant helps companies overcome issues of data fragmentation and availability, which ultimately leads to better research and decision-making in healthcare. The importance of connecting clinical trial data with real-world data to improve research outcomes, and the challenges of semantic interoperability in data sets. Here we go.
Ross Katz: Vera Mucaj, welcome to the Data in Biotech podcast.
Vera Mucaj: Glad to be here.
Ross Katz: Awesome. Well, just to get us kicked off, could you give us a brief overview of your background and career to date?
Vera Mucaj: Sure, happy to. I’m a scientist turned technologist. I’ve trained in molecular biology as a cancer researcher and currently and for the past six years have been Chief Scientific Officer at a company called Datavant, which is a technology company that helps with data connectivity and exchange.
Ross Katz: Awesome. What—why do you do what you do? What motivated you to get into this field?
Vera Mucaj: Sure. I think it starts from a love of science and seeing scientific innovations being translated to the betterment of human health. Starting as a bench scientist, I could see a lot of technology be used to accumulate a lot of data, from single cell to the omics, and to really uncover the beauty of science. But there was a huge disconnect between all of the things that we were learning in the lab and how population health was being implemented. How the interactions between a patient and a doctor, how interactions between pharma and insurance and providers were happening in the real world. Moving from the basic science to technology gave me an opportunity to translate some of the things that we were learning as scientists to things that can improve public health.
Ross Katz: Yeah, you’re getting—you’ve gone from the narrow microscopic view of biotech up to the macroscopic and ecosystem view. With that as the jumping-off point, can you tell us about your role? What does the Chief Scientific Officer at Datavant do?
Vera Mucaj: As a technology company, we serve as a platform that helps protect, connect, and deliver healthcare data. You can think of us almost as a data logistics company. There’s a ton of healthcare data that’s being generated on a day-to-day basis, but only about 3% of it is estimated that it is useful and usable today. What happens to the rest of that data? It’s not there where it’s needed, and it’s not being used by who it’s needed. It could be anyone from physicians to scientists to policy makers to patients themselves. At Datavant, we provide both a platform and technology to do two things. One is to move medical record data through request of information and retrieval of information in the United States. The second one is to provide privacy-preserving record linkage technology, which is software that allows you to share data in a de-identified way, but in a way that’s still linkable and usable for longitudinal and large-scale research. As Chief Scientific Officer, I oversee product development for privacy-preserving record linkage and all of the technology that we needs to bring to play to ensure that you can exchange data, but you can do so in a safe and secure way. Beyond that, I’m also engaged in research as part of the company and with our ecosystem partners. I work very closely with public entities and government agencies in the US and outside of the US as well.
Ross Katz: You’ve got this privacy-preserving role where you’re helping organizations to acquire and share data. So there’s producers and consumers of data. Can you give us a sense of what are the challenges that organizations have been facing prior to the existence of Datavant that causes Datavant to be so useful?
Vera Mucaj: I would say the primary challenge has been one of data fragmentation, and the secondary one has been one of data availability. When we think about fragmentation, let’s take, for example, a R&D researcher at a pharma company. They might be involved in running a clinical trial and in analyzing the results of that trial. But if you think about it, that trial is only 4% of an individual’s life. Let’s say in a clinical trial for 18 months, that researcher wouldn’t know what happens to those participants in the clinical trial before they got to the trial. How did they get to be in a clinical trial to begin with? What those individuals are doing as part of their daily lives outside of the clinical study. And then more importantly, once the clinical trial is over, they might not have a way to know what happens to those patients longitudinally. Is the intervention that was tried in the clinical trial still useful? Is it effective? Is it safe? If you think about the data sources and the data users, we’re seeing a market dynamic that’s not being satisfied. Before Datavant was in place and providing this technology, you had data sources that were actively or passively gathering certain information that could be useful to research. And you had data users that had no ability to get that data. Through privacy-preserving record linkage, you can aggregate a number of these different data sources that we just mentioned: electronic health record, medical claims, lab, wearable, etc. You can aggregate them so longitudinally. In the US, this is allowed and allowable through HIPAA, if and only if that data is de-identified. How can you ensure that you are aggregating all of these data around an individual or around a population while still preserving their privacy? The technology that we offer does exactly just that. You might hear about it through the term that I used: privacy-preserving record linkage, or another shorthand for it is tokenization. It is a way to cryptographically be able to have an identity of a patient without necessarily knowing who that patient is.
Ross Katz: Interesting. When you talk about privacy-preserving medical linkage and the tokenization, you also mentioned cryptography. So there’s this way in which you’re doing two things: you’re understanding who the patient is from data set to data set and linking them together, but then you’re also using cryptographic methods to mask who that patient is so that it’s not discoverable by anybody who might want to connect the dots and re-identify a person that’s been de-identified. Am I thinking about that right?
Vera Mucaj: That is exactly right. The benefits of this are that now you can do large-scale, long-term research, but without really putting an individual’s privacy at risk. If we think about healthcare data, they’re probably the most sensitive and the most precious data for any single individual. We all want to make sure that we get the best care possible and we contribute to research. Every patient that I’ve talked to, every researcher that I’ve talked to says the same thing. But at the same time, everybody wants to ensure that this is done in a way that doesn’t impinge their rights to privacy. When we think about de-identification and re-identification risk, obviously tokenization is a big part of it to ensure you can have that identity resolution without necessarily knowing who the person is. But there’s also a number of layers of technology that go on top of it. Because as you can imagine, the more information you add to a data set, the more ways to re-identify an individual or group of individuals there could be.
Ross Katz: Interesting. I’ve heard data used—a metaphor for data as oil, as being particularly valuable. But I’ve also heard the metaphor of nuclear waste, that it can potentially be highly risky to store it and use it. By using Datavant to de-identify it and connect it together for you, there’s this opportunity to get the data you need to get the value, but also not have to take on the risk of making sure that it’s all where it needs to be.
Vera Mucaj: Yeah, absolutely. We were talking earlier about that statistic that I mentioned, which I believe is from the WHO, that 97% of data is not being utilized today, 3% is. One of the reasons for that is this risk components to it. Should we share data? In what cases should we share data? What kind of data user agreements should we have in place in order to do so safely?
Ross Katz: I’m thinking of your role as being this kind of two-sided marketplace, or multi-sided marketplace where you’ve got the producers of data on one side who have this data that’s potentially valuable that they want to share. And then there’s also the consumers of data that have value that they want to get about these patient populations, about these particularly disease categories, and they want to connect the dots between the two. Is that an apt metaphor for understanding of your model and the role that you play?
Vera Mucaj: Yeah, very much so. We consider ourselves a two-sided network. In fact, even a number of companies and clients of ours who work and use our technology and build on top of it, they’re also two-sided networks. When you’re thinking about data sources, the term that we often use in the industry is this concept of real-world data. Data about an individual that is anywhere where they could potentially touch healthcare outside of a randomized controlled clinical trial. This includes anything from medical claims data, which are generated through interactions between your healthcare providers and the insurance companies that reimburse for that care; electronic health record data that is usually generated through a patient’s interactions with healthcare; laboratory data from blood work and lab tests. But there’s also up-and-coming types of data sets like genomics data, metabolomics, proteomics, etc. There is data that is often thought of outside of the four walls of healthcare, so think, for example, socioeconomic data, financial data, consumer data, that would also be considered data sources. And your medical record itself, which I alluded to as electronic health record data, that is also a big source of data that is often exchanged for a number of use cases. You can think of data sources as companies or even individual patients that generates data. That data could be collected for a primary use case—care, reimbursements, research—but it could also be usable for secondary use cases: analytics, more research, policymaking, population health, etc. Lots of sources on the left-hand side and we’re very proud at Datavant to have more than 700 sources that use this universal technology to be able to exchange data with each other. On the right-hand side, you have the data users and the use cases for which this data could be exchanged. Some of these players could be the same players. Pharma can generate data, as I mentioned, clinical trials. Pharma also uses data. You could use it for commercial analytics, for health economics outcomes research, for improving your next clinical trial. Pharma in life sciences is a big user of healthcare data. Payers and providers themselves are also huge users of data. Others could be the patients themselves who can get back their data and information for their own decision-making with their physicians. It could be research institutions who are working as part of either individual researchers or as clinical research networks. It could be governments themselves. In the United States, which is where most of my expertise is, we’re seeing a lot of registries and a lot of research databases that are, let’s say, National Institute of Health funded or CDC funded that use healthcare data to allow a number of researchers to pursue important scientific questions.
Ross Katz: There have to be lots of organizations out there that have these underutilized data assets that they’re protecting because they’re really concerned about security, or their worst nightmare is that this data gets out there and it comes out that their organization was the source of that data that got out there and got re-identified. I’m just curious how you see potential producers thinking through that problem space.
Vera Mucaj: First, they think about whether they have the data use rights to exchange that data in a de-identified way. If you have that right and if you’re doing it in a way that’s compliance with the law and compliance with your contractual obligations, the concern of am I sharing that data shouldn’t be the case because you have the right to share that data. De-identification and other protections on top of it make it easier to exchange that data. A number of data sources have been doing so for years now. Obviously, there are business considerations on whether we should exchange this data or not. There are strategic considerations as well. One of the things that we’re seeing more and more with the years is, there’s some data sources that have always exchanged data, but there’s others that have never really thought about it. In part because some folks are not aware of the need for that type of data. We’re seeing this anywhere from research to individual patient registries or nonprofit organizations. What we’ve seen happen is, throughout the process, before anybody even agrees or decides to exchange data, we offer technology to do essentially a little bit of a matchmaking, so overlaps. Let’s say I am a patient organization that has data on a thousand patients and I have consent to share that data in a de-identified way from the patients in my organization. How do I know that data on these thousand patients is useful to a researcher or a pharma company? How do I know that there’s other data that can enrich the data set that I have if I wanted to bring data in rather than share data out? If you can imagine, I tokenize my thousand patients, this other data source tokenizes millions and millions of claims records. You can run an overlap. You’re not sharing any information. You’re not even sharing the identity of your patient, and you shouldn’t. But you can see that through that overlap, let’s say you have a 98% match. When you have a 98% match, you know that there’s information about your participants in both parties, in both sides, and you can decide at that point whether it makes sense to exchange data. A bit of what we’re doing right now is to essentially go and share this message with all potential types of data sources and say: A, an education component, there are reasons for why your data is important and valuable. B, here’s how that data can contribute to positive outcomes that are in line with your mission. And C, here’s some opportunities for you to exchange data, and to make the right decisions on this data exchange without actually going into a contract, as a first step.
Ross Katz: Interesting. As you’re going out and you’re acquiring all of this diverse data that can be linked at the patient level, I’m imagining that the opportunities for the value that can be derived from the data grow exponentially from there. I’m interested, what are the biggest use cases that you’re seeing from the consumers of data and the value that they’re able to drive when they’re getting data through Datavant?
Vera Mucaj: If we’re taking medical claims data as one of the data sets that’s the most exchanged when you’re talking about de-identified data, that information has a lot of cost component to it, obviously, just by virtue of what medical claims are. We’re seeing a lot of health economics and outcomes research, or HEOR work, in some cases, folks will just call it real-world evidence generation, that’s done with that data. That data becomes super valuable when it is combined with electronic health record data, because you could have some high-level information in claims, as well as cost information, and with electronic health record, you have the deeper medical history of the patient. When you put it together, you can tell a more complete story. This type of use case can help make the case for why we should continue to pursue a potential asset, why we should reimburse a particular intervention, and why certain policies should be in place so that more patients can get access to that particular intervention or drug, for example. HEOR is a main use case. The second one that I alluded to earlier that we’re seeing a lot more over the past few years is this use case that we internally call trial tokenization, but essentially connecting your clinical trial data to real-world data about those participants so that you can do more longitudinal studies. If you look at some of the trends of clinical trial approvals and you’re seeing a lot more accelerated approvals that require longer-term studies, in many cases, those longer-term studies involve starting a new trial, a larger trial, potentially a more expensive one. But what would it look like if you’re taking your Phase 3 participants and you’re continuing to passively learn more information about their outcomes even after the trial is completed? This is a cost-effective way to bring some of that research forward. If, god forbid, there’s any safety issues with a particular intervention, you’re learning earlier. If there’s continued efficacy, you’re getting from accelerated approval to full approval hopefully. That is a use case that’s been very interesting as we’re seeing more and more interventions come to market. When it comes to some more novel interventions, so think cell and gene therapies, which are being tried on patients for probably the first time ever through these clinical trials, we’re hearing both the FDA and the EMA saying that we’d need 10, 15 years of follow-up studies for these patients. Being able to do so in a way that is both active and passive is important. A lot of cell and gene therapies are done on pediatric populations, for example. Assuming they work wonderfully, those patients are going on to live their lives, they’re going to college, they’re moving somewhere else to be in a job. In 15, 20 years, do you really want them to go back to the academic medical center every six months to provide information when really they are providing this information wherever they’re being seen as patients? How can we use tokenization technology to make these longitudinal studies valuable for regulators, optimal for the sponsors who are pursuing them, be they pharma or academia, and also patient-centric? How do we reduce the burden of the clinical trial participants?
Jason: Are you a biotechnology company looking to unlock the potential of your business data? CorrDyn can help. We’re an enterprise data specialist that helps companies working in life sciences make smarter, strategic decisions. From developing the right data strategy that starts with our data maturity assessment to building and delivering bespoke technical solutions, we are equipped to tackle the most complex data challenges. We have partnered with dozens of high-growth organizations, from manufacturers of custom oligonucleotides to molecular diagnostic companies to achieve data competency. Whether you need to supplement existing technology teams with specialist expertise or launch a data program that lays the groundwork for future internal hires, you can partner with CorrDyn to unlock the potential of your business data today. Simply visit connect.corrdyn.com/biotech to learn more. Now, back to the show.
Ross Katz: I’m curious, you mentioned the long tail of value that people could get from the data. Are there use cases in that long tail that you think are unseized opportunities from the data sets that you all have available or expect to have available soon?
Vera Mucaj: I think generally, building robust patient registries is one that I would love to see in the next five years happen both in an active way—so patients provide active information through patient-reported outcomes, through visits—but also their medical records and their medical history are automatically incorporated. Tokenization is a component to this, but I think just medical record exchange, so the logistic of moving data to where it needs to be, I think will be super important. If you think about any industry that we touch in our lives right now, we have our smartphones, we have the ability to order something online and for it to arrive the next day at your house. You can’t do that today with your medical record. There’s no one magical button to bring all of your medical history in one place where it’s needed. It’s Datavant’s mission to build that hypothetical button. But today, that process is very, very difficult. Every time a patient gives, at least in the US, their HIPAA authorization for their medical records to be included in research, there are a lot of barriers to that today. Both some privacy and security issues around unauthorized disclosure, but also very much logistical ones. Not every piece of data can be gathered electronically even if you connect to APIs directly to an electronic health record. Some information might be found in just another part or another system in the hospital. For oncology data sets, we see imaging is found elsewhere from your medical history, your pathology is found somewhere else. Bringing all of this together to build registries is something that I consider long tail right now because it’s not being done at scale as I would like to see, but also where I see some of the largest opportunity in the next five years.
Ross Katz: Just to help me understand the business model from the producer’s perspective and from the consumer’s perspective, is it that the producers are presenting their data set through the Datavant platform and having it de-identified and then selling the data set to the consumers and then receiving revenue from the data set that they sell through your marketplace? Is that how it works? Or if you could give a little bit more information about how that connection works.
Vera Mucaj: By and large, one of the things that we decided to do strategically from the very, very beginning of Datavant is to let the data sources—data producers as you’re mentioning them—have full control for how they should exchange their data. We actually don’t gather all of that data in one marketplace. Each source owns their own data, and they can choose whether they want to engage in an overlap, whether they want to be seen as a source on the Datavant platform, and how they want to have the relationship directly with a potential data licenser. Different groups do it differently. Some sources are very open to say, here’s my logo, I want to exchange data, here you can do an overlap on two clicks. Some others could say, let’s have a conversation first about what agreements we have in place for exchanging data versus not. In some ways, there’s no two sources that choose to do this the same way. But we’ve put together the platform and the technology that anybody can use whatever components are relevant for them to be able to do this exchange. It could go from exploring for either sources or buyers, to doing some of these data assessments like overlaps, to contracting, to distributing data, etc. Who use all of it? Who use us also as the privacy tools to make sure that once the data are together, the re-identification risk remains low, all the way to distributing that data or holding and distributing that data on their behalf. There’s a little bit for everyone depending on what your needs are, in a way.
Ross Katz: Interesting. It sounds like a really custom transaction model where you can choose the partners that you want to have access to the data, depending on the particular contractual agreements that you’re already in. What is Datavant’s business model in the context of that? Do you receive—is your revenue coming from the consumers or the sources/producers or—
Vera Mucaj: It’s coming through from both parts of the network. The benefit of us remaining neutral is how you get so many sources to want to be part of your ecosystem. For the graph theory enthusiasts in the audience, you can think of it almost as nodes and edges and links. The bigger that network becomes, obviously, the more valuable Datavant becomes. Our business model is, obviously, we provide the tools for the tokenization and the privacy preservation and there’s a cost to that. But once two parties decide to exchange data with each other, that’s how we’re growing. The more links, the more data exchanges there are in the ecosystem, the better Datavant does.
Ross Katz: You’re doing this de-identification work on behalf of the producers in a way that allows the consumers to link up the patient records. How do the producers know in advance what the quality of the data is or how well the data set will integrate with the existing data they have available?
Vera Mucaj: Yeah. Some of that is through technology that we provide, but a lot of it is coming as information from the data sources themselves. As we talked earlier, a few of these companies in the ecosystem have made a business out of aggregating, normalizing, and building very rich data assets and building analytic tools on top of it. As part of their work, as part of their go-to-market, they also provide quality measures on that data. In terms of how we help both parties make the decision for that data exchange better, on top of just the overlap components, so do I have patients in my data set that are also found in yours? We also provide other profiling information. Still in a privacy-preserving way, there’s a number of metadata that you can provide to make that decision better and faster. If there’s a particular data element that I’m super interested in, let’s say geography, I want to have a diverse geography in my data set, what coverage do we have at the state level, at the zip 2 or zip 3 level, for example? Ultimately, in the end, there’s a lot of specifications around how the data was collected, how the data was normalized, were there any imputations on the data that happened before the data sources and data users, quote-unquote, meet in the Datavant ecosystem. A lot of that work does happen outside of the Datavant platform, but we’re happy to have our technology help wherever that’s possible. One thing that we haven’t talked about is how is artificial intelligence helping with some of this process right now? Both to generate data but also to analyze and potentially create some imputation of data. We’re seeing some of that work be emerging right now, in terms of how do you make your data sets more high quality by using technology around it and by enriching with other data? But also, how can we use AI to bring more types of data into the realm of usefulness? This whole time we’ve talked about some data sets. They all fall into the category of what we call structured data. Data that you can find in a table, to be simplistic. There’s lots of information like doctor’s notes, genomic information, medical imaging that fall in the realm of unstructured data. It is very difficult to standardize, normalize. Frankly, it used to be very difficult to even de-identify that data. We use AI technology at Datavant to de-identify unstructured data and we’re moving into the imaging and genomics realm as well so that that data can be exchangeable now in a way that was not the case five years ago. When I started at Datavant, most folks would just drop off their doctor’s notes from a data set rather than exchange it because there was this fear that you might have somebody’s name or somebody’s address in the doctor’s notes and you weren’t able to de-identify it properly. AI has led to huge advances in doing that properly. It’s also led to huge advances in almost structuring that unstructured data. How can we abstract the right information out of a sea of words? And then obviously with generative AI tools now and with large language models, you don’t even need to do that abstraction. You could ask a question to unstructured data in layman terms and be able to get the answer back in simple sentences as well, all powered, trained, and fine-tuned by healthcare data. Super fascinating, super interesting. It leads to two questions for me, and I think this is where the industry and the ecosystem will have a lot of thinking to do. Number one is, how do we make sure that we have all of the right data and all of the representative data to train this model to begin with? There’s that concept—you hear the term “garbage in, garbage out.” If you don’t have representative data training your models, you’re not going to have representative answers and you’re not going to have answers that are applicable to all your patient populations. Number two is, going back to where we started, the privacy question. When you have single data sets and you’re adding them together, you know what information you’re adding, you know what you need to redact, what you need to roll up, in order to still maintain privacy. How do you do so when you’re having models that are training on the entirety of internet and then are fine-tuned on healthcare data? That data could walk in to the model de-identified, but there might be questions or prompting that could increase re-identification risk. A big question for our privacy experts and I think for the industry is: how do we ensure that if these new AI models are using healthcare data as training or fine-tuning elements of the model, that we still maintain privacy? How do we de-identify the data going in, and also how do we protect the information coming out so that it doesn’t risk re-identification? Datavant and others in the industry right now are thinking very hard and doing a lot of research and technical development to make sure that the new wave of analytic companies that are built on AI and are AI-native are also doing their work in a way that is representative and privacy-preserving.
Ross Katz: Interesting. It’s an open research question that requires a lot of experimentation to get to the bottom of, to what extent is training models on this data lead to adverse outcomes. I want to turn back to the clinical trial use case that you talked about earlier. It seems like using this long-term longitudinal data that you provide in the context of a clinical trial is a paradigm shift in the way that clinical trials would think about analyzing the results, monitoring the results, reporting, or thinking about the results. How do you see clinical development teams, clinical trial operations teams adjusting to the availability of this data or getting up the learning curve of putting this data to use?
Vera Mucaj: Yeah. It’s going to be education for everyone. It’s going to start at the clinical trial sites, the clinical research coordinators who have the engagement with the potential trial participant candidates. Education for the places where the clinical trials are being run, and Datavant does quite a bit of this. Education and communication with the patient themselves, with the trial participants. If five years ago you went into a clinical trial and you signed some standard informed consent form that said “my data might be used for secondary research,” now we’re actually being very proactive and saying: Hey, if you wish—and only if you wish, trial participant—to also have other passive information about you be incorporated in the future in a de-identified way, here’s what that means. The word tokenization is not a common word, so we try not to use it. There’s patient videos and there’s other things that we’re putting in place to explain what will happen, for how long, and also making it very clear that this is voluntary and participants may revoke consent if and when they wish. There’s a lot that’s happening at where the clinical trial happens to educate those groups. Then there’s a lot that needs to happen to educate the researchers, so the folks at a pharma sponsor or at an academic medical center who will be linking clinical trial data now to real-world data. If you’d asked me five years ago when I was first talking to some of these folks, they’d say, “Do not touch my clinical trial, nothing must happen around it. It is very important that we don’t accidentally unblind that information, that we don’t accidentally incorporate information that’s not correct or that we don’t know how to interpret.” There’s been a lot of work that’s gone into bringing all of that ecosystem up through the growth curve. How the process of privacy-preserving record linkage maintains a blinded study, what kind of data you could link—so knowing that there’s medical claims and EHR data out there that you could link—and then knowing what kind of questions you can ask of that data. Oncology, as I mentioned earlier, you can finish a trial with a progression-free survival endpoint, but five years later, you might need an overall survival endpoint. What is the mortality endpoint? Knowing “how can I add mortality, death data to my clinical trial data to really build those survival curves in a way that I haven’t been able to in the past?” When you explain what these data points could be and how you can use them in your study, there’s almost a penny drop, a click moment that says, “Hey, this is something that I wasn’t able to do before and I am now.” We’re actually seeing a lot of enthusiasm out of the folks who actually will benefit from this linkage. Finally, there’s the regulators. There’s been a tailwind in this space because in the US and beyond, there’s been a lot of funding that’s gone into having our regulators think about how real-world data is used for regulatory decision-making, 21st Century Cures Act in the US as an example. We’ve had lots of guidances that have come from FDA, EMA, and elsewhere that says, “Here’s how we think about real-world data, here are some of the pros and cons of working with it, and here’s where the bar is.” As I mentioned earlier, there could be lots of ways to collect and connect real-world data. That type of diversity is not the kind of thing that regulators are used to. Clinical trial data, everything is collected in a very standardized way, and you have the patient there if you need to get more information. But medical claims only have so much information. How do you make sure that you’re not inferring something that is absolutely wrong, for example? There’s lots of guidance that has come from regulators and lots of two-way conversation that’s going between researchers and regulators and pharma and regulators to make sure that the data that’s being used is fit for the purpose that it is being used in. The advice that I usually give folks who want to connect clinical trial data to real-world data is to actually go and have the conversation with your regulators if you’re doing this for a regulated use case. Have the conversation first. Say, “This is the statistical question that we have, this is the research question that we have, this is the type of data that we’re thinking about linking together. Let’s get some feedback. Will this be sufficient to get to a conversation, or is there something else that we should do at the beginning?” Rather than going with all the work done at the end and then having the other party tell you that there’s a lot more information that would be needed in order to make a decision.
Ross Katz: There was a quote from Datavant’s Chief Product Officer, Shannon West, at your conference a couple months ago that really jumped out to me that I felt like was really indicative of where things are, and so I would just love to get your thoughts on this quote. She said, “We can figure out how to technically exchange the data, but that doesn’t mean we have semantic interoperability between the data sets.” That sounded to me like it was the next frontier of where the ecosystem goes. Can you talk a little bit about what she might have meant by semantic interoperability and where you’d like to see the ecosystem go?
Vera Mucaj: Absolutely. First of all, shoutout to Shannon and her team. They’re the folks who are going to put the fax machine out of business. It’s a joke within Datavant, being able to do that click of a button medical record exchange, and that’s the primary goal of Datavant today. Where she’s going with the semantic exchanges: Even if I have that magical button and all of my medical history and your medical history goes into a registry, what do you do with it next? How do you standardize it? How do you normalize it? What is the right common data model to exchange it across different researchers? There’s a lot of work that happens slowly and then very, very quickly in order to get to truly interoperable data. The logistic component of delivering information from one place to another is difficult, but I think how do you put that information together and analyze it in a way that is proper is a research in and of itself. Shannon mentioned this at the Datavant Summit because we were gathering there folks who sit and think about the exchange of data, and when you have payers, providers, pharma, research, government, all in one place, she wanted to throw the challenge out there. What happens if we actually deliver all the data? How can you make a really, really good use of it on the other side of the delivery?
Ross Katz: Another thing that I noticed in just looking at some of the services that are available in Privacy Hub was this idea of expert determination. If you could just explain what expert determination is and how it’s classically done and Datavant’s lens on expert determination, I think people would find that interesting.
Vera Mucaj: Absolutely. I’ve alluded a little bit to the fact that tokenization alone is not what makes a linkage privacy-preserving. If I’m de-identifying my data set and you’re de-identifying yours and we put it together, but I removed some information here that your data set is now adding back, you’re increasing the risk. When you’re looking at HIPAA saying, “Look, you can share de-identified data,” they’re saying you need to remove a lot of information for it to remain de-identified. HIPAA puts two ways to do this. One is called the Safe Harbor method, which removes a ton of information, and some of that information is really useful for research. The second method that they recommend is something called the expert determination method. Having experts, statisticians who are trained in healthcare data, take a look at that data, analyze individual data sets and combined data sets for re-identification risk. HIPAA doesn’t tell you what’s a high versus a low re-identification risk, it just says make sure it’s very low. These experts are individuals who’ve worked in this space for years and years, who’ve done the risk assessment process for looking at one or two data sets together to see if there’s other things that you need to do to the data before it’s actually exchanged. Datavant has brought a number of these experts together under the banner of Privacy Hub, where we continue to do research on privacy preservation, but we’re also building technology around it. How can you take some of these work that was done as a service, as a process, as a long-term analytics, and make the data available and linkable much faster? How can you use some of the technology so that you don’t have an individual or two, you don’t have a number of data scientists looking at a data set, you actually have continuous monitoring for risk, and you have those expert determinations still led by experts but being done in a way that is faster and more efficient?
Ross Katz: Yeah, so scalable expert determination that’s of equal or higher quality is the vision of where that’s going. Well, great. As we come to the end of our conversation, I want to give you an opportunity if there’s anything we haven’t touched on that you want to make sure that you share about the arena that Datavant works in. There’s just so many parts of the real-world data ecosystem that you all touch. I’m certain there’s something we might have missed.
Vera Mucaj: Yeah. I think this is more of a call for researchers. If there’s something that you’re looking at right now and saying, “I wish I could do X, but I don’t have Y,” is it possible that this work could be done if the data was there and if the data was there in a timely manner? If that’s the thing that’s your barrier, talk to us, talk to other data sources, look for that information because, frankly, one of the things that I’ve learned is the data are there, maybe they’re not there at the right time or at the right place for the right person to work with them. We want to make sure that that exchange is happening a lot more seamlessly going forward. Keep asking the questions, don’t stop where you think you might not be able to do something, because it is possible that today or in the very near future that data and information might be available to you in a way that wasn’t before.
Ross Katz: Very interesting. Just for anyone who’s following your field, are there any things you’re reading or listening to or resources that are out there that you would recommend to people to better understand this arena or keep tabs on it?
Vera Mucaj: In particular, if you’re interested in how to use real-world data for regulatory decision-making, FDA and the EMA have what’s called guidance documents that are publicly available. You can get quite a bit of education that way. Datavant of course has a number of publications and blog posts where we essentially educate folks on a real-world data ecosystem, including one blog post called “The Fragmentation of Health Data” that I always use as a primer for everybody who’s new to this industry. In general, just follow the same news sources or publications that you would to follow up on clinical and scientific research and think about them from the context of healthcare data. Every piece of research that I look at on a day-to-day basis has a healthcare data component to it. Once you start thinking about it, you’ll see it everywhere.
Ross Katz: Well, Vera, thank you so much for joining us today. It was great to have you on and look forward to connecting down the line.
Vera Mucaj: My pleasure. Have a good one.
Jason: And that’s it for this episode of Data in Biotech. If you enjoyed the episode, please subscribe, rate, or leave a review in your podcast platform of choice. See you next time.





