Skip to content
Jonathan Eads — Unlocking Genomics Intelligence with Jonathan Eads
Data in BiotechEpisode 26

Unlocking Genomics Intelligence with Jonathan Eads

Jonathan Eads of Genomenon discusses balancing software-driven processes and human curation to make genomic evidence actionable.

40:59Full transcript below
JE

Jonathan Eads

VP of Engineering at Genomenon

Overview

The exponential growth of genomic data presents a critical bottleneck: millions of genetic variants remain “of unknown significance” because extracting and standardizing information from scientific literature overwhelms traditional methods. Data leaders and executives face the challenge of transforming this deluge of unstructured data into actionable insights without compromising accuracy or incurring unsustainable costs, directly impacting drug development timelines and patient care.

Host Ross Katz speaks with Jonathan Eads, VP of Engineering at Genomenon, who offers a clear perspective on addressing this problem. His extensive background in biotech data systems informs Genomenon’s mission to make genomic evidence actionable. He explains how their platform, Mastermind, processes over 9.5 million full-text publications to index genetic variants and their associations.

This conversation explores the technical and operational complexities of normalizing genetic nomenclature, the essential role of human curation in validating critical insights, and the strategic integration of AI to enhance efficiency without sacrificing accuracy. Eads highlights how this hybrid approach directly supports clinical diagnoses, accelerates pharmaceutical research, and informs regulatory decisions.

Key Takeaways

Inconsistent Genomic Nomenclature Creates a Critical Data Bottleneck.

The exponential growth of genomic research yields up to a million new papers annually, yet millions of genetic variants remain unclassified. This is largely due to widespread inconsistencies in how variants are described across literature, often ignoring or extending standard nomenclature. This challenge prevents automated indexing from effectively making these critical insights actionable for patient care and drug development.

Human Curation is Irreplaceable for High-Stakes Variant Classification.

While software can broadly scan and extract information from scientific literature, determining if a genetic variant is benign or pathogenic demands human expertise. These classifications follow stringent American College of Medical Genetics guidelines and carry significant ethical weight. AI’s primary role is to enhance curator efficiency and precision by surfacing relevant context, not to make definitive diagnostic calls.

Maximize Data Recall to Prevent Critical Misses in Scientific Evidence.

When building a thorough genomic knowledge base, prioritizing broad data capture (high recall) is crucial, even if it initially means tolerating more false positives. Missing a rare, clinically relevant variant (a false negative) can have severe consequences for patients and research. Subsequent processes can then refine precision through granular classification of articles, such as distinguishing clinical relevance or human vs. mouse studies.

Actionable Genomic Data Delivers Tangible Business and Patient Impact.

Beyond basic variant searches, organized genomic evidence directly supports strategic business outcomes. Pharma companies use it to determine disease prevalence for regulatory submissions and expand clinical trial eligibility. For patients, it can lead to diagnoses, justify insurance coverage, and connect families with expert researchers, underscoring the real-world value of transforming unstructured data into clear insights.

Related: CorrDyn helps clients build reliable data engineering pipelines and define effective AI strategy that augments, rather than replaces, human expertise. In the biotech and life sciences industry, ensuring exceptional data quality is paramount for clinical and research outcomes.

Full Transcript

Jason: Hi everyone, this is Jason, producer of Data in Biotech. Before we get started, I wanted to let you know about our latest white paper. It’s a comprehensive guide to implementing machine learning models in biotech manufacturing. It’s a complete overview of all the potential problems of ML adoption and, more importantly, how to solve them. To download it, simply visit connect.corrdyn.com/biotech-ml. We’ve also dropped the link in the show notes of this episode. Okay, let’s get into it. Welcome to Data in Biotech, a podcast from CorrDyn where we explore how companies leverage data to drive innovation in life sciences. Every two weeks, we sit down with an expert from the world of biotechnology to understand how they’re using data science to solve technical challenges, streamline operations, and further innovation in their business. In this episode, we talk with Jonathan Eads, a bioinformatics expert at Genomenon. Jonathan shares how the company makes genomic evidence actionable for patient diagnosis and precision medicine development. He explains how their software, Mastermind, scans genomics literature and indexes genetic variants to provide critical insights for stakeholders, including pharma companies, biotech firms, clinicians, and patients. Jonathan highlights the role of human curators in variant classification, the challenges in normalizing genetic data, and the development of their NLP pipeline. Finally, he discusses their use of automated quality control processes and the AI methods that they’re using to enhance data accuracy and efficiency. Here we go.

Ross Katz: Jonathan Eads, welcome to the Data in Biotech podcast.

Jonathan Eads: Happy to be here.

Ross Katz: Awesome. Well, just to kick us off, would you mind giving us a brief introduction to who you are and your background and what brought you here today?

Jonathan Eads: Yeah, sure. I’m Jonathan Eads. I’ve worked in a number of different biotech companies. I was the senior director of software engineering at Synthetic Genomics, a Craig Venter company. I was the VP of software engineering at a startup called Denovium. We were acquired by a biotech company called Absci. And there I was the senior VP of informatics and led the bioinformatics software engineering, data science, and IT efforts there. That brings us ultimately to the company I’m in now, which is Genomenon, which is a whole new deal for me. Most of my career has been working with NGS data or some kind of experimental data, surface plasmon resonance data, transcriptomics data, proteomics data. This is working with data coming out of the literature, which is an interesting switch for me in my career.

Ross Katz: Yeah, that’s fantastic. So with that as a background, can you just introduce Genomenon as a company and then what your role is within it?

Jonathan Eads: Sure. Genomenon’s mission is to make genomic evidence actionable. We are all about providing actionable insights for patient diagnosis and precision medicine development. My role at Genomenon, I’m the VP of engineering. I lead a couple of smaller teams that would include the clinical engineering team, curation engineering team, and platform teams, as well as the data science and AI engineering efforts.

Ross Katz: Awesome. So making genomic evidence actionable. Can you just dig a little bit deeper into what that means in practice for the different stakeholder groups that you support?

Jonathan Eads: Yeah. Making it actionable can mean, say, accelerating an NGS variant interpretation. You’ve got a patient that’s got a particular genetic variant. Is that variant pathogenic? Is it benign? Is there a disease associated with it? That’s one type of actionability. We do that, in part, by scanning the full corpus of genomics literature. We’ve scanned about 9.5 million full-text publications. We look for genetic variations and their associated genes, and we index them and provide a very sophisticated search interface, one that specifically can deal with different structured forms of variant nomenclature, which is really critical to get that part right. Another part of making it actionable is we’ll partner with biotech companies or pharma companies where they want to learn more about the genetics of a disease of interest. They have a couple genes they’re really interested in and we have a services part of what we do, where we do curation for hire in that case. And that can really accelerate a pharma or biotech company’s efforts in a couple different ways.

Ross Katz: Yeah. I think it would be helpful just for the rest of the conversation to understand, both on the drug development side and on the pharma side and then also on the clinician side and on the patient side, what are the different ways that the different jobs these different user categories are trying to accomplish using the genomics data that you’re surfacing through Genomenon?

Jonathan Eads: Absolutely. It’s a little complicated. A good way to answer that to start with would be to look at what we offer. We have three offerings: software, data, and services. On the software side, our commercial product is called Mastermind. That software application allows you to search across the indexed genetic data from 9.5 million full-text publications. On the data side, you’ve got a couple forms of high-value data. One would be curated variants. We have a team of highly skilled genetic curators that curate genetic variants. They also curate gene-to-disease relationships. Both of those piles of data are part of our offering. You can also access just the data indexed from the full-text publications that hasn’t necessarily been curated. And then the third offering is services, where our expert curators can work for a very specific purpose. A pharma company might have a gene they’re investigating. Maybe they want to look at identifying an evidence-based threshold for determining how many people out there are in need of help for a particular disease. That’s an area where we can provide a lot of value. We can provide value in regulatory decisions — what should this variant be included in this clinical trial or not, or should it be excluded from a regulatory board. Those sorts of value-adds are at the pharma and biotech level. At the patient and doctor level, it’s much more hands-on in that there’s a kid with a variant and some horrible symptoms. It’s likely some rare disease. A doctor is trying to figure out what’s going on. They can come to our interface, search by that variant, and they might find the one and only publication out there associated with that variant. That could lead to a diagnosis and, hopefully, a treatment. Just the diagnosis by itself is really high value in some cases. If you’re a parent in that situation, that diagnosis can lead to connecting with a university professor that’s an expert in that topic. And we can also build evidence for that threshold where there’s enough people that have this — let’s put some resources into finding a treatment.

Ross Katz: Interesting. Just to summarize what I’m hearing to make sure that I understand. On the pharma side, if you’re in the process of drug development, there’s an understanding that genomics underpin a lot of disease. There are genetic components of a lot of the different diseases that are out there, either known or unknown. As you’re exploring the way that disease comes to be and the disease pathway and different options for intervening along that pathway to alter the disease, you want to understand the variants that are associated with it, and then also any research that’s been done to determine how those variants influence what’s happening inside of a person over time. Am I thinking about that right on the pharma side?

Jonathan Eads: Yeah. That’s exactly it.

Ross Katz: Okay, fantastic. And then on the clinician and patient side, it’s more along the lines of, we have a problem and we know what the phenotype looks like, we know what the presenting symptoms are, but understanding the genetic variants that are present in the person can help to diagnose, help to understand what the disease is. Has the disease been researched before? Are there indications inside of this person that might help us to better understand where those symptoms are coming from and what the real source of the disease is, because that helps to understand the entire life cycle of it and think about treatment options for a particular individual. Am I thinking about that right?

Jonathan Eads: Yeah. That’s it. A backdrop to that is to keep in mind that if you go to ClinVar, the public database where most of these genetic variants are deposited, the vast majority of them are what are called variants of unknown significance. Nobody knows — are they benign? Are they pathogenic? They’re not classified. And that’s a result of the sequencing technology, the actual physical devices that are able to determine genetic sequence. That technology has exponentially grown. We’re pretty good at it at this point. We are generating lots and lots of biological sequence. It has created a problem: we’re finding all these variants, there’s millions of them, we don’t know what they do. The means of figuring out what they do is still very manual and labor-intensive. Like you described, you have a patient, they have a set of symptoms, you discover a variant. That variant might be a variant of unknown significance. What do you do? We’re trying to put a dent in that space.

Ross Katz: And so the solution — by curating all of this research, you’re able to get more leverage on the experiments that are being done, the research that’s being done to understand how each of these variants influences a person’s biology or a disease pathway, so that clinicians or a pharma company or any biotech company can better understand what research has been done and then generate hypotheses about what could be done or what additional research could be done that builds on the collective intelligence that’s been developed over time. Am I thinking about that right?

Jonathan Eads: Yeah, absolutely. Another thing you can get out of that is evidence-based information on disease prevalence — how many publications are out there about this? We also know how many times that variant has been searched. That can be some really important information.

Ross Katz: Yeah, awesome. How is this done historically? Was it literally just people — scientists or doctors were expected to read as much of the literature as they possibly could? Or were there other methods available to understanding how these different genetic variants might relate to given diseases?

Jonathan Eads: Before we had these custom search indexes, initially it was: you’re expected to read a lot of papers. In genomics, you’re talking 700,000, probably pushing a million in some cases, publications a year. So that’s a lot of reading. Not all of those are going to have genetic variants in them, but a subset will. Obviously when search engines came along, you can search across them. We have things like Google Scholar you can search. But if you’re a clinician and you have the variant you’re interested in, you can’t go to that search interface and get precise results utilizing a variant in formal nomenclature. That’s a really important step forward. Without that, you’re just going to miss stuff. That’s exactly what’s happened historically. This is your only resource to get at some of what’s being seen in the clinic. Even if it’s one publication and the variant is described utilizing natural language, it’s not going to come up in a search. The odds of that are really low.

Ross Katz: So it sounds like one of the challenges, even if, hypothetically, a human being could read the 700,000 or a million genomics papers that are coming out a year and could find the variants that they’re looking for, those variants aren’t always named the same way or described in the same way. Can you just elaborate on that a little bit and how that influences the kind of work you have to do at Genomenon?

Jonathan Eads: That gets into some of the crux of the fundamental problem. If you look across those 9.5 million articles, it’s a tsunami of poorly adhered-to nomenclature standards. Some of it predates the standard — go back far enough and there was no standard. The main standard for describing human variation comes out of the Human Genome Variation Society, the acronym HGVS, and in most modern publications you’re going to see a variant described in that nomenclature. But even within that nomenclature we see problems. The author might use a mix of natural language and partial nomenclature to describe it. In some cases, there might be some genetic feature or anomaly that isn’t handled well by the nomenclature and they’ll extend it in their own way. None of that is predictable — it’s not like you get a manual for detecting this stuff in advance. And this is just the situation for single-nucleotide variations. Structural variations — insertions, deletions, translocations, inversions — the syntax is even more complicated and messy and there are more historical anomalies. All of that is a challenge to get indexed. That’s one of the big reasons why there’s a lot of value in trying to convert that information into a structured, searchable format. You can’t get there without doing it.

Ross Katz: Can you walk us through the process — what it looks like on your end for how you get it into that structured searchable index?

Jonathan Eads: Sure. You’ve gotta come up with a way of internally mirroring the full corpus of publications. PubMed is a public resource that gets you the metadata associated with a publication: the title, the abstract, the authors, that sort of thing. But it doesn’t get you the full text. Once you get that, then you’ve gotta figure out how you’re going to go get the full text. And this is another area where things get really messy. You have supplemental materials to deal with. There’s no standardization for the format of supplemental materials. We’ll see rasterized images with tabular data in them — how are you going to get the variant out of that in a PDF? We’re lucky if we get a CSV file. But sometimes it’s a Word document. We’ve gotten videos before. It’s a messy problem. It’d be great if the publishers could get together and come up with some standards for how to represent supplemental material, but I don’t think that’s going to happen for a long time, if ever.

Ross Katz: That makes a lot of sense. On the software side, when you’re exposing this information to people, I’m assuming you’re linking people to the full-text version of the paper, not giving them the full text directly. You’re saying within this paper these variants are here and this is some metadata about how this variant relates to a particular disease pathway or protein function. And then if you want to dig into the details, you can direct them to where they can find the full text. Is that how it goes?

Jonathan Eads: Absolutely. We do not provide direct access to the full text. We link out to it, but we do record a bit of text right around where the variant was found. And that’s super valuable in some cases. We can build actual evidence that can support a particular variant classification by doing that. That’s very useful for our customer base, and being able to go right to it saves a lot of time.

Ross Katz: Yeah, absolutely. Having all of the resources related to the variants you care about in one place, being able to at a high level understand what was written about these different variants within the papers, and then if you need to, dig deeper into a given article or a given piece of research, you can go off-site and do that. That makes a lot of sense. I think I’m starting to understand why you need human curators to be part of this. Can you talk a little bit about where you use software processes, where you use people processes, and how you combine them together to get the most accurate research outputs you can?

Jonathan Eads: Sure. The software is used for scanning the full corpus of full text — that’s beyond human capacity. We’ll do our absolute best using software to extract the genes and variants. But there are false positives. Gene symbols — every gene in the human genome has a symbol, a string of characters that represents it. Sadly, some of those strings are very short, two or three characters. They’re also not unique to the genes. In some cases they’re reused and it could be the same name as a cell line in biotech. You’re going to get a bunch of those papers coming in. It could be a plasmid, something just not related to the gene of interest. Worse than that, they can come up in completely other disciplines. The Journal of Optics might have a symbol that matches the string for a gene and it has nothing to do with genetics. All of that we try to address with software. Where the curators come in, they make that call of variant classification. Is this benign or is this pathogenic? There’s a set of guidelines provided by the American College of Medical Genetics, ACMG, for how you do that. It’s pretty structured and it needs to be. It’s a formal process you have to go through. You don’t want to get that wrong. We don’t want to say something’s pathogenic when it’s not, and we don’t want to call something benign when it’s pathogenic, because that can have some real negative consequences. Those calls are always done by human curators. The software’s job relative to their responsibilities looks like: how can we improve their efficiency? How can we make their job easier? How can we put the right information in front of them with more precision so they can make a decision faster?

Ross Katz: Let me just see if I’m thinking about this correctly. You’ve got a pipeline with all of the articles coming in — obviously that’s oversimplifying, because you need to gather them from a variety of different sources, some of which have APIs, others where there’s just a variety of ways the text enters the pipeline. After the text enters, you’re trying to assess: does this article contain any genetic variants, and if so, which ones? Is that where the software part ends, or are there other things you’re doing inside the software system?

Jonathan Eads: You are thinking about it right, but there’s a lot more we’re doing inside the software system. I was trying to leave out some of the complexity. It’s not just about getting the variants and the genes. Is this the canonical transcript? Do we have multiple transcripts? Are we describing a different transcript? Splice sites is another level of complexity. And then once you get that stuff, can you determine from what was written what build of the human genome they’re using? You have different versions of that build and the coordinates completely change. They might present you with the protein coordinates for a variation. Those are going to be normalized from the starting amino acids. That doesn’t get you to the genomic coordinates. cDNA, similar situation. You’ve gotta have your own code to map that back to genomic coordinates. You’ve gotta have your own code to figure out: is this the canonical transcript? Are there multiple transcripts? All of that the software has to do. Then the code has to do its own QC as well. Let’s say there’s a publication with multiple genes and multiple variants, and you’re pulling it out of full text. Did you get the association of those variants to the right gene? What if you got that mismatched? What if you got the wires crossed? It’s complicated.

Ross Katz: That makes a lot of sense. There’s an entity resolution problem, and then a metadata extraction problem, and all of that needs to be QC’d. What I’m hearing is that you have an automated version of QC happening inside the software pipeline so you can minimize the burden on your human annotators — keeping those highly skilled people focused on the things that actually require their expertise. Is that right?

Jonathan Eads: That’s exactly correct. One other thing to mention: we do index additional information — we’ll try to map terms to the phenotype ontology that we can get out of the paper at the same time. There’s a variety of things like that we’re leveraging the search index for, so they can be provided as additional metadata that comes along with the variant.

Ross Katz: Yeah, that’s metadata that’s already there in the paper, and the author and journal would have a vested interest in having accurate metadata associated with the paper as well. That makes a lot of sense. As you start thinking about making changes to this pipeline and incorporating more advanced NLP methods, how do you think about testing the system to make sure you’re moving it forward rather than backward?

Jonathan Eads: This is a real challenge. The way we’re thinking about it is, we need validation data sets where we can compute recall, precision, F-score, accuracy repeatedly in a software development context. Building up those corpuses is not trivial. We’ve got all of these full-text articles at our disposal, but which thing are we going to test? We need a corpus where we test specifically getting the variant-to-gene association correct. What happens when there’s no variant and it’s just genes? Well, that’s a whole other situation — particularly prone to false positives because of those reused gene symbols. That’s another validation corpus you’d want, where you can recompute precision-recall accuracy after any kind of change so you can see, are we doing better or worse? One of our observations is there’s just a long tail of anomalies. The bulk of recent publications will roughly fit into HGVS nomenclature, but then you have this tail where the frequency of individual occurrences of a particular syntax anomaly isn’t very high, but there’s a ton of different ones and they go on and on. That’s a challenge because you’re optimizing for one-off cases. At the same time, you add them all up and it’s a significant number — we’ve got to put a dent into that. This is an area where AI is particularly promising.

Ross Katz: Awesome. To the point about the long tail, the beauty of large language models — the ChatGPTs of the world — is that they’re really good at identifying semantic overlap and resolving entities that previous NLP systems couldn’t have resolved before. But the challenge there, I would imagine, is that you’re going to get new forms of noise introduced to the system, new forms of false positives that your validation sets had never seen before. Do you have ideas about how you approach that problem?

Jonathan Eads: It comes down to getting that validation data set right and being able to catch whatever AI model we’re using, some of its shortcomings, early on. We’re always trying to get some curator time to further label and examine AI model output and make a determination on what’s happening. It can get really nuanced — the gene symbol occurs in a full-text article, but it’s in the context of the protein and not the context of the gene and is not really indicating any kind of variant, but the symbol is correct. That kind of context, amazingly, AI can pick up on. But it’s not always right. It’s something where you’ve got to do some fine-tuning a lot of times. Another thing with the GPTs of the world is the context window size compared to, say, a BERT model, is ginormous. You can fit the whole full-text article in there, which opens up the doors for — I don’t know what the right term for this is — distance associations. You have genomic coordinates in the caption for an image that are relevant to natural language being used in the results section. Try to do that with old-school NLP and you’re going to really be scratching your head.

Ross Katz: So there’s an opportunity for those associations that you would have missed previously to get surfaced. And since you have a people-based curation process at the bottom of the funnel, there is some tolerance for error there, which to me means you’re set up for success in adopting these AI tools. What do you think?

Jonathan Eads: It has to end in human curation at this point. AI is not ready for making variant classification calls. It just isn’t. That’s not to say it won’t be one day, but it’s not today. So it always ends in a curator double-checking stuff. We do have some tolerance for false positives, hopefully not too many false negatives. A general philosophy in our infrastructure is to cast the net as broad as possible. We’d rather tolerate false positives than miss something. That has been a philosophy since the very beginning in our engineering architecture. It’s served the company well — our recall’s pretty good; precision, we can improve that.

Ross Katz: That makes sense. You want the output you provide to your customers to be as comprehensive as possible — capture all of the known associations, because your customers are using it for research anyway. They’re going to apply the appropriate level of scrutiny to whatever the results are, and providing more perspectives on the roles and functionality of these different genetic variants is just more information for them.

Jonathan Eads: Exactly. We’d rather do that, and unfortunately there are going to be some genes with lots and lots of variants where the false positives become a bit of obfuscation — there are too many of them. That said, if we went the other way, you’re going to have a false negative and you’ll never know it. And that’s really not good. The false positives — we’re coming up with better strategies to manage that. We can come up with better granular classifications of the articles themselves. Is this a clinically relevant publication, or is it a review article? Is it human genetics or mouse genetics? Is it somatic or is it germline? We’re building systems that can assign those granular classifications, which makes it a lot easier to sort through the junk.

Ross Katz: As you think about developing the evaluation or validation corpus you’re using as you’re making changes to the NLP pipeline, is there any change in your processes for your human annotators to try to improve the quality of the data that’s being fed into your system so that these AI systems can learn over time?

Jonathan Eads: Yeah, we have our internal tools that the curators use where those labels of article classification — we’ll have them give us a yes or no, give us feedback that gets recorded, and we can then incorporate that back into whether the model got it right or not. We can fine-tune based on the curator continually providing us with more information. That’s an area where we’re constantly looking for light-touch ways of getting more high-quality curations to improve the models we’re building.

Ross Katz: So there’s a short-term, long-term trade-off. There are so many variants of unknown significance out there, and you’re constantly swimming against the tide of all the annotation that needs to be done to create data to feed the pipeline. But you also have the long-term benefit of additional annotations your people can make that help teach the system and maybe automate more things in the future. How do you manage that trade-off?

Jonathan Eads: That’s a tough one. The curators’ time is going to services where we’ve got a real pharma company wanting this set of genes and all their variants curated. As we dig into this, we’ll leverage interns, graduate students, anybody we can get. But I think at a certain point we’ll probably need dedicated curation resources for AI efforts. We’ve been able to get by so far with squeezing resources through the cracks, but we’ll see how long we can push that for.

Ross Katz: You’ve got this really great team of expert curators, but not all of the annotations that need to be done require quite that high a level of expertise. And for the long-term benefit of the company, you need at least some of that high-level expertise to move the AI forward. So some combination of less experienced people and your smaller team of experts sounds reasonable.

Jonathan Eads: It’s a mix. Some of the AI labeling efforts — you don’t need a lot of skill. You need some, but not a ton. Other AI labeling efforts, you need a ton of skill. Like getting the nuance of AI model output — the example I gave where the gene symbol is referred to in protein context, do we want to capture that? Is the model getting the context right? Some of those decisions get really complicated. It gets really complicated when we get into copy number variants. Then we need a very high skill level — we need the best we’ve got. So it’s a broad spectrum of skill level that we require to get through all of the labeling needs for getting AI right.

Ross Katz: It’s almost like you need a classification model for how much expertise is required to annotate a given example. This discussion of the NLP aspects has been really interesting, but I just want to zoom back out to some of your customers and the way they’re using Genomenon. Do you have any examples of success — reducing the time, effort, or money needed to advance research efforts?

Jonathan Eads: We have a lot of examples. We’ve got examples at the doctor-patient level, at the pharma company level, and at the biotech company level. At the pharma and biotech level, some of the key questions we’ve been able to help companies answer are: how many patients can be positively impacted by going after a particular therapy? How many folks need help? What we can provide is an evidence-based, fit-to-purpose strategy for calculating genetic disease prevalence. There’s a big need for that with different types of diseases. That evidence-based threshold we can help derive can be taken to make a case to the FDA, and we’ve had that happen before. We’ve had some really nice success stories determining prevalence of rare disease. Another area: our data and platform has been used to get variants on and off regulatory lists, to expand clinical trial labels to include additional variants that would otherwise not have been considered. There’s also more pragmatic value where a pharma company has a gene or collection of genes that’s very relevant to a disease treatment they’re going after and they want expert-level curation of every variant out there. That’s something they can come to us for in our services. And then some of the stories that are very impactful — and at times a bit gut-wrenching — are the doctor-patient stories where we’ve got a patient with a rare genetic variation and all they’ve really got are the variations and the phenotypes; they don’t have a diagnosis. They can come to our platform and identify that disease and at least get to a diagnosis, ideally get to a treatment. That’s not always available, but getting a diagnosis is the critical first step. One value in that area that I was not aware of until we had customer stories about it: they can use our platform to justify getting a variant covered by insurance. That can have a huge impact on a patient’s life, and it’s happened multiple times. One publication for a variant of unknown significance can tip the scale in a variety of different ways. That’s why our platform is so important.

Ross Katz: Awesome. Well, as we head toward the end of our time together, I just want to take a little bit of time and look toward the future. You’ve mentioned some of the AI enhancements you’re doing in your NLP pipeline, but what do you see as the major advancements you expect in literature search for genetic variant analysis over the next five years or so?

Jonathan Eads: My hypothesis is that we’re going to see an end-to-end linkage of information that we haven’t been able to do previously. Going from variant to gene to disease to phenotype all the way out to EHR — electronic health records — and actual patient population-type information, and being able to traverse that linkage and do inference in both directions. Meaning I could come with a set of symptoms and get all the way to variants associated with it, or I could come with variants and get all the way to symptoms associated with it and maybe a patient population. That kind of association in the data would bring a level of value we’re not quite at yet. It’s something AI is capable of doing, and the information is there. It’s going to take some work to get it, but that’s what I’d like to see happen.

Ross Katz: For sure. And I heard that Genomenon recently announced an acquisition of JAX Clinical Knowledgebase. Would you just talk a little bit about the rationale for the acquisition and what excites you about it?

Jonathan Eads: It’s a really exciting acquisition. The Clinical Knowledgebase is probably one of the best resources for curated variants associated with cancer, specifically somatic cancer — cancers that are specific to a tissue type. Genomenon’s main corpus covers germline cancers pretty well. With this acquisition we’ve got germline and somatic under one roof. That opens up some doors to provide an integrated experience that customers just haven’t had. The potential there is enormous for providing a one-stop cancer resource for the community. I’m really excited to work on that. There’s a lot of talk about ways we can integrate the two platforms and I think there’ll be some exciting stuff to come there.

Ross Katz: Awesome. And where can people go to find more about you and about Genomenon?

Jonathan Eads: The website’s always a good place to start for Genomenon. We just released a new website, so that’s fun. A number of the success stories I was referring to — you can get more information there. You can find me on LinkedIn. Happy to chat with anybody about this stuff. There’s a lot of exciting work going on at Genomenon and in the space in general.

Ross Katz: Fantastic. Well, Jonathan, thanks so much for joining us today. It’s been a really interesting conversation and I’ll look forward to connecting more down the line.

Jonathan Eads: Thanks so much, Ross. It was a pleasure chatting.

Jason: And that’s it for this episode of Data in Biotech. If you enjoyed the episode, please subscribe, rate, or leave a review in your podcast platform of choice. See you next time.

Frequently Asked
Questions

How does Genomenon ensure data accuracy when processing a vast amount of unstructured scientific literature?
Genomenon combines broad automated extraction with essential human oversight. Software scans over 9.5 million publications to identify genes and variants, prioritizing comprehensive recall. However, critical variant classification (benign or pathogenic) is always finalized by expert human curators following strict American College of Medical Genetics guidelines, ensuring the highest level of accuracy for patient and research outcomes.
What are the biggest technical hurdles in extracting and normalizing genetic variant data from research papers?
The primary challenge stems from a 'tsunami of poorly adhered-to nomenclature standards' where authors frequently deviate from established genetic syntax or use natural language. This is compounded by the lack of standardization in supplemental materials—often containing critical data in unsearchable formats like rasterized images or PDFs—making reliable, automated extraction difficult.
How does this platform directly accelerate drug development and expand clinical trial possibilities?
Genomenon helps pharmaceutical companies by providing evidence-based insights into disease prevalence, which is vital for regulatory submissions and justifying therapeutic development. The platform also enables the identification of additional, relevant genetic variants that can expand the scope of clinical trials, thereby accelerating research and development pipelines.

Need a data partner for life sciences?

CorrDyn helps biotech and pharma companies build the data infrastructure that accelerates research and operations.

Book an intro call