Skip to content
Data in BiotechEpisode 14

Challenges of NGS Data Sets with Joseph Pearson

Joseph Pearson of QIAGEN OmicSoft discusses challenges of NGS data sets, the value of APIs, and selecting the right bioinformatics tools.

35:00Full transcript below
JP

Joseph Pearson

Global Product Manager, OmicSoft at QIAGEN

Overview


Biotech organizations often face a hidden cost: their data scientists spend 80% of their time preparing and cleaning NGS (Next-Generation Sequencing) datasets, not analyzing them. This heavy lift in data acquisition, curation, and unification directly impedes drug discovery, slows therapeutic development, and siphons resources from critical research. Without a reliable system, answering even a ‘simple’ gene expression question can take weeks, delaying crucial insights and hindering competitive advantage.

This episode addresses that core problem. Ross Katz speaks with Joseph Pearson, Global Product Manager for OmicSoft at QIAGEN, a company focused on sample-to-insight solutions. Pearson brings a unique perspective, having transitioned from an academic background in gene regulation and bioinformatics support to product management. He understands both the scientific rigor required for NGS data and the operational challenges of making it useful at scale.

Pearson details the significant effort required to transform raw public datasets into reliable, unified resources. The conversation covers the complexities of managing metadata, the pitfalls of internal build-it-yourself solutions, and how a specialized platform like OmicSoft accelerates time to value. This discussion provides a clear framework for data leaders and executives considering their bioinformatics infrastructure strategy.

Key Takeaways

Data preparation consumes 80% of bioinformatics effort

Biotech data scientists spend a disproportionate amount of time—up to 80%—on finding, cleaning, and unifying NGS datasets before any analysis can begin. This significant overhead delays research, inflates project costs, and limits the pace of scientific discovery. A reliable solution must address this preparation bottleneck directly to accelerate insights.

Manual data curation is essential for NGS data quality

Public NGS datasets, while abundant, frequently contain subtle errors such as switched sample labels, inconsistent terminology, or misreported clinical parameters. Automated pipelines often miss these issues, leading to flawed conclusions. Dedicated manual curation, like OmicSoft’s multi-year protocol, is critical to ensure data integrity and reliable downstream analysis.

The ‘build vs. buy’ decision must account for perpetual maintenance

Companies considering an in-house NGS data infrastructure often underestimate the ongoing burden. Beyond initial development, challenges include maintaining undocumented systems after key personnel depart, integrating disparate freeware components, and managing unversioned data in disconnected files. These issues create a ‘curse of success’ where scaling data means diverting resources from analysis to infrastructure upkeep.

Flexible data access accelerates diverse user needs

To serve varied bioinformatics users—from biologists needing quick lookups to data scientists building custom models—data solutions must offer multiple access points. Providing graphical user interfaces, SQL-enabled APIs, and direct flat file downloads allows teams to choose the most efficient method for their specific task, preventing data from becoming ‘walled off’ and unused.

Related: CorrDyn works with biotech and life sciences organizations, providing data engineering and data quality services. For considerations on building in-house versus external expertise, review our insights on hiring an external data team.

Full Transcript

Jason: Hi everyone, this is Jason, producer of Data in Biotech. Before we get started, I wanted to let you know about our latest white paper. It’s a comprehensive guide to implementing machine learning models in biotech manufacturing. It’s a complete overview of all the potential problems of ML adoption and, more importantly, how to solve them. To download it, simply visit connect.corrdyn.com/biotech-ml. We’ve also dropped the link in the show notes of this episode. Okay, let’s get into it. Welcome to Data in Biotech, a podcast from CorrDyn where we explore how companies leverage data to drive innovation in life sciences. Every two weeks, we sit down with an expert from the world of biotechnology to understand how they’re using data science to solve technical challenges, streamline operations, and further innovation in their business. This week, we’re excited to be joined by Joseph Pearson, Global Product Manager of OmicSoft at QIAGEN, a global provider of sample-to-insight solutions that enable customers to gain valuable molecular insights. During this discussion, we dive into OmicSoft, a powerful NGS analysis suite with the ability to quickly explore and compare across 500,000 curated OmicSoft samples from disease-related studies. Joseph outlines the challenges of acquiring and analyzing NGS datasets, discusses the different ways customers can interact with OmicSoft data, and unpacks the build-versus-buy debate when selecting new bioinformatics tools. Here we go.

Ross Katz: Joseph Pearson, welcome to the Data in Biotech podcast.

Joseph Pearson: Thanks Ross. Very happy to be here.

Ross Katz: Awesome. Well, just to kick us off, would you just give us a brief introduction to your career and what brought you here today?

Joseph Pearson: Sure. I started out as a long-time academic, focused on model organisms. I was working in fruit flies, really interested in gene regulation. This brought me into my affinity for basic bioinformatics tools, enhancer bashing and interpreting DNA sequences to understand gene regulatory networks. Through that work, I developed a bit of a knack for explaining bioinformatics concepts to biologists, and vice versa. After a long academic career, I joined a small company called OmicSoft, a local bioinformatics company that had started up, and joined their tech support team, which was a plum job interacting with customers to help them answer their scientific questions. It was the best part of bioinformatics, and I continued from there.

Ross Katz: And now you’re in a product management role at QIAGEN. How did you end up in product management and what drew you to that part of the tech landscape?

Joseph Pearson: I really just grew into the role. I was on the tech support team for several years, grew into the manager role, but it was the experience interacting with customers that helped me understand what their needs were. I was a pretty good fit for the Global Product Manager role for OmicSoft, and stepping into it has been a great opportunity to learn a lot of new skills that I would not have developed otherwise.

Ross Katz: And why do you do what you do? What motivated you to go down this path in the first place?

Joseph Pearson: It’s a genuinely exciting opportunity to figure out, on a day-to-day basis and from a more strategic level, how we can make our data and our software more useful for the scientists I’m working with. There are a lot of little decisions and a lot of really big decisions, but there’s always a new challenge every day — something to look forward to.

Ross Katz: Can we talk a little bit about QIAGEN and the OmicSoft product? Maybe just orienting us to QIAGEN as a company and then how OmicSoft fits into the portfolio of solutions that you all provide would be great.

Joseph Pearson: Yeah, that makes a lot of sense. I suspect a lot of people are familiar with the QIAGEN name. It’s a big global company, very well known for wet lab reagents and the big blue and red boxes you’ll see everywhere. QIAGEN Digital Insights is a small division in the context of the large global company, but it’s a division that was purpose-built to cover the spectrum of bioinformatics needs, including bioinformatics pipelines, data analysis, interpretation, data warehousing and management, as well as the underlying resources of high-quality manually curated data and knowledge for basic research as well as for drug discovery. OmicSoft is one part of the QIAGEN Digital Insights group. QIAGEN Digital Insights has multiple products and product lines, and OmicSoft is one of them. Our strength is focused on NGS analysis, especially RNA-seq, and the organization, management, and unification of those NGS datasets so scientists can actually use those data after they’ve been generated. That’s what we’ve focused on over the years.

Ross Katz: What I hear is that you’re providing NGS datasets that have been cleaned and prepared, but also a set of tools for analyzing those datasets and getting value from them.

Joseph Pearson: Yeah, that’s basically it. The actual origin of OmicSoft was as a microarray analysis platform, grew into NGS during the NGS revolution. It developed organically through interactions with our early customers, who were initially asking the OmicSoft team to pre-analyze public datasets under a unified pipeline and clean up those datasets so they could use them for larger meta-analyses. This grew over time to the point that we built up a substantial database of unified datasets and a robust curation protocol, analysis protocol, and statistics protocol to bring these data together so people can get right into the data analysis. We’ve built up the data as well as the tools to interact with those data over many years now.

Ross Katz: From the perspective of a customer that uses OmicSoft, the time to value is one of the key things you get out of it — you’ve gone through all the work of preparing the data for analysis and serving it to researchers in a way that makes it easy to get insights. Who are the typical customers you see coming to use OmicSoft?

Joseph Pearson: I can usually tell in a room who are the people who’ve actually gone through this painful process. If you look at a data scientist’s day, about 80% of their time is spent finding the datasets to analyze, about 20% is spent actually analyzing those datasets. When we explain our curation protocol, the data analysis we’ve done, the unification, and show how much time we’ve spent ahead of time so they can get to that data analysis step, you can see certain people just nodding their head: ‘Yes, that is what I wish I had five years ago.’

Ross Katz: Maybe it’s worth grounding this in an example of a team in the absence of OmicSoft, without the pipeline in place for them. What would be the burden on a group starting from zero and trying to do everything you’ve done for them?

Joseph Pearson: It really varies on how big the question is and the type of disease they’re looking for, but the standard place to go is one of the public repositories — either a website with a built portal to look for these data or just a data repository. You find a GEO or SRA dataset, download the dataset, analyze it, QC the data, do the secondary analysis to interpret the dataset, and then try to match it up with some other comparable data. At that point you might realize: was it a good quality dataset? Was it a good experiment, or was it iffy? Once you’ve gone through that process, it’s several days to weeks of analysis. Most of our customers are relatively small bioinformatics teams serving a bunch of biologists, and the biologists have very simple-sounding questions. Is this gene expressed in such and such tissue? Is it upregulated in a disease? Does it have a mutation in a cell line? These are simple-sounding questions, but to actually answer them properly takes a lot of work — this whole process. You have to find the dataset and make sure it’s the right dataset to answer that question. The difference is that our customers can, much of the time, just mine our data and use our interfaces to answer those questions in minutes as opposed to weeks.

Ross Katz: A couple of things come to mind. The scope of ad-hoc requests that could come into a small bioinformatics team is relatively broad, so having a comprehensive dataset already cleaned and ready for them — the way OmicSoft does — is a key value proposition. And if it’s a question about gene expression in a particular tissue sample, I’m imagining you have all of the metadata around each of these NGS datasets ready for the bioinformatics team to easily answer that question based on the tissue it came from.

Joseph Pearson: For very basic questions, a lot of these could be answered using other tools. If you’re just looking for an overall view of tissue or disease, you can find public resources that relatively quickly can be mined from a single source. The moment you get to a more detailed, subtle question — like a treatment response question, or trying to mine additional metadata describing the clinical parameters of the patients or subjects, or anytime you’re trying to merge multiple datasets — that’s where you really run into trouble because everybody uses their own terminologies and ontologies. You have to do all that cleanup. In our experience, there are frequent instances where there are clear issues with the data as submitted to the public repositories: sample labels switched, spellings or switching of treatments, that would give you the wrong answer if you just took them at face value.

Ross Katz: So you’re going through that quality control process and that metadata validation process on behalf of the clients, dataset by dataset, making sure that the labels are consistent within the ontologies you’re providing and that none of those tags have been flipped, as you described.

Joseph Pearson: That is one of our core value propositions — we’ve already done this work so customers don’t have to worry about it. This has been built up over time through the years talking to our customers about what is important. They’ve indicated that it’s valuable to have these detailed metadata curations with definitions of the metadata fields and quality control and unification to save them time. There are customers who will take our data and reanalyze the actual raw data on their own pipeline, because it’s easy to automate that. It is not easy to automate data curation to make it unified, and there are lots of pitfalls in that. People like to trust that we’ve built up the manual curation protocol to catch those errors.

Ross Katz: When a potential customer sees a problem where they know they need to go out and acquire a dataset to answer either a particular question or they’re growing their bioinformatics capabilities and they know that they need more ad-hoc capability to respond to these types of questions for which data doesn’t exist yet, how do they evaluate OmicSoft and the tools that you provide when they’re considering using it?

Joseph Pearson: The first question is: do you have data that are relevant for my question? The nice thing about building up the databases for so many years is that, through working with our customers and filling their requests, we’ve built up a broad collection of data. There’s a good chance we have robust data useful to answer their question, and then we can come back to them and clarify. When you said inflammatory diseases, which of these 57 different specific diseases did you mean? We can start to talk about the curation protocol. Maybe it’s more detailed than they need and they realize there are only a couple of datasets out there, or maybe they realize there’s actually a significant amount of data and they’re willing to go through the effort of mining it. We provide the data and the tools, but they still have to use their bioinformatics expertise to pull out the relevant insights and use their judgment. If they decide there’s enough data and we’ve saved them time, they’ll subscribe.

Ross Katz: Interesting. Where does the data underlying OmicSoft come from? Are these public datasets that you’re cleaning and organizing, or proprietary datasets that you’re licensing?

Joseph Pearson: Almost exclusively public datasets — retrieving the same datasets that one could find for themselves, but we are doing a lot of additional work on those datasets to actually make them unified. We do have some partnerships with proprietary datasets, for example the collection of cell lines that are certified. A nice value from that is that people can mine all of our other cell line data that we’ve curated and unified and then discover the ATCC verified cell lines that match those same data and check at the metadata level and at the gene expression mutation level whether they actually match up. Because these ATCC data have been generated under ISO standards from validated lots, you can guarantee that the aliquot you get from ATCC will match the ATCC gene expression profile that we have in the Lands.

Ross Katz: So that’s a team of people in-house monitoring the new datasets becoming available, any research being released, and then manually cleaning and integrating the datasets into OmicSoft. How does that process work from your perspective?

Joseph Pearson: Yes, we do have a team who are monitoring for new interesting datasets and routinely do searches for keywords from past requests from customers, but we strongly encourage our customers to send in requests of their own. We like to prioritize those and get them in as quickly as possible because those are the data that are relevant to them. We’ll use that to guide our future searches to find new datasets that are important.

Ross Katz: Is there a turnaround time that your customers generally expect or hope for when they make a request for a new dataset, or does it depend on the complexity of the dataset and the number of people requesting it?

Joseph Pearson: They would always like it to be faster, but it does take a little bit of time to curate these datasets and unify them into the next release of the database. We have a well-established quarterly release cycle, but we absolutely can do accelerated turnaround for specific customer requests as a custom service.

Ross Katz: From the perspective of the bioinformatics team or the data science team, how are they interacting with these datasets? You mentioned a database. Are they querying a database, or using the programming language of their choice to interact with an API? How does it look?

Joseph Pearson: It really depends on the user. We provide a couple of different ways to access the data, ranging from pre-built graphical user interfaces that have been purpose-built around the structure of the data — both heavy clients and web-based access. That’s great for quick lookups of a gene’s expression or to find a specific dataset for the occasional lookup. For the more advanced use case, people are usually asking for API programmatic access to get into the data, and that’s where you can really slice and dice through the data, find datasets of interest, build custom cohorts across the datasets, do custom analysis, aggregate datasets and do statistics on the server as opposed to having to download everything. We also provide full flat files. We have customers who are just using us as a data provider and bring it into their own database. It’s really about the use case. We have customers doing all three because they have different user groups, or even just within a single day: do I want to do a quick lookup, do I want to do a robust analysis, or do I want to build a custom model on these data?

Ross Katz: So if there’s an ad-hoc request that just needs the fastest possible answer, you go to the user interface and get the answer as quickly as you can, but if you’re building data infrastructure and data capabilities for the long term, you want that more programmatic integration where you’re bringing the data down into your own environment and aligning it with the datasets you’re collecting as a biotech organization. Am I thinking about that right?

Joseph Pearson: Yes. That’s really been the feedback from customers — they want that flexibility. They don’t like things to be walled off, but they do like the guardrails and the simplified interface for those quick lookups.

Ross Katz: Have there been any sort of surprising integrations or surprising applications you’ve seen of either the tools that that you’ve provided for analysis or the datasets that you’re providing?

Joseph Pearson: Some of the more interesting integrations were using the APIs. We were surprised how quickly certain people really glommed onto the idea of the more powerful APIs to use SQL-based queries to slice and dice and aggregate on the data. These were customers that had had access to the full flat files for years — they could have been doing these analyses just by downloading gigabytes of data as tabular storage — but once we provided access through APIs querying on our hosted side, they could really focus on the scripting side as opposed to the database management side.

Ross Katz: That’s interesting. So there’s a SQL interface integrated with the API, where you can do your aggregations and filters directly in the API call, and it’s flexible to all of the fields you have available in the database that you expose via the API. Am I thinking about that right?

Joseph Pearson: Yeah, that’s basically right. One of the fun inventive aspects of these new APIs is that we’ve been able to reapply some of the visualizations and analytics that were built up from customer requests in the heavy software, and rethink how to answer the customer question — not just reimplement the algorithm or the tool, but ask: ‘What is the customer question they’re trying to answer?’ Our very talented developers in QIAGEN have come up with clever solutions to efficiently do this with a simple API query so you can search over lots of data at the same time and pull out aggregations and correlations, differential expression analysis, mutation frequency summaries, all kinds of ways to aggregate the data. This basically builds up a little toolkit that our customers can take and plug and play.

Ross Katz: There must be a significant customer education component to what you have to do, because there are so many questions you can ask of the data. And when you talk about exposing a SQL-based API, there’s teaching people how to use SQL, how to query the database, what the limitations are on what they’re allowed to query and how much they’re allowed to query based on their plans. How do you approach that customer education problem?

Joseph Pearson: It’s the QIAGEN Digital Insights approach to have a dedicated team of very talented scientists focused on engaging with those customers, talking to them. First of all, they are our eyes and ears — when a customer has a complaint or requests a new feature, these application scientists bring it to the product managers to prioritize. They are also the ones presenting how to answer specific questions to the customers. It’s really about that high-frequency interaction to make our customers feel like it’s a partnership, not just a vendor-client relationship.

Ross Katz: That makes sense. When an ad-hoc request comes into the bioinformatics team and it’s a question where they don’t know how to access the data they need, they’re talking to someone at the Digital Insights team who’s helping them wade through all of the data available and find the right query to get them the dataset they need to analyze.

Joseph Pearson: Exactly. I don’t think there’s an application scientist on the team that doesn’t love to have the answer at their fingertips. It’s a point of pride that if they don’t know how to do something, they are going to find the person who will help them understand, and then they can explain it to the customer and make the customer feel like they’re a superstar.

Ross Katz: You’ve already shared an example of an ad-hoc request that came through, but can you share some practical use cases companies have executed using the OmicSoft product?

Joseph Pearson: Sure. I don’t think it’s unique to the OmicSoft product — it’s anytime a customer wants to ask a question of omic data. It’s really one of two types of question: ‘Where is my gene expressed, mutated, amplified, deleted?’ — whatever your omic metric is, in such and such dataset, treatment, tissue, disease, cell type, whatever the cohort is. The second question is: ‘What are the important genes for my tissue, treatment, disease, cell type?’ Each of those questions starts with a big matrix of data that you whittle down to find what’s going to answer your question, and then the details come into how you’re specifically going to mine the data to decide: is this actually giving me the answer that’s helpful to go to my next step? Because of the unification of any database — and this is why people build an integrated database in the first place — you can do analysis across. Whether you’re building your own or subscribing to a database and software solution like OmicSoft, the big benefit is the downstream analytics you can perform on these data. It’s not just finding that slice of data, but looking for correlations, doing new custom statistical analysis, and then sending to the next step in your pipeline.

Ross Katz: I can imagine a startup or scale-up biotech organization building up its bioinformatics capability, focused on a particular therapeutic area, trying to decide whether to build the capabilities that you have in OmicSoft in-house or to buy into the platform and leverage all of the time and effort and energy that went into the development of your platform. Can you help us think through how a company like that makes the build versus buy decision in that context?

Joseph Pearson: To try not to be too glib about it, it comes down to whether the person in charge of data management, warehousing, and mining has been burned before by a build-it-yourself approach. If they’ve tried it and realized it’s a lot harder than they thought to maintain, then they start to look for custom solutions. We’ve seen people go both ways — start build-it-yourself and then come to us and say, ‘Okay, actually we could use this solution.’ Sometimes they’ll come in and say they’d like to save time up front by using a turnkey solution, and then realize they need a bit more customization and would rather just use the raw data. That goes back to the feedback we’ve gotten from customers for years about flexibility. They need access at different levels: flat files to integrate into their own database all the way through a pre-built solution that just visualizes the data, no coding required.

Ross Katz: What are some of those nightmare scenarios that someone starting up a bioinformatics practice might not be aware of — that you’ve seen people come back and see the value of OmicSoft to avoid?

Joseph Pearson: The telltale signs are either that they have a complex internally built solution using a bunch of freeware components and the person who built it left without documenting enough, so they’re stuck with a solution they can’t put any new data into or update. Or they’ve gone the old-school route and they’ve got folders full of Excel files.

Ross Katz: And where does that leave them with the folders full of Excel files?

Joseph Pearson: They’re basically stuck with all of these data and they can’t actually do anything with them unless they open up the Excel file and discover that many of their genes have been converted into dates, or there was a drag and drop mistake, or there’s just no version control. All sorts of problems that you discover once you start to actually try to integrate these data.

Ross Katz: I can imagine there’s also a curse of success here — if you’re one of these smaller biotech organizations that managed, with a small bioinformatics team, to build the data platform you needed, now you’re responsible for maintaining it in perpetuity. There’s a constant flow of new data you want to bring in, and now you’re spending all of your time just managing and maintaining data infrastructure when what you really want to be doing is driving value from the data that’s running through that system. Am I hearing that right?

Joseph Pearson: That is absolutely a path we’ve seen. Some people will decide to bring in some sort of solution. Maybe it’s not even a full redesigned solution, but just work with the QIAGEN Digital Insights services team on modifying or integrating our data to work better with their internal pipeline. There are lots of different ways to freshen up an internal data lake or interface.

Ross Katz: Does the size or maturity of the organization play a role in the decision to build versus buy in this context?

Joseph Pearson: Maybe it’s a little bit of an hourglass or some multimodal shape. If you’re too small, it’s not worth the effort of actually subscribing — it’s better to just do an ad-hoc download and analyze the data. There’s a certain size where you realize you’re actually building up some internal data resources that you want to integrate with public data, and it’s going to be quite a lot of effort. You may have to hire another person, or you could subscribe to a software and free up the existing team to do the work. As you get to a big enough organization, you could make the decision to build an internal framework that’s custom built towards your solution. But what we see frequently is that even the big organizations are very siloized, very fragmented, they don’t talk a lot, they have different opinions. They’re less unified than you would imagine looking from the outside.

Ross Katz: Interesting. As a product manager, I imagine you get exposed to a lot of customer feedback. What are some of the biggest challenges or criticisms of the platform that you hear from customers?

Joseph Pearson: The data are complex, but occasional users want the data to be simple. There’s always a learning curve to introduce to any client, and you can either build more and more customizations to make it simple for the new user, focus on bringing in new data and data types, or add more advanced capabilities for the advanced users. That’s one of the challenges, but also one of the joys of being a product manager — understanding what’s actually going to help people do their job better and move targets down the discovery chain faster.

Ross Katz: I can imagine you have to think through these problems and the different ways they can be solved. You’ve mentioned this strong service and support organization that helps with the customer education component. I’m imagining there’s always this decision criteria of: is this something that needs to be built into the technology platform, or is this something where you just need to do a better job of educating your customers about how to leverage the tools you provide today? Is that right?

Joseph Pearson: Yes. It’s really easy to lean on the crutch of our fantastic field application scientist team and tech support team and say: ‘User education is all that’s needed, we just need to train, train, train,’ maybe improve the documentation. But we do need to update the tools and make them actually easier to use for new users, while at the same time bringing in more capabilities for the advanced users who want to mine every last bit of useful information from the data.

Ross Katz: OmicSoft is part of this portfolio of products and services that QIAGEN provides. How does OmicSoft fit into the QIAGEN portfolio and what differentiates it from some of the other tools that are available?

Joseph Pearson: That’s actually part of the trick. Each component of the portfolio came in initially as its own independent company and was brought in because of its unique strengths, but there are always going to be overlaps. There are some overlaps and then some distinct features, focusing on different user personas, target markets, and so forth. Every day there’s some decision about a roadmap item or some feature that could overlap with another feature from another product, and we have to decide: ‘Is this something we want to continue to improve, or should we start thinking about a better way to integrate?’ That’s really one of the focuses right now in QIAGEN Digital Insights — to bring the data together, bring the software together to interplay a lot better and really focus on our strengths to bring a more unified solution for users so they don’t have to decide what’s the best path or which product to bring in, and we definitely don’t want them buying redundant products.

Ross Katz: Is that part of the sales intake process — routing the customer problem to the right product — or is it also a challenge for the marketing organization to figure out the specific pain points and customer problems that each of these products is designed to solve?

Joseph Pearson: That is the responsibility of the sales team, the marketing team, the product team — how to effectively communicate clearly what the value is so that there’s as little customer confusion as possible, because the confused mind says no. And even if they said yes, if they come in and didn’t get exactly what they thought they were buying, there’s going to be frustration. We want to make sure they know exactly what they’re getting and that what they’re getting is going to solve their problem.

Ross Katz: That makes sense. You come from this experienced biology background, you’ve done bioinformatics research work, and you’ve been on the front lines dealing with customer problems. How have these experiences colored your view of the challenges you face in your role in product management?

Joseph Pearson: On the positive side, I really enjoy building the use cases and telling stories about what you can actually do with the data and the software, and I can communicate that pretty effectively. On the other side, I really want to include everything. If there’s an exciting new technology, a new application, new data types, I can see how cool that is and how we could integrate it — let’s add these new capabilities. But you have to be very disciplined, have that ruthless focus. What is the thing that is actually going to bring the most value if we bring it in? Because we do not have infinite resources to make a true Swiss Army knife that covers everything.

Ross Katz: What’s on the roadmap for OmicSoft? What comes next?

Joseph Pearson: We’re still very excited about the new APIs. We’ve been very successful at our data offerings just by offering flat files and the graphical user interface, but the APIs we’re building are designed to make it as easy as possible to mine the data in ways that you normally wouldn’t be able to with home-brew solutions. There are a couple of key capabilities, including the graceful handling of missing fields and tables across databases, so that instead of getting an error and spending hours troubleshooting, it’s just going to handle those missing values and you get a bigger table — the ideal version of a big integrated database. There’s a lot of improvements we can add in there. We have some new data on the roadmap, again responding to customer requests. We can spend more focus on enabling new users to get into the data with new graphical user interfaces, but really the biggest impact is continued integration of our different data offerings. OmicSoft is one slice of the data portfolio, focused on curated omic data, but we also have a massive knowledge graph of curated biomedical relationships that underlies the IPA software. Data scientists need these kinds of relationships as a knowledge graph, and so we can improve the integration there as well as with our curated variant databases.

Ross Katz: Are there any other visions for the future you’d like to share, either from the perspective of OmicSoft and QIAGEN and where you’re going, or where you see the biotech landscape going?

Joseph Pearson: I think it’s just important to remove as many of those barriers as possible towards good data. That’s what we’re trying to move forward and make it easy for people to get real insights out of these massive repositories of data.

Ross Katz: Well Joseph, really appreciate you joining the podcast today. Are there any other places where people should look to find you?

Joseph Pearson: The QIAGEN website is a good place to find out about the product. It covers the QIAGEN Digital Insights product line and the webinars we’ve got routinely going on showing what you can do with these data.

Ross Katz: Well thank you so much for coming on. Look forward to connecting down the line.

Jason: And that’s it for this episode of Data in Biotech. If you enjoyed the episode, please subscribe, rate, or leave a review in your podcast player of choice. See you next time.

Frequently Asked
Questions

What are the biggest risks of building an in-house NGS data solution for analysis?
Joseph Pearson describes common issues: unmaintained systems after key personnel leave, reliance on a mix of freeware components, or unversioned data in disparate Excel files. These problems lead to unusable data, hinder scalability, and often incur significant rebuild costs, diverting resources from core research.
How does OmicSoft address the critical issue of data quality in public NGS datasets?
OmicSoft employs a rigorous manual curation protocol, developed over years of customer feedback. This process identifies and corrects common errors like switched sample labels, inconsistent spellings, and incorrect metadata, ensuring that the unified datasets provide reliable inputs for analysis and prevent flawed conclusions.
How does OmicSoft accelerate time-to-insight for bioinformatics teams?
By pre-analyzing and unifying vast public NGS datasets under a single, quality-controlled pipeline, OmicSoft reduces the 80% of a data scientist's time typically spent on data preparation. This allows research teams to answer complex questions—like gene expression in specific tissues—in minutes, not weeks, directly speeding up discovery.

Need a data partner for life sciences?

CorrDyn helps biotech and pharma companies build the data infrastructure that accelerates research and operations.

Book an intro call