Listen on
Overview
Pharmaceutical R&D faces a paradox: an explosion of scientific data creates immense potential for discovery, yet the sheer volume makes it impossible for human experts to keep pace, risking missed therapeutic targets and slowing pipelines. This challenge demands a strategic blend of advanced computational methods and pragmatic external partnerships.
Jesper Ryge, Director of Computational Biology at Merck (Germany), spearheads a lean team that addresses this directly. He outlines how Merck manages the complexities of neuroinflammation and neurodegenerative diseases by prioritizing data-driven target discovery. This includes deploying knowledge graphs to synthesize vast scientific literature, enabling link prediction for novel target identification, and strategically using external platforms for data curation and infrastructure.
In this episode, host Ross Katz talks with Jesper Ryge, who details how his team bridges the gap between computational insights and experimental validation, discussing the criteria for prioritizing druggable targets, the role of new multi-omics technologies like spatial transcriptomics, and the critical importance of interdisciplinary collaboration. His insights reveal a clear path for data leaders in life sciences to accelerate discovery while managing internal resource constraints and data quality.
Key Takeaways
External Knowledge Graphs Accelerate Discovery & Reduce Internal Burden
Merck, a major pharmaceutical firm, uses a lean approach for computational biology. Instead of building internal knowledge graph infrastructure from scratch to manage exploding scientific literature, they pilot and license pre-curated external platforms like QIAGEN’s Omicsoft. This strategy quickly provides thorough, high-quality data integration, freeing internal teams to focus on analysis and validation rather than curation and development.
Druggability and Feasibility Must Drive Target Prioritization
Identifying potential therapeutic targets through computational methods like knowledge graph link prediction yields dozens to hundreds of candidates. The crucial next step involves prioritizing these based on real-world factors: druggability (is it amenable to small molecule design?), safety profiles, and experimental feasibility. Integrating these criteria with internal expertise and public databases like OpenTargets helps narrow down prospects to the most viable and impactful for drug development.
Spatial Omics Reveals Critical Biological Context, Demands New Data Generation
New spatial transcriptomics technologies promise unprecedented resolution for understanding complex biological systems, particularly in challenging areas like neurodegenerative diseases where cellular microenvironments are critical. While these technologies offer deep insights into cell interactions and disease mechanisms, their novelty means a significant gap in publicly available data. Organizations must plan for substantial internal data generation or targeted partnerships to capitalize on their potential.
Interdisciplinary Teams Bridge Computational Insight to Experimental Reality
Effectively translating computational predictions into validated drug targets requires seamless collaboration between computational biologists and bench scientists. These interdisciplinary teams openly discuss experimental limitations, computational capabilities, and potential pitfalls in proposed models. This ensures that the questions asked are experimentally tractable and that the data generated directly addresses the core biological hypotheses, ultimately accelerating the validation process.
Related: CorrDyn provides deep expertise in biotech and life sciences data, helping organizations with data assessment to evaluate external platforms and build reliable data engineering foundations. We also assist with AI strategy to apply advanced computational methods for discovery and development.
Full Transcript
Jason: Welcome to Data in Biotech, a podcast from CorrDyn where we explore how companies leverage data to drive innovation in life sciences. Every two weeks we sit down with an expert from the world of biotechnology to understand how they use data science to solve technical challenges, streamline operations, and further innovation in their business. In this episode, we sit down with Jesper Ryge, Director of Computational Biology at Merck, to explore the evolving intersection of neuroscience, immunology, and AI-powered discovery in the pharmaceutical industry. He discusses how his team uses tools like knowledge graphs, single-cell data, and dashboards to support target discovery and pipeline stage projects. Jesper emphasizes Merck’s lean, partner-first approach, validating external platforms via pilots and applying generative AI for tasks like literature mining and patent analysis. Finally, he highlights the importance of interdisciplinary collaboration and being mindful of data’s limitations. Let’s get into it.
Ross Katz: Jesper Ryge, welcome to the Data in Biotech podcast.
Jesper Ryge: Thank you. Thanks for having me.
Ross Katz: Awesome. Well, just to kick us off, would you mind giving us an introduction to your background and what brought you into the field and what brought you here today?
Jesper Ryge: Yeah, sure. I originally started as a biophysicist back in the day. Trained in Copenhagen at the Niels Bohr Institute. I’m originally from Denmark. Before there was something called bioinformatics, I was very interested in both topics and at the time I started also being interested in genetics and neuroscience. The first exposure to experiments was actually electrophysiology, measuring electrical signals from neurons and trying to model that in a simple way. That was the physics at the time and I think is still valid. Try to build simple models that enable understanding of whatever you’re exploring, in this case, biology. I’ve really always appreciated that. Even though models are becoming more and more complex, I think we’re all struggling with how to interpret them and understand what is actually happening and how does it help us, for instance, understand biology and diseases. I’m trying still to find a balance, but that was my starting point. I continued with the neuroscience. I thought it was extremely exciting and interesting. I did a neuroscience PhD in Karolinska Institute. I was there for several years working on the spinal cord, since it’s actually a quite good model system for understanding how neural networks work. You can record electrical activity in vitro while it’s oscillating in a similar way as if it was sending signal out to the muscles when you’re moving or walking. You could actually access and record from the cells in vitro in this system and it was a good learning experience as well. We were trying to connect that at the time with the microarray technology that was emerging, molecular aspects, like looking at different cell types that we could label, for instance, with fluorescent markers, and then collect several of them, not single cells, but populations of cells, and then look at the molecular profile and try to relate that to electrical signals we were seeing. It was very challenging, and not as simple as we had thought. We also looked at a disease model, spasticity, which was in the tail model of the mouse, so it was actually more gentle than the classic models that were used at the time because it’s just the tail that is affected and develops this kind of spastic phenotype. We were looking at motor neurons specifically that project out to the muscles and we were hoping by doing microarrays that we would understand the molecular mechanisms and we had a good hypothesis that there was some calcium channels that were driving this. Again it turned out that biology is much more complicated and we did not find a simple explanation. It was probably the whole network that is changing, the balance between inhibition and excitation was probably shifting as well. But it triggered a lot of appetite for trying to bridge different types of data to understand diseases. I moved on to do a postdoc in EPFL in Switzerland as part of the Blue Brain Project that was actually building in silico models for the cortical column of rats, but based on data from the lab. I was fortunate enough to have a collaboration with one of the pioneers in single-cell transcriptomics, Sten Linnarsson at Karolinska Institute, because I bumped into him there before I left. We were looking then at mouse in the cortex, but also in the dopaminergic system and started working a little bit on neuromodulation and that led into the exploration of Parkinson’s disease and mechanisms there, but just characterizing all these different cell types were emerging and is still moving forward. I think all this single-cell data has really given us a much better understanding of the molecular profiles of different cell types and the big challenge in neuroscience is just how do we make sense of the billions of cells. Are there subtypes that are very characteristic and we can think of them as interacting as groups of cells or are they very individual and you can base it on morphology, like where do they project, and then you have the molecular profiles as well that is really helping in driving this understanding. When it comes to disease, maybe also looking at the dynamics. That was the next step, right? To see can we have reliable disease models where we can look across time? Because the big challenge in humans is that for brains, we don’t have access to tissue samples. You just get them postmortem and that I think makes the brain quite unique in terms of other tissues and other diseases. We still rely heavily on good models in the mouse and translating them to the human. I was working at the time on that but I also felt like I wanted to be closer to the translational aspect. I transitioned into the industry. First in a company that was working on clinical diagnostics on sequencing, so it was quite distant, but it was a good learning experience to come from academia where you have a very pragmatic approach to scripting and programming and it can be a bit messy, but whatever works, it’s fine. To a production system where it really has to be neat and you have to have version control and everything has to be structured and organized in a completely different way. That I have carried with me since then and that was a good learning experience. But I missed a little bit the scientific part. It was a very engineering approach. I continued into the pharma industry where I’m still located and that was really like in early discovery and translation of biomarkers at the time. I was fortunate to join a medium-sized company and that can also make a difference, right? Big companies tend to maybe be a little bit siloed. You get expertise in one topic but your colleagues might sit far away and you might not get insights to all the activities in the company. In a medium-sized company, you see all the different departments, even commercial, legal, and you get a better understanding of how all these things play together and how things are prioritized in a company. You might be interested and passionate about a certain disease and then you bring that to the management and they just say for commercial reasons this is just not interesting for us, right? You start learning, okay, there are different aspects that are important in prioritizing what direction you go in terms of the diseases and targets that you are continuing and move forward and explore further. Now I find myself in Germany, in Merck Germany. Maybe I should also clarify a distinction here that was a lot of confusion for myself as well when I joined. The original Merck is here in Germany. Was established more than 300 years ago. One of the family members created a branch in the US more than a hundred years ago and for historical reasons, that branched off into a separate entity. But with the same name. They insisted on keeping the same and now the complication is we entitled to call ourselves Merck everywhere apart from US and Canada. In US and Canada, if somebody remembers the Merck company, it is the US part and there we’re known as EMD Serono. That would be the reference. I always need to clarify this aspect. When I say Merck Germany or even if I say Merck, I’m referring to the German part. I was tasked to basically set up a computational biology team within the research units of neuroscience and immunology. I’ve been doing that for one and a half year now, and it’s super exciting. There’s so much happening within Merck, and also outside of Merck in biotech companies, developing data-driven AI tools for target discovery. It’s really amazing how fast the field is moving forward.
Ross Katz: For sure. Can you give us a little bit of an overview of some of the- you mentioned research science and immunology, what are some of the projects that your computational team works on?
Jesper Ryge: Yes, we’re tasked with basically supporting the pipeline projects until the clinical stage with whatever questions they have, which is typically mode of action, understanding better how is the drug actually exerting its effect in vitro and in model system. But also translation. We need to translate the mouse models and in vitro systems to human. I would say that we have been incredibly good in curing and treating animals and mice in particular and not so much always humans because the models don’t translate very well. That’s also something that we are challenged with and are focused on. I think the big chunk of what we are currently doing is actually early discovery. Implementing different data-driven methods to really improve the discovery because the challenge is that until a few years ago, it was actually an immunological research group, right? They were focused on autoimmune diseases and then they realized that there’s a big opportunity in neuroinflammation and neurodegenerative diseases by bringing in that knowledge but moving into the CNS space. That basically means that we have to build a new knowledge space both on the computational and on the lab side, and infrastructure. That has been the challenge for the last two years. Supporting the lab part by characterizing the in vitro system with omics. A lot of it is omics-based, right? The other part is maybe on the machine learning, NLP on the literature and trying to bring in the relevant datasets, analyze them, and bring insights for target discovery, target validation, and then also supporting the pipeline projects as they move forward. And also making these accessible to bench scientists, right? They don’t necessarily program in Python and R, so we need to have dashboards and other types of interfaces where they can go in and answer maybe more simple questions.
Ross Katz: There’s so much that I want to unpack there. I want to make sure that we get to translating mouse models into in vitro systems into humans, and the concept generation and the target discovery. I also want to make sure that we spend time on both the multi-omics and multimodal datasets and what you’re doing with NLP on the literature search. Let’s take concept generation and target discovery to start. Can you just explain in your world what concept generation means in practice and how you as a computational team support that?
Jesper Ryge: For us, concept generation is really proposing new targets for treatment of a disease. Traditionally, it was done by bench scientist experts looking at the literature and say, ‘Hey, there’s a phenomenal new paper that came out. There’s a clear causal correlation between this gene and this disease. We should pick up on this.’ Then they start validating those experimentally and moving forward with that. We still do that to some extent, but it’s clear that scientific literature is just exploded and nobody can keep up with all of that. We really need computational tools to basically leverage on existing knowledge. That’s one challenge, right? One of the key things of concept generation is if they still bring this proposal forward, this target is relevant for this disease, how can we capture all the existing knowledge in an easy way that doesn’t require this scientist to go and read like 500 papers? One thing that we’ve been doing is looking a lot at knowledge graphs and seeing if we can capture all this information in a knowledge graph where this gene is associated with this disease and this drug is targeting this protein and bridging all that together and then pulling out that information so we can say for this target, this is the knowledge space around it and this is the diseases that are associated with this. That is one attempt to making it easier and to harmonize the data retrieval to support and validate these proposals. I think the other aspect is also to have computational data-driven proposals, right? We use the data omics or also the knowledge graph because there we can maybe do link prediction and say are the patterns in the topology of the graph that suggest that not everything has been explored, right? These are incomplete. We’re still learning. We’re still adding information to this knowledge graph in a way. There might be certain genes that are important for a disease, but that link has not been established yet scientifically. But the graph might help you because they are connected in a certain way and then they say there’s a very strong likelihood that this gene is really involved in this disease and then we can also maybe experimentally try to validate that. That’s one method that we’re exploring when it comes to concept generation where you build on the existing knowledge.
Ross Katz: That makes a lot of sense. Can you provide a little bit of insight into how you integrate all of the new literature that’s coming out all of the time into the knowledge graph itself? Especially given the reproducibility issues that exist out there, there’s a data quality issue. You could be adding a lot of links to your knowledge graph and introducing a lot of noise, but also, expert time is really expensive. I’m just interested in what sort of processes have you developed for constructing the knowledge graphs so that all of these downstream tasks are achievable.
Jesper Ryge: Actually we’re cheating a little bit because we are very pragmatic. We are a small team and we need to really be quite lean. We can’t develop everything in-house. I think that’s an insight and a learning I think the pharma industry has done in the last few years that it doesn’t make sense to develop everything in-house. There is a lot of companies out there that has a lot of knowledge and expertise that create some of these infrastructures for us. I think that became clear, first of all, from the omics data that my team was spending too much time on bringing in the datasets, curating them, making them ready for analysis rather than just analyzing them and generating insights. We started looking for external providers that could allow us to just access that and externalize that in a way. Now we’re using QIAGEN’s OmicSoft product for that. Then in that process, I also started exploring their knowledge graph. They’ve been building up that resource for many years. Maybe some people are familiar with this IPA, Ingenuity Pathway Analysis engine, and in principle everything that’s under the hood is curated information they’ve been building over the years that at some point they decided why don’t we just license that, enable data scientist that don’t care about our user interface, that just want access to the knowledge graph to do with it whatever they want. That’s what we’ve been playing around with for the time being. They are in a way providing it to us and building it for us. We created some pilots around that as proof of concepts, also because for us it would be huge effort to first build a graph without knowing that there’s actually value in this type of analysis. What we often do is to create pilots and see if we can find these resources elsewhere and then once we’ve shown that there is value, we can then think about is it good enough that we can use it as it is or we would make an internal effort to improve on that now that we know that there’s value in going forward along this path. Currently, our knowledge graph is not- we’re not building it.
Ross Katz: That’s great. We had an episode on OmicSoft earlier that we can link to in the episode notes on this if people are interested in learning more about that. You’ve got this knowledge graph and you mentioned link prediction and also selecting which links in that knowledge graph you want to experimentally validate. How do you- either in the context of a specific problem you’re trying to tackle, disease area, a concept that you’re in the process of generating, how do you decide which of these aspects of the knowledge graph that you want to experimentally validate and layer your own internal knowledge on top of the knowledge graph that you have at your core?
Jesper Ryge: That’s a great question. Basically, from the link prediction and the existing knowledge in the graph, we can have different types of focus. We tried one where we said, let’s try and detect the most novel relationships between genes and a particular disease. Then with all this type of analysis, you typically end up with at least 50 or 100 genes that you need to rank and prioritize which one do we follow up experimentally, which ones do we think are valuable. I think that’s the other aspect that is very important for the pharma industry also when we interact with external partners that is also doing something like this or academia and they come with this phenomenal proposal: ‘This gene is really, really valuable for this disease.’ They say, ‘Yeah, but is it druggable?’ Is this something that is worth exploring for us, right? There’s additional features that you need to consider, right? There we say, ‘Okay, but what is known about this gene in addition?’ Or set of genes, right? Is it druggable? Do we have the structure? Is it amenable for a small molecule design? Transcription factors, for instance, have been very tricky, right? Maybe we would down-prioritize those. Safety issues. Is it expressed specifically in the cell types we’re interested in or is it expressed everywhere? Maybe it’s not a roadblock, but we have to think of mitigation strategies because we have to then target-ly deliver a drug to those cell types if it’s expressed everywhere not to get off-target effects. There’s a lot of types of information that we can then include to make a ranking and prioritize the gene targets and then we can with the biologist often dive into a handful of them. Then decide on which ones looks most promising and then start validating experimentally. It’s a long process. Even that part we’re also looking at leveraging knowledge graphs but knowledge that might be coming from databases. Open Targets is a great resource. Again, we’re not reinventing the wheel. We’re trying to grab as much as we can from either the public space or seeing if we can license things that can really move us forward fast. Then for the gaps that are left, we fill in those as much as we can.
Ross Katz: That makes a lot of sense. You mentioned with the biologist, you dive into a handful of them. I’m imagining that at that point when you have a handful and they’re in the hands of biologist, there’s custom experimental regimes that need to be constructed around a given concept in order to collect the data that either validates or invalidates the target that you’re trying to get to. Am I thinking about that right? Or is there more of a standardized process of just running all of the targets through a lab-automated process with certain assays that you’re doing over and over again, certain multi-omics panels that you’re doing over and over again and then that data alone is enough to get you a lot closer to the answer that you’re trying to get to?
Jesper Ryge: We can certainly, as much as that data exists, try to leverage on existing omics data, right? If we’re looking at a specific disease and there are transcriptomic data single-cell or bulk that we can look, ‘Is that gene upregulated or is that pathway upregulated?’ that would at least give additional support that that’s worth pursuing. I think on the in vitro side, the big challenge is to have the right models, right? Is it expressed in a neuron? Is it expressed in a microglia? If it’s in immune cells, but their effect on the CNS, it can be very challenging to combine all of these things into one in vitro system. That can be a selection criteria as well where we say, ‘We’re just focused on microglia partially because they seem like important players and partly because it’s feasible experimentally to validate targets that are in these.’ We can create aspects of the disease in vitro, right? Then we can see if we knock it down with CRISPR or silencing, does that have the effect that we’re looking for? Those are the type of things that you need to consider as well when it comes to feasibility. How easy is it to follow up on this? Do we have the right in vitro models? You can also have broader experimental system, right? Perturb-seq is a new technology that’s come out that we’re looking into, which is actually single-cell CRISPR. In one disease model, if you have a set of 50 genes, you can actually screen all of the 50 genes at the same time and then you can actually from the data itself figure out which cell had which gene knocked out, for instance, from the CRISPR panel, and then see if that brings the cell state back to normal. That also gives a good idea. You can in one experiment get a good understanding of which of these are good candidates and which are less good candidates. There is a lot of new powerful techniques that are emerging as well. I think this has even been bridged so you can do it in animal models. You have spatial resolution with spatial transcriptomics. This is very exciting, that we can accelerate the discovery process with these technologies.
Ross Katz: That makes a lot of sense. You mentioned earlier that you had the opportunity to work with one of the early pioneers of single-cell transcriptomics and now you’re talking about spatial transcriptomics data. My understanding is that these approaches are much more computationally heavy, but then also richer in terms of the insights that they can give you. I’m curious, what have you seen in terms of your ability to generate concepts or validate targets with these new approaches? How has it changed your workflow?
Jesper Ryge: For now, it hasn’t, but we are very curious about exploring these technologies. I think it might also be depending on the disease, but for the brain that is so complex and we are looking at immune cells that are innovating this organ in disease, it would be really crucial to have good quality spatial transcriptomics so you can see, okay, what immune cells are innovating the brain and what is the microenvironment around these? Which cells is it interacting with and what is also happening in these neighboring cells? That would be extremely insightful, I think, in terms of understanding disease mechanisms in neurodegenerative disorders, for instance. I think the platforms have evolved rapidly, right? But the spotted ones where you had overlap of multiple cells that you had to deconvolve was a bit challenging to use. I don’t know how many insights it has generated for us. But now that you get image-based spatial transcriptomics platforms coming out where you have true single-cell resolution but, okay, you might be limited by panels, but the panels can be up to several thousand genes. You can maybe leverage on single-cell data to make sure that you get the genes you’re interested in captured in the spatial transcriptomics. We are actually exploring collaborations that we can generate data, right? Because there’s very little out there. Fortunately the tradition in academia has been that if you publish an omics study, you share the data. But since these technologies are so new, there’s not a lot out there. We have to generate this data ourselves. I think the big promise was that there’s a lot of biobanks with, for instance, postmortem brain samples. If we can use this technology on these samples, I think that could generate a lot of insights in terms of disease mechanisms. We are of course very, very interested in that. But it is early days, and we will have to see how many insights we can generate with this, but I’m very excited and I’m very optimistic about what we can learn from this.
Ross Katz: When you look at partnerships to either acquire data directly from someone who’s already generated for you or partnerships to generate data on your behalf, I’m curious, how do you- you mentioned the pilots earlier that you do when you’re onboarding external tools like OmicSoft and seeing whether the value is there or whether you can create that value more effectively in-house or whether onboarding the external tool is there. I’m just interested in how you think through the process of onboarding these new partners or of bringing in these new datasets, and how those pilots tend to work.
Jesper Ryge: That’s also a great question. I think we have transitioned a little bit from a situation where we would focus extensively on a small handful of diseases and we would just scout everything that was out there in terms of omics data and just ingest it and it was manageable. I think now we are pivoting quite fast between different diseases. I think that’s a typical aspect of immunological assets, right? That you target a B-cell or a T-cell or a mechanism that can be relevant for multiple diseases. The competitive landscape changes over the years, so what was relevant today might not be relevant five years from now and then when it goes into clinical development, there’s a push to say, ‘You might have think you wanted to go into this disease, but we’re not sure that’s the right commercial potential. You need to go into maybe X, Y, Z,’ and then you need to ingest that data to see, ‘Okay, but is that supporting basically the mechanism of action for this?’ That is also reflected early on, that we are exploring more and more diseases and we need to bring that data in quite fast and quite dynamically. We have had some efforts in scouting and identifying the datasets, but then we’ve looked a little bit also of course at which companies can then provide that service for us. That is actually not as easy as it sounds. Because there’s a lot of companies out there and then you Google, you try to look at the homepage- but if you don’t have your network in place or you haven’t been exposed to that. Fortunately in a big company, we actually have the business development that supports these activities. I’m fortunate to have a colleague that actually helps me scouting, so if I define and say, ‘We have this challenge. We clearly have gaps. We have looked for the datasets, they are not there in this indication. We want to generate them. Can you help me find the right partner for us?’ They will actually come with some proposals. They will look for the companies and then we will reach out to them and start a conversation to see if they can help us with that. The same for the curation, right? Like I mentioned earlier, who’s the right company for that? There’s also a lot of different players. Then you have to be very clear in what your expectations are, what is it that you need from them. Then you can try and engage in these partnerships in this way. But it can be a challenge to find the right partner, let’s put it like that.
Ross Katz: I can imagine that when you talk about expanding the scope of the data that you’re gathering and the diseases that you are targeting or might be targeting in the future and building the data platform for all the different use cases computationally across the organization and then you’re thinking about onboarding a new partner, you’ve got to pick the right problem and the right narrow dataset to initially trust that partner to bring the data in and then you get your hands on the data inside of a specific use case so that you could understand how would this apply to the rest of the scope of the work that you’re doing. Am I thinking about that right or how would you update that?
Jesper Ryge: Sometimes you can get a sample datasets in another disease, but it’s the same method and you can validate that, ‘Yeah, this is good quality.’ It can also be matter of curation and metadata- and again examples of datasets that they have generated or curated can be valuable, so that’s one starting point. Then like you also mentioned, we can do pilots. Just to say, ‘Do they deliver what they promise?’ right? Then you can stage the engagement as well and say there’s certain milestones and we basically do an initial validation and then once you have delivered that and we have shown that it satisfies our criteria, then we move on to the next stage. You can build that in as well.
Ross Katz: That makes a lot of sense. I want to zoom out a little bit and talk more about the computational workflow as a whole. How do you think about the role of a computational team in biotech research? Maybe it’s worth talking about how your computational team complements the rest of the research that’s happening in the organization and what is the interplay like between your team and people who are in the wet lab or other stakeholder groups that touch your work?
Jesper Ryge: I think there’s two aspects. There’s supporting questions coming from the different teams, using data, right? Typically omics data or like I said, text mining from the scientific literature. Having a more supportive role, and simple questions sometimes can require quite a lot of work to onboard the right data or do the right analysis. The other aspect is actually driving projects from the computational team to generate insights. Especially in the target space, right? Think about new ways of integrating datasets. Like I mentioned the knowledge graph link prediction, something that would not have come out of maybe a request from a lab person. We do both things. That’s the challenge a little bit to then build up the data assets that can support these things, right? Like you say, we build up data assets omics in different indications, but that also enables us to look more globally at some point and say, ‘But what is the commonalities between diseases? Can we actually in a smart way find targets that are maybe related to an immunological process that is shared across several diseases?’ I think there’s also a benefit in that sense from the computational point of view to look more globally and more holistically at these processes, and not just answer individual questions that are coming out of these research teams. These is the two aspects that we are dealing with.
Ross Katz: That makes a lot of sense. You’ve got the reactive question-answering aspect of what you all are doing and you’ve got people who are studying very deeply a certain disease category or certain therapeutic approaches and so you’re answering questions that support them. But then also- another thing I heard was that you’re building the platform to make answering those questions relatively easy and making it so that you have the data in hand to answer the most important scientific questions across the organization. But then once you’ve built the platform for answering these reactive requests across the organization, you also have the opportunity to do that proactive exploration of the datasets that you’ve developed to establish these commonalities across different diseases, for example, the immunological process that you just shared. Am I thinking about that right and if so, how do you prioritize between the different aspects that your team is doing or allocate resources across those different aspects?
Jesper Ryge: It’s not so difficult, really. We have these two accountabilities, in a way, that’s pretty well defined. But of course, they can change over time. But we have to support the teams in answering these questions. I think the other aspect is also just to promote the digital culture within the organization. Some bench scientists are very excited about all this computational stuff and are really happy if they, for instance, get access to a dashboard. They might not start programming, but they’re very eager to dive in if they have tools that are a little bit more intuitive to use. We also focused together with other teams in the organization to build these self-service, as we call them, platforms and dashboards. Then you have the people that are more slow to adapt these things. But then as they see examples in other projects where we are generating insights and answering questions that are similar to the ones they’re faced with with a computational approach or an omics data that was out there, they also maybe get inspired and then reach out to us and say, ‘You know what? I have a similar question. Can you help me with that?’ That’s a way also to improve a little bit the impact we have on the different projects. Then the discovery part is the other mandate, right? That we try to really drive new computational strategies. In a way, I see it a little bit as two sides of the same coin, right? The experimental side and the computational side. In a way, we’re just generating datasets that are so big that you can’t just do a T-test like you used to do to see what is significant or is there an effect, right? You need to have a bigger toolbox of computational tools to deal with those datasets. But in principle, it’s still an experiment that is designed to address a certain question. There is the experimental biology aspects and then there is just the computational aspects, but it’s still biology in a way. We are not doing in silico modeling or something really purely theoretical in that sense. I think that’s the mindset that’s also changing a little bit. I think maybe younger generations will not think so much about because they probably grow up with some sort of skills in programming and will not find that so challenging even if they’re in a lab setting. A lot of the work is really like do we generate different types of phenotypic screens? It’s an experiment that you have a disease model, and then you screen for, for instance, the drugs that bring that model back to normal. You can do that for thousands, thousands of compounds. It just becomes a very massive dataset. There’s really a close relationship between the computational and the experimental side. I don’t see it as so separate. Then we’re just supporting these kind of experiments because it can seem intimidating for some biologists to engage in these things and say, ‘I don’t have the skill set.’ But once we have a team that I know is dedicated to this, all of a sudden becomes a very exciting collaboration in a way, right? Then you bring people to the same table that have these skills that can basically push these projects forward. I think that’s super exciting. We have a few of those also starting up in the team.
Ross Katz: That’s really interesting and it makes a lot of sense that as the assays are generating more and more data, even if the experiment was designed in order to generate the data and that experimental context is critical and also follows a logic path that came out of the scientist’s brain, the computational processing of that data is still going to require somebody with a computational mind or computational skill set in order to get the resulting insights out of there. That partnership seems like these two sides are being driven together or attracted together as much as they’re being forced to work together. One of the things that you mentioned earlier that I’m glad you brought up again is the idea of self-service platforms and dashboards for scientific users. I think you mentioned maybe this is related to the knowledge graph and the link prediction and stuff like that, but would just be really interested in what are some of the- if there are any success stories you have with self-service platforms or dashboards or things you have in the wild right now that are really useful to your end users.
Jesper Ryge: I think a lot of them are conceptually rather simple, right? But if you have a lot of omics data, it might not be easily accessible to non-computational people and you just have a dashboard where they can look up, ‘Is my new target of interest expressed in this disease? Is it upregulated?’ I found an experiment in a publication, very recent one, and the data looks interesting but they don’t look at my favorite gene. Can we bring that data in? Then we can put it basically in these self-service platforms and make it available for them. Then they can at least do some of this simple analysis on their own and I think that’s the biggest success stories, where they can go in and independently explore and answer questions. Then the more complicated stuff we might not build into a self-service platform, we just do that together. I think also we have a team that is building up an internal knowledge graph. I think that’s also becoming more and more powerful because you can really in a very nice way integrate different data resources. It can be a licensed platform for competitive intelligence. Then it can be some open resources like Open Targets. Since you can integrate it all together, for instance, now we’re working on a dashboard that just summarizes in a more harmonized way all the information that’s relevant for evaluating an early target, right? Where is it expressed at a single-cellular resolution? What is the competitive landscape? Is there a lot of things happening there already? Are there safety concerns? What is the effect when you knock it out in an animal? Are there CRISPR screens that have looked at this? A lot of different aspects. Not everybody was aware of all the platforms that are licensed internally or exist for free externally. People were looking in different places providing different types of information or not simply being aware of some of them. Now having one access point that integrates everything is becoming increasingly valuable, I think, across the whole organization, not just for us.
Ross Katz: That makes a lot of sense. Are there any applications- you mentioned summarizing in a more harmonized way, so obviously that triggers the idea of generative AI or bringing LLMs into the mix. Are there any interesting ways that you’re using AI or LLMs internally at this moment?
Jesper Ryge: Well, there’s just the standard use, right? We have our own internal platform, it’s called Merck- like ChatGPT version, right? We are building agents around that. I think that’s quite interesting. It can be really simple stuff like it can be related to non-scientific aspects, HR or other things. But it can also be like people were looking into patents, right? That these patent applications are hard to read and extract information from. If you can just get an LLM to digest that and spit out the results, that would be very valuable. This is not really fully developed, but that’s a use case where we would see the value, right? Then there’s just the day-to-day stuff where people are just asking instead of going to PubMed, they are using ChatGPT instead and saying, ‘What’s known about this disease? Is this target relevant?’ But you have a conversation with it to challenge it a little bit. It’s not giving you a final answer. I think also the other learning is that it’s not accurate all the time. Sometimes it hallucinates and you have to challenge it a little bit to see. You might only become aware of that if you’re an expert in the field you’re asking it about and you see the mistakes. But then if you go into another disease area where you’re less knowledgeable, it might not be as transparent. That’s also an insight and a cultural change. But people are embracing it across the organization and I really see that we have projects where you then try to build agents that can interact with each other because it can’t just magically solve a big question. But you can build an agent that’s a specialist in omics data. Then you can build one that’s a specialist in something else and then they can interact with each other and cross-check each other and then finally produce an answer. Of course, it’s extremely powerful for non-tech users as well as the data science community to use these tools to integrate all these diverse datasets to bring you insights. Then also for programming. I think we are getting more efficient in a way, right? But there’s also caveats. You can ask it to review certain parts of your script and improve it or make it more memory-efficient because you did something quick and dirty and now you have a bigger datasets and everything is crashing and then they say, ‘Okay, can you optimize it? I’m not really sure how to do it.’ It also works quite well for these things, right? I think on the programming side, I also see a big benefit to using these tools.
Ross Katz: That makes a lot of sense. As we come toward the end of our conversation, I just want to zoom out a little further and ask you some questions about the future and what things look like. Are there any emerging technologies that you’re really excited about transforming your work in computational biology? That might be on the assay side or the lab automation side or the data processing side or AI, any of the things that are out there.
Jesper Ryge: I think we’ve touched on them, but spatial transcriptomics for me is super exciting and I think will generate a lot of insights into disease mechanisms. Multi-omics aspects is just crazy how much data you can generate. I think the big challenge is that like you said earlier, you generate a huge amount of data from a very, very tiny area of the human brain that is quite big. If you want to cover a bigger part of the brain, it’s still challenging. In pharma, typically you want to have many patients, right? You want to look at a hundred patients and a hundred controls. That’s a big challenge with these type of technologies. There’s still a long road ahead, but I’m very excited about these new developments. The other thing is generative AI, right? It’s just again crazy how fast that is going and I clearly see that as a powerful tool in the future to bring different assets data types together and generate new insights and support a lot of the activities we have on the experimental side and on the computational side as well.
Ross Katz: That makes a lot of sense. Data integration and harmonization is one of the places where generative AI is already demonstrated the capacity to do what we need it to do, so it makes a lot of sense that that just has the capacity to make it easier to get that full picture of all of the data that you’ve collected and how the pieces fit together. Are there any questions that you think people in pharma R&D should be asking about their data, but probably aren’t right now?
Jesper Ryge: Hm. I would maybe just caution that you have to be aware of the limits of your data. What can it not answer and what is the quality of your data? There was a lot of hype and for good reasons about single-cell, right? But there are technical limitations and sometimes a computational scientist might not be aware of those because they don’t have hands-on experience. There might be certain cell types that just don’t get a signal for various reasons. We had an example in psoriasis where in bulk samples, there was a clear signal from neutrophils, but we were not seeing them in the single-cell data because basically the way the platform works is that they get captured in a droplet, they get lysed, and then because they have these internal lysosomal organs that are phagocytic, they will basically digest themselves and the RNA and you will not get any signal from them. You just don’t see those cells appearing in single-cell. Then you make conclusions about disease mechanisms and all this kind of stuff and you’re missing an important player. I think that you need to build in sanity checks as much as you can. Generate bulk and single-cell and see if, maybe make pseudo-bulk from the single-cell and does it correlate with the bulk that you’re seeing? Are there signals that are not appearing in the single-cell? This is just one example, right? But there’s many of these where you try to build in some sanity checks and be aware of the limitations. There’s also a lot of interpretations on single-cell in terms of proportions of cells. They might also have biases. You have to be aware of these because you often say, ‘This cell population was increased and this was decreased.’ I think it’s also been shown that they might be just more fragile and they might fluctuate for various reasons, but they can be confounding factors that are hard to decompose in these datasets. But it’s good to be aware of them. Every technology, I guess, has this kind of problems and it’s important that the data scientist is aware of these when they make conclusions, so they don’t end up going down the wrong path.
Ross Katz: I think that makes a lot of sense and I guess it’s difficult to be an expert in absolutely everything. I’m wondering how you or your team think about building up the healthy skepticism and the internal knowledge about the inferential boundaries of these different datasets and what some of the assumptions are and how they can be used and what the sanity checks look like. Is it just a collective conversation or how do you approach that kind of problem?
Jesper Ryge: It depends a little bit on the expertise you have in your team. I think for our team, a lot of people emerge from labs, they have lab experience, they trained in the lab, and at some point they had an interest in computational aspects so they were exposed to omics data that they were generating and wanted to analyze themselves. They have a good understanding of the experimental aspects as well that might not always be the case. I think best is interdisciplinary teams. That’s by far if you don’t have that in one person, bring these people together at the same table and have a discussion around, ‘What is the pitfalls experimentally? What can you do computationally?’ and then try to understand what you can and cannot do, right? What is the question ultimately that you want to answer and how do we get to that? What are the limitations in whatever you’re proposing and is it ultimately going to give you that answer or do we need to do it in a different way? What people need to be part of that conversation to make sure that we do it the right way.
Ross Katz: Awesome. Well, it’s been a real pleasure talking to you. Are there any final thoughts you’d like to share before we let you go?
Jesper Ryge: No, it’s been a really nice conversation. I really appreciate being here.
Ross Katz: If people want to get in touch with you or follow your work, what’s the best way for them to do that?
Jesper Ryge: Well, they can always find me on LinkedIn. I find that that’s the one thing that stays, no matter where you move. I would just suggest that people find me there. I don’t reject any connections and you can message me, so just look me up there.
Ross Katz: Well, Jesper, thank you so much for joining today. It’s been a real pleasure talking to you and look forward to connecting down the line.
Jesper Ryge: Thanks a lot for having me.
Jason: And that’s it for this episode of Data in Biotech. If you enjoyed the episode, please subscribe, rate, or leave a review in your podcast platform of choice. See you next time.






