Skip to content
Hugo Shi & Ilya Burkov — Scaling AI Infra in Biotech with Saturn Cloud and Nebius
Data in BiotechEpisode 55

Scaling AI Infra in Biotech with Saturn Cloud and Nebius

Hugo Shi of Saturn Cloud discusses how specialized AI infrastructure solves compliance, scale, and security challenges for modern biotech teams.

52:17Full transcript below
HS

Hugo Shi

CTO and Founder at Saturn Cloud

IB

Ilya Burkov

Global Head of Healthcare and Life Science Growth at Nebius

Overview

Biotech companies face immense pressure to accelerate drug discovery and research, but their AI workloads are rapidly growing in scale and complexity. This expansion frequently hits a wall with infrastructure management, GPU scarcity, and spiraling costs on general-purpose cloud platforms. Scientific teams find themselves dedicating valuable time to managing Kubernetes clusters, optimizing hardware, and ensuring compliance, diverting focus from their core mission of innovation.

In this episode, host Ross Katz speaks with Hugo Shi, Founder & CTO at Saturn Cloud, and Dr. Ilya Burkov, Global Head of Healthcare and Life Science Growth at Nebius. They offer perspectives on how specialized platforms overcome these barriers. They outline the unique demands of biotech AI workloads—from stringent security requirements and integrating complex academic code to the critical need for efficient, accessible GPU compute. The discussion clarifies how a strategic approach to AI infrastructure can significantly reduce operational overhead and accelerate scientific breakthroughs.

Listeners will learn how specialized cloud solutions provide not just compute power, but also managed environments tailored for biotech’s specific challenges, delivering substantial cost savings and enabling teams to concentrate on impactful research rather than infrastructure.

Key Takeaways

Hyperscaler GPUs are up to 70% more expensive for AI workloads.

General cloud providers like AWS, GCP, and Azure present a significantly higher cost for GPU-intensive AI training and inference. Specialized platforms, designed specifically for AI and high-performance compute, typically offer between 60% and 70% cost savings on GPU usage. This price difference stems from the dedicated focus and optimized infrastructure these specialized providers offer, making them a more cost-effective choice for sustained AI development.

Biotech’s unique AI infrastructure demands go beyond raw compute.

Biotech AI workloads require more than just powerful GPUs. They necessitate tight security and compliance (HIPAA, GDPR, ISO), the flexibility to integrate diverse, often complex academic code, and the ability to publish interactive dashboards (Shiny, Streamlit) for collaboration with non-coding biologists. Infrastructure choices must support these specific requirements without compromising scientific agility or data integrity.

Offloading AI infrastructure management is critical for innovation.

Managing AI infrastructure—from hardware maintenance and Kubernetes orchestration to MLOps optimization—is an intricate, time-consuming task. Biotech teams that try to handle this internally divert resources from scientific discovery. Specialized platforms abstract this complexity, allowing computational biologists and data scientists to focus on experiments and model development, thereby accelerating research cycles and reducing time-to-market for new therapies.

Specialized cloud platforms deliver superior performance and operational models.

Unlike general hyperscalers, dedicated AI cloud providers design, own, and operate data centers optimized for AI training and inference. This includes tight integration with advanced NVIDIA hardware and high-bandwidth networking, ensuring ultra-low latency and efficient memory optimization for large models. These platforms also offer flexible operational models, from self-service GPU access to white-glove MLOps support, adapting to diverse team needs.

Related: CorrDyn helps organizations with data engineering and machine learning implementations. For biotech companies, we also offer specific guidance on realizing data value and data cost optimization.

Full Transcript

Jason: Hey everyone, this is Jason, producer of Data in Biotech. I’m excited to announce that this podcast is now an AAPS media partner. Join us on-site at PharmSci 360, November 9th through to November 12th in vibrant San Antonio, Texas. We’ll be interviewing leading scientists about the data trends that span the pharmaceutical development and manufacturing spectrum. Data in Biotech thanks AAPS for this opportunity to share emerging science with the global pharmaceutical community.

Hugo Shi: Like, if you are running GPU workloads on AWS, GCP, and Azure, you are basically fine with lighting money on fire. So if you are a platform engineer and infrastructure engineer and you have time, you should be focusing on the remaining 10% that doesn’t exist on off-the-shelf tools that will actually provide differentiation and value to your business. Don’t solve the problems that are easy for people like me to solve, solve the problems that I can’t solve.

Jason: Welcome to Data in Biotech, a podcast from CorrDyn where we explore how companies leverage data to drive innovation in life sciences. Every two weeks we sit down with an expert from the world of biotechnology to understand how they’re using data science to solve technical challenges, streamline operations, and further innovation in their business. Here we go.

Ross Katz: Hugo Shi and Ilya Burkov, welcome to the Data in Biotech podcast.

Ilya Burkov: Thanks for having us.

Hugo Shi: Thanks for having us.

Ross Katz: Awesome. Well, since we have two of you today, just to get us started, Ilya, would you mind kicking us off by introducing us to your background, what brought you here, and a little bit about your organization?

Ilya Burkov: Yeah, sure. My name is Dr. Ilya Burkov. I’m the global head of healthcare and life science growth here at Nebius. I’m based in the UK, but leading this vertical for about 15 months now. I joined with experience in the healthcare and life science sector, come in as a subject matter expert in the field. Before Nebius I spent about three years at AWS cloud services, which was a really great experience into general cloud. Earlier on in my career, I worked in the NHS, in orthopedics at Addenbrooke’s Hospital in Cambridge. I worked on biomarker identification of early disease onset, in particularly looking at osteoarthritis and osteoporosis, looking at regenerative medicine and dabbled in some biomedical engineering as well. What I love the most about my work today is bridging the world of biology, data, and the cutting edge AI infrastructure that comes with it. I love to help the biotech and healthcare teams accelerate the discovery that they’re working on, streamline R&D processes that they have in place and bring new therapies to the patients a lot faster. The company that I work for, Nebius, it’s a next generation cloud platform. It’s built specifically for AI and high performance compute. We provide the infrastructure that helps a lot of the organizations train, deploy, and scale AI models much more efficiently, including those that are driving breakthroughs in biotechnology, drug development, genomics, and beyond.

Ross Katz: Awesome. And Hugo, if you wouldn’t mind doing the same thing, giving everyone an introduction to your background and then a little bit about your work.

Hugo Shi: Yeah, my name’s Hugo Shi. I was once a signal processing person in the field of medical imaging, working on statistical medical image reconstruction. I was a MATLAB programmer for a long time. Since then, I’ve transitioned into a bunch of things. I started off leaving grad school and working in quantitative finance, quantitative trading, and that’s where I fell in love with Python. I’m also one of the founders of Anaconda, but in a much more minor capacity compared to Peter and Travis. I am the founder of Saturn Cloud. At Saturn Cloud our goal is, similar to what Ilya said, to remove barriers and give data scientists, AI engineering teams and machine learning teams the freedom to get their work done without having to worry about managing infrastructure or Kubernetes, things like that.

Ross Katz: Awesome. I can see the overlap between the two organizations in terms of you’re trying to do something really complex that requires high performance infrastructure, that requires optimization for that infrastructure, but there are a lot of complications that come with that. What are some of the challenges that biotech and healthcare data teams and companies face when they’re thinking about expanding the AI workloads that they’re trying to train or infer, the different use cases that they’re driving at that require this infrastructure?

Ilya Burkov: Yeah, Ross. AI workloads in biotech are exploding in size and complexity. Models are larger, the datasets are multimodal and pipelines span genomics, imaging, text, there’s a plethora of information there. A lot of organizations have very valuable data in both on-prem and also a mix of cloud and other solutions. A lot of these things are not easy to move. Flexible and high performance infrastructure is needed. It supports this hybrid environment, but also this is a gap that Nebius is addressing. The biggest pain points that we see are access, efficiency, and scale. GPU scarcity is a thing, and Hugo mentioned it a couple times with the hyperscalers that they faced. Many biotech teams can’t access reliable infrastructure to the compute that they need, especially for larger AI models, for example in protein design or image analytics and so on. We as a company don’t discriminate against what their workload is and what the size of the company is. Even a smaller group can have access to the Blackwell generation of GPUs. Imagine going to a hyperscaler and asking for a couple nodes of B300s. You can imagine their reaction. I think Hugo can confirm. Second of all, cost and utilization. Even when companies secure GPUs, they’re often run inefficiently. There’s idle times, there’s data transfer bottlenecks, fragmented workflows that eat into the budgets, that we can help manage, offset those costs and so on. The final thing that I mentioned was complexity. Managing the entire AI infrastructure across these environments, it takes teams away from the science. It takes them away from what they’re really trying to do. What we’re seeing is that companies want high performance compute, but they need it to be as simple and as elastic as the cloud. That’s what we do best.

Ross Katz: Hugo, would love to hear for you similarly from the data team facing perspective, what are the challenges that you see and how does Saturn Cloud intervene to abstract those away?

Hugo Shi: Yeah. I’ll say the first thing, this might be a little boring, but the first thing that they face is tight security, because the data is very sensitive and the field is highly regulated. If you’re working at a company like that and you just hand AWS keys to computational biologists, a lot of bad things can happen. We focused a great deal on how to set up the right guardrails, so that both the IT security team is happy and they can be comfortable that the data science, the computational biologists or the titles — there’s so many titles in that world right now — can have freedom to do their work without it being possible for them to do something dangerous. I will also add that I really think most of the products that are out there don’t have the necessary support for the kind of compliance that these organizations need. Oftentimes we’re asked about integrating with customers’ security stacks, for example, CrowdStrike, Falcon or Orca Security or SentinelOne, and those are things that we have built into or we’ve made possible to build into our stack. The other thing that I would add is two other things that are unique about biotech is there’s a lot of very interesting code coming out of academia. That code is often very difficult to install and has a lot of strange requirements. Or maybe strange is not the right word, but very unique requirements. It’s really important if you are going to build a platform or a stack or something to make it very open so that you can support all these different requirements, because the last thing you want to do is say, okay well this is our special proprietary runtime, I want to install this new project that just came out of MIT, oh I can’t because I don’t have access to something something something. That kind of stuff. Finally, a lot of the biotech teams that we work with also have to, or a lot of the computational biologists that we work with also have to collaborate with people who are technical in the sense that they know biology very well, but they’re not technical in the sense that they don’t code. For that we find it very important for them to be able to easily publish shiny or Streamlit dashboards to help those teams consume the work that they’ve done.

Ross Katz: That’s really interesting. What I sense is that on both ends you’ve got the need for this infrastructure, the need to understand how the infrastructure works and then optimize the workloads that you’re working with on that infrastructure. Then you also have these unique computational approaches that are coming from a variety of different areas that have probably infrequent or never before been optimized for that infrastructure and then also you have the tight security and the highly regulatory landscape that requires a lot of control from the IT perspective. The goal from end to end is to enable the data teams to do what they need to do without having to think about the regulatory stuff, without having to think about the optimization of the code for the infrastructure that they’re working on. Am I thinking about that right, or where would you all add to my overview?

Hugo Shi: I think that’s absolutely correct. If you just think about the stack that we sit on today and how complex it is. Nvidia’s putting out new GPUs all the time. When they do, those have different hardware requirements. This is out of my expertise but those racks are really hard to put together correctly and to manage. Managing those data centers is an incredibly complex and difficult task. I’ve heard of so many failure conditions. Someone at a conference recently told me about a situation where a customer was using — this is not a Nebius thing, but this is just to give you the scope of the problem — someone was using liquid cooling and they decided to use water instead of a mix of water and propylene glycol. What happened was that algae built up inside the tubes. Those tubes at the narrowest points are 50 microns. They had to take everything offline, filter the water to remove all particles that are bigger than 25 microns so that two can’t block the pipe and then turn the system back on. That’s just the hardware stuff which Nebius handles which is amazingly complex. In some ways I get the easy job because it’s okay, well we deal with Kubernetes. But Kubernetes — that field is moving so fast too. On top of that, how fast is AI moving? The model space is moving so fast. If you think about all of that stuff and how complex it is, one, it’s amazing that people get anything done at all. But two, you have to offload that. It’s crazy to try to get one person or one team to do that. You have to offload it. The smaller the surface area you can put on the AI ML team, the better off you are. Without restricting their freedom, because that’s the other thing. There’s usually a tradeoff between, well I can create a really safe environment, it’s just that you can only do this one specific thing. That’s not good either. You want to give them a situation where they don’t have to think about all this stuff below the stack, but they have still the freedom to innovate. Because if they’re not innovating and not living on the cutting edge then there’s no point. They’re not going to be successful.

Ilya Burkov: Yeah, 100%. I fully agree with what Hugo said. Ross, going back to what you said about compliance and things like that. In life science, security and compliance are really non-negotiable. Patient data, proprietary research, it can’t be exposed. Which is why when we’re building Nebius architecture and the hardware, we built it with that in mind. We support the compliance with frameworks like HIPAA, GDPR, various ISOs that are required and other life science regulations. That means that a lot of the biotech teams can run AI workloads confidently. They know that their sensitive patient information and intellectual property remain fully protected. It’s really simple. Nebius is secure by design and by default. We’ve got a trust center that you can have all the access to. You have all of the compliances that we adhere to. You have the PDF certificates that go into in-depth information on this. That’s what builds trust between the customers that rely on the infrastructure and also rely on where the data resides. Without that, it’s very difficult to progress and to grow and to scale at this level.

Ross Katz: That’s really interesting. What I want to do is ground the conversation from the perspective of a customer. To Hugo’s point just a minute ago, the AI landscape is moving really quickly, just in the larger ecosystem and then also inside of the biotech space specifically. We just had an episode about the federated training of OpenFold 3. OpenFold 3’s preparing to be released, it’s going to be the open source version of AlphaFold 3. I’m thinking about this from the perspective of the biotech company. There’s two situations where you might want to use Saturn Cloud and Nebius or one of the other. You’ve got this new model that’s just come out, you have an existing workload you want to run it against. You can’t get the GPUs you need, or as a scientist, or as an organization you don’t have the devops capability in-house to be able to get up and running to where you need to be so that you can just solve the problem you’re trying to solve from a scientific perspective. In that example, how would you all recommend they utilize — let’s start with you Hugo and then you Ilya — Saturn Cloud and Nebius to enable this workload to come to fruition.

Hugo Shi: I think it would be pretty simple. You sign up for a Nebius account. Maybe you have to talk to support to increase quotas or something, but probably not. Saturn Cloud we’re actually working right now to get into the Nebius application, it’s not a marketplace but it’s managed Kubernetes applications for Kubernetes I believe is what it’s called. Right now it’s just a simple Helm chart. Spin up the Nebius managed Kubernetes cluster, install the Saturn Cloud Helm chart and then you’re all set. I guess the next — once Saturn Cloud’s installed, the next step would be to configure a Docker image that has the necessary dependencies to run OpenFold. If one already exists then great. That’s it. Then you launch it. If you’re not on Nebius, I would say the first step is to get on Nebius. Actually, one comment I will say that’s a little inflammatory is that if you are running GPU workloads on AWS, GCP and Azure, you are basically fine with lighting money on fire. That’s another thing — you should definitely onboard a neo-cloud like Nebius in addition to your hyperscaler usage.

Ilya Burkov: Yeah, in terms of the infrastructure, it’s ready to handle all of the workloads of OpenFold and beyond. It’s the case of scaling up to whatever the workloads are. If they want to start off with a couple of nodes and they don’t need us, they can do a self-service account and just get up and running themselves. It’s a couple of clicks away. If they want more of a white glove solution, they can approach us directly and say this is what we’re trying to solve, can we do a POC, can we have that environment, can we incorporate Saturn Cloud solution on top of that and so on. We’re happy to do that as well. We’re happy to be as hands-on or hands-off as needed. Sometimes people just want access to the GPUs and they say leave us alone, we know what we’re doing, then they’re happy with that. Other times when they’re working with a new model like the one that you mentioned, the OpenFold, they might want a bit of hands-on experience, they might want some time with the DevOps team, ML ops team and they don’t have to spend time fighting on infrastructure and can actually focus on the actual science that they’re doing.

Ross Katz: That’s the white glove service that you’re providing — the ML ops, the DevOps, the infrastructure optimization side — if people want that kind of service, that’s what you’re providing on top.

Ilya Burkov: Yes, exactly. With partners like Saturn Cloud that can be achieved even quicker and even faster and they have that additional layer of support and additional level of experienced people looking after them, helping them to remove a lot of the bottlenecks that plague the teams.

Ross Katz: That makes a lot of sense. Hugo, back to what we were just talking about, if one of your customers was already on Saturn Cloud and utilizing AWS, GCP, Azure, Oracle, and they want to try out the workload on Nebius instead, what does that process look like for them?

Hugo Shi: Yeah, we actually just did this with the flagship pharma company that I just mentioned. They were primarily working on AWS. They were not happy with the GPU pricing and availability. They created a Nebius account and then we installed Saturn Cloud into it and then they’re good to go. I guess I should add that we typically, similar to what Ilya said about the white glove, have two operational models. The one we prefer and the one most of our customers prefer is the white glove or fully managed approach, which is they give us some sort of access to their cloud account and then we set up the Kubernetes cluster and we install the software and manage it. But for people who want to be a little more hands-off, they just install the Helm chart.

Ilya Burkov: One of the things that Hugo mentioned is that neo-cloud approach. A lot of people ask us what is the difference and who we are as a company. I would position ourselves somewhere in the middle between a hyperscaler and a neo-cloud because we offer the flexibility of hyperscalers, but the performance of the bare metal GPU neo-clouds that are out there. It’s not just access to the GPUs, it’s access to the full stack solution. We design, we own and we operate our data centers specifically for AI training and inference. We have a very tight integration with Nvidia’s latest hardware. That includes the high bandwidth Infiniband networking as well as all of the GPU clusters and memory sharing and ensuring that it gives us the ultra-low latency, but at the same time you don’t need six certificates to be able to use it like you do with a lot of the hyperscalers. You can get up and running within a few minutes on the environment that you need without having to buffer around with all of the messy details. You can get access to the GPU within minutes.

Ross Katz: That makes a lot of sense. Also, because you’re focused so specifically on the AI workload and the AI hardware, you don’t need all of these general purpose permission schemes that have to manage all of these different services and how they work together. You can streamline the workflows that are most important to people right now. My understanding is that’s what allows you to deliver this experience for the kinds of AI workloads that we’re talking about. Am I thinking about that right?

Ilya Burkov: Absolutely. I don’t like to use the word bloatware, but there is a lot of bloatware that exists out there. What we want to do is make sure that they have what they need and not more than that. If they want to have an additional package that they need installed or a particular notebook or a library and so on, they can do that manually and we can help them get that up and running either through the internal solution that we provide or working through our partners, to provide that as an extension of the infrastructure.

Ross Katz: Ilya, one of the reasons I’m excited to have this conversation is because you come at this from a healthcare and life sciences angle and you all have a lot of interesting customer studies that you’ve done with specific biotech and life sciences companies doing this kind of work on Nebius. Would you mind sharing one or more of those case studies? Would love to understand a little deeper how companies are using Nebius.

Ilya Burkov: Absolutely. No company is using Nebius the same way as another company. Everybody is using it in their own unique and very interesting ways. I brought in three examples. One would be the work that we did with Stanford University on the CRISPR-GPT model with Professor Le Cong and his team. They built that for automated gene editing. Researchers from Stanford, Princeton, and Google DeepMind developed this CRISPR-GPT model, which is a LLM powered agent. It’s a system that automates gene editing experiments. This system streamlines the entire process, from selecting the appropriate CRISPR system to designing gene editing RNA and analyzing all the data. Essentially they’re giving the power of chat and communication to a very, very complicated process. Which means that a master student can have a post-doc level equivalent of assistance in their hands. By leveraging the Nebius infrastructure the team very efficiently trained and deployed these models, to basically further the advancement of gene editing.

Ross Katz: It said automating gene editing experiment — experiment, excuse me. Were those happening in silico where you’re running the experiments and then getting predicted results, or was it more of a lab automation type platform, or are we thinking of this as more like a co-scientist, a co-pilot for gene editing experiments?

Ilya Burkov: It’s like a co-pilot, but they do a mix of wet lab and actual computational side. You typically look at these models, you understand how it works and then you test it in the wet lab. It’s that collaboration that they incorporate. These are different scenarios. You’re looking at knockout genes, base editing, prime editing, there’s so many things and we don’t need to go into the details, but they run these models to be able to communicate with a person who is doing the research, having an interaction. The reason why it has CRISPR-GPT is because a person can ask it questions and with reasoning it can reply and say, okay are you using the right method, are you doing this, I would like to understand this gene, can you tell me a little bit more about this particular gene set, and so on. You can have that interaction between this agent that they built basically.

Ross Katz: They did the training and then they host the inference all inside of Nebius?

Ilya Burkov: Yeah, yeah. They did the training, they ran all the models, like I said, they use a combination of Stanford, Princeton, and Google DeepMind. It was a joint effort and we provided the compute for what they were doing.

Ross Katz: It’s interesting that even Google DeepMind would want to participate in something like this to leverage Nebius for it. Would love to hear your other two case studies. Thanks for digging a little deeper.

Ilya Burkov: Yeah, sure. The other one was Convergence Bio. They’re basically redefining precision medicine by combining single cell RNA sequencing with large language models. They use that to unlock the patient level therapeutic insights. They built a model for single cell RNA sequencing, which was far more accurate than scGPT, which was the Zuckerberg and Chan foundation model. They trained a full transcriptome foundation model Convergence sc, which is completely done on Nebius infrastructure. That was I believe one or two nodes of H100s for a month that they needed to do that training. It’s capable of processing more than 20,000 genes per cell and it has outcompeted compared to all of the other single cell models out there. A state-of-the-art accuracy. It’s much more explainable for drug discovery and clinical trial development, as well as looking at the efficiency of treatments on particularly oncological pathways.

Ross Katz: That’s interesting. One of the things I remembered about this case study was that one of the challenges with training any of these large language models is figuring out how to get all of the computations done efficiently while also holding the appropriate weights, the appropriate optimizer states etc. in memory. There were some memory optimizations that you had to help them do in order to make this possible, like distributed data parallelism, fully sharded data parallelism and then tensor parallelism. I’m wondering if you can give us some insight into how you help with those kinds of memory optimizations to really get the most out of the GPU infrastructure that you’re using for this kind of exercise.

Ilya Burkov: Absolutely. We have the expertise in-house. We also have the experts in those teams that we sync on. They have a problem that they want to understand, they want to have a brainstorming session where they understand the challenge that they want to face. We sit down, we discuss all of those problems that they might be facing and we work through it together. In this case the team were extremely competent to do a lot of this work on their own. Hats off to them. They were working with I think it was 36 million cells in the dataset, with I believe it was more than two terabytes or something like five terabytes and trillions of tokens within that. We came in where we had to come in, but I would say 90% or 95% of the workload was handled internally. We just provided them with the stable infrastructure to be able to not even think about that.

Ross Katz: It sounds like for teams that have the idea of what they want to do but maybe don’t have the in-house expertise to do those kinds of optimizations, that’s part of the service that you’re able to provide as well.

Ilya Burkov: Absolutely. Yeah. That’s right. Time is limited, but I’ll say xAID is another very interesting healthcare — rather than drug discovery or drug development — side. They built an AI assistant for medical imaging. It was developed specifically looking at improving diagnostic accuracy and efficiency. They looked at training cycles that lasted over five days on very noisy clinical data. They have imaging data that tends to be quite noisy. They relied on our infrastructure for the interrupted, very high performance compute at a scale, plus support from our MLOps team to be able to clear this up, to be able to work through their datasets. This collaboration really ensured that they could handle all of the medical imaging tasks efficiently and effectively. That’s what they did. They provided a lot of these models to healthcare providers, and those are being used by medical imaging teams across different hospitals and academic environments.

Ross Katz: Just from those three examples we really get a sense of the multimodality that you mentioned earlier of the datasets that need to be brought together and the diversity of use cases that are out there for biotech and healthcare to train and run inference on these types of models.

Ilya Burkov: That’s right. Completely different workloads and that’s why I tried to make it a balance of the different segments within the healthcare side. But the GPUs at the core are the same. It’s just how they’re used is very different.

Ross Katz: Awesome. Well, uh Hugo, I want to I want to turn the conversation back over to you and just, you know, hearing some of these examples of the types of of the types of workloads and the types of models that are being trained. I’m just interested in from the perspective of a Saturn of like a team that’s onboarding onto Saturn Cloud and wanting to attack some of these use cases. What does a workflow look like for getting up and running with Saturn Cloud on top of Nebius?

Hugo Shi: Yeah, so generally you start by spinning up a development resource within Saturn Cloud. So that’s generally a server that has JupyterLab installed, but also with SSH access so that you can use VS Code, PyCharm, Cursor, or any desktop IDE. Um, and so we we always recommend that people write their code and develop and test it interactively. Um, because we do have some customers that just go through like job launching loops and that’s slower because you’re spinning up a new machine to, you know, test something. Um, but once you’ve got something working well in development then we would encourage you to clone it into a different type of resource. Either a deployment if you want to like deploy a model, um, or a job which is what most people are using for experimentation. And so um that’s and that’s when you start scaling, right? So typically it’s you develop, um and then once you’re happy with it, then maybe you spin up 1,000 or 2,000 parallel jobs to do some parameter scan. Um once you have the results of those experiments then you can turn it into a dashboard or something to or deploy a model for your end users. And that’s what the typical workflow looks like. Um, I unfortunately do not have a biotech background so I can’t get into the depth like Ilya can about what people are doing specifically, but I can talk from a high level about what their workflow patterns look like.

Ross Katz: Ilya, it’s pretty clear to me how Nebius sort of uh can compete with the with the hyperscalers and the neo-clouds, but like I’m interested in hearing from you Hugo, you know, my understanding is that when when companies are looking to alternative, you know like alternative for this kind of workflow they might already be on Google Cloud, so they’re thinking about Vertex AI or they might already be on Azure, they’re thinking about Azure ML or they might already be on AWS, they think about SageMaker or, you know, they’re sort of rolling their own Kubernetes cluster to uh to do this hard work for them. I’m just interested in hearing from you, you know, how um where do you view the value proposition that Saturn Cloud brings relative to the other alternatives that are out there?

Hugo Shi: Yeah. Um. So there’s a lot, so that’s a very there’s a lot and it’s something that I thought about a lot and there’s a lot of subtle things that we I could talk about, but I’m I’m not going to because I think there’s something clearer that will make that that will just address it. Um, I will say the subtle things are that Saturn Cloud requires less DevOps and infrastructure work. Um, generally if you’re using a tool like SageMaker, you also have to have a DevOps team configure and sort of build what you need on top of that. Um, and also I would say that like our stuff is a lot easier to use and easier to troubleshoot. Um, that is not a very satisfying answer because it’s very qualitative and I can’t like prove it. The only thing I can say is like you could look at our G2 reviews to get a sense. Um, so that’s why I’m not going to spend a lot of time on that stuff. The the key thing, the key distinction is that we’re multi-cloud. So if you’re using SageMaker, you’re stuck on AWS. You can’t leverage Nebius. If you’re using Vertex, you’re stuck on GCP. You can’t leverage Nebius. So like the clearest value prop here is that we let you use other clouds. Um, and you can use GPU capacity and pricing where you can get it.

Ilya Burkov: This might sound strange, but we are cloud agnostic as a company. Nebius understands that some people might have a workflow with a hyperscaler. We understand that they might have reliance on on using one of the the the big five or multiple of the big fives. They get credits and so on. We’re not against that. Feel free to do that, but bring your generative AI workflows, the ones that require very AI-centric requirements to Nebius and that’s it. Um, you don’t need to move your entire or migrate your entire workflow into Nebius. We’re very comfortable and we work with so many companies that that have this multi-cloud approach.

Hugo Shi: Oh I just realized that I didn’t fully answer the question because the other part was about like rolling your own stuff on top of Kubernetes. And I’d say there’s there’s a couple areas that I want to address. One is that people who roll their own infrastructure on top of Kubernetes, and this is just something that I want to say because I’ve seen a lot of people make this mistake and I want to make sure people don’t do that. Like if you don’t want to use Saturn Cloud, no problem, but if you’re going to roll your own infrastructure don’t don’t make these mistakes. So the first one is that often times people will just go deploy JupyterHub. Great product, you know, works very well and it’s very good at what it does. The downside of doing that is that I’m I’m personally against platforms that push people into just using notebooks. Um, the reason I’m against that is because AI and ML practitioners have been leveling up their technical capabilities over the past 10 years. Um and I think that trend is a very good trend is very important and I think trying to force people to use notebooks is a disservice to that effort, because then you don’t get to leverage a lot of coding best practices, unit testing, splitting up things into multiple files, um and also you don’t get to leverage a lot of the features that full featured IDES bring. Because JupyterLab is pretty good but VS Code is better, PyCharm is better, Cursor is better um for for writing code, right? Those things are better at writing code. And so use notebooks for experimentation, but if you say well you can’t use like a full featured coding thing, that’s a very big disservice to your users. That’s the first thing. Um, the second thing is that I think a lot of people feel like it’s a one-time cost, you just build the thing and then you’re all set. Um that’s unfortunately I wish that were true, because then we would have just built software and then just like went to the beach and not worked anymore. Um when the pattern that we see is that platform teams will spin up some infrastructure and it’ll work. Um and then requests start to come in, like oh well I need to be able to do multi-node training. And so then they have to build that into their platform. Uh and then you know we actually do need some cost controls because last month’s bill was like way too high. So then they need to build that. Uh and so over time you’ve gone from like I was going to do this one thing to I’ve got people working on this full time. And so if that’s the if that’s your competitive advantage then then great, but if it’s not then don’t do it. Um and the last thing the last thing that I’ll say is um the stuff that we focus on, which is development workspaces, deploying um containers, deploying jobs, those things are very reusable and very core to a lot of workflows. Uh and it’s very easy for us to do that very well. I treat that as basically you shouldn’t be doing that, right? Because those things are very those support 90% of the workloads. There’s our solution handles that well, there’s other solutions that can handle that well, you should not build stuff to handle those things. So if you are a platform engineer and infrastructure engineer and you have time, you should be focusing on the remaining 10% that doesn’t exist on off-the-shelf tools that will actually provide differentiation and value to your business. Don’t solve the problems that are easy for people like me to solve, solve the problems that I can’t solve.

Ross Katz: Would you mind just touching on what BioNeMo is and then, you know, how Nebius enables biotech teams to, you know, leverage BioNeMo more effectively inside of the Nebius environment?

Ilya Burkov: Yeah, sure. I mean, BioNeMo is Nvidia’s domain-specific, very generative AI platform for life sciences. Um, it’s part of the NIMS packages, which are the Nvidia microservices. Um, so it’s the inference side. Um, it provides a lot of pre-trained and customizable models for things like molecular design or protein structure prediction, generative chemistry. There’s there’s a whole suite of software, um and solutions that are out there. And what Nebius does is it makes BioNeMo much more accessible and scalable. Um, it’s a one-click to deploy solution. We run Nvidia’s optimized infrastructure so that, you know, biotech teams can train, they can find tune or infer, um, with these very massive models, much more efficiently. Uh without needing to to manage the high performance compute clusters or looking at things like GPU scheduling and so on. Um in short, I’d say that BioNeMo brings the intelligence and Nebius provides the engine that powers it. Uh it it even though it’s an open source framework, um it’s incredibly powerful. It provides all of these things and fills the gap, um where Nebius can provide this turn-key approach, um you know, that AI-native infrastructure that’s easy to use and optimized for very, very complicated processes. Uh without having to download look at how stable that open source model is, there’s no level of support in in the open source community other than relying on getting some, you know, help from somebody who developed it, but they’re very busy. So with these BioNeMo packages, uh internally you can get the level of support from Nebius plus, uh if you have the enterprise license, from Nvidia as well.

Ross Katz: Yeah, that makes a lot of sense. And also so my understanding is that BioNeMo has sort of open source models that have been pre-optimized so, you know, optimized for the for the Nvidia for Nvidia compute and, you know, uh on Nebius uh rather than sort of having to verify directly inside of your own compute environment that you’re actually getting the benefits that are alleged in the open source package, you can just sort of see you you can know that Nebius has has validated that the that you’re getting those performance improvements directly. Am I thinking about that right?

Ilya Burkov: Correct, completely. And and what why Nvidia package it in this way is because they’ve accelerated those license so those individual components uh to an extent that is two to three to four times, sometimes it’s you know, 20 times more performant than the open source version of the software, because they have the expertise in-house to be able to get the absolute maximum out of that GPU.

Ross Katz: Awesome. As we’re heading to the end I just want to finish with a few last questions. Ilya, looking toward the future, where do you see the biggest opportunities for AI acceleration in life sciences over the next couple years?

Ilya Burkov: Next couple years? I guess the biggest opportunities would be in drug discovery, bioprocess optimization, personalized medicine as well. In drug discovery AI models can predict molecular interactions, they can look at designing novel compounds, prioritizing candidates much faster than traditional methods. For bioprocessing, again this takes a long time, this needs a lot of expertise, but AI on a whole can help optimize the cell cultures, fermentations, purifications, constant check-ins, looking at improving yield. There’s so many different aspects that it can accelerate and scale up. For personalized medicine, it can analyze patient-specific data almost in real time, can have those validation points to guide the therapy design and choose the right doses, get the clinical trial much quicker, really approach it on a person-by-person basis rather than a generic disease or a generic condition basis.

Ross Katz: Awesome. Hugo, as GPUs become more accessible and as more of these workloads start to require GPUs, the supply and demand escalate in parallel, do you see AI and ML workflows evolving or changing as more of these models need to be trained, but also foundation models become more mature, and how do you see Saturn Cloud playing a role in that?

Hugo Shi: Yeah, there’s going to be a ton of changes. That’s the understatement of the year. I would say there are two GPU trends or two things in the GPU space that I think are interesting. One is that Nvidia has been focused not only on generative AI, but they also have been working on a lot of projects that are not specific to generative AI that help leverage the GPU. These are packages like Rapids, for example. Tools that can help accelerate things that you would traditionally not run on the GPU, but now you can and you get massive performance benefits. Rapids has been around for a while but it’s getting more and more popular and we’re seeing more and more of that usage on Saturn Cloud. The other thing that I’ll talk about — and this might be too nerdy — but Nvidia a couple years ago acquired a company called Run:ai. Part of that acquisition was the Run:ai scheduler which they have open sourced as a project called the Kai scheduler. That’s K-A-I and then scheduler. We’re working on integrating with that. We think it’s pretty cool. The two reasons why we like it is one, it handles batch scheduling. Batch scheduling is a situation where if you’re launching a multi-node training job, let’s say you want to use 10 nodes, it’s not useful to start that training job before all 10 nodes are available. If you have five nodes available and you start five then it’s just going to hang out and do nothing until the other five come online. Batch scheduling is one of the things that it offers. There are a lot of products that do batch scheduling so that’s not really that special, but it is important. But what’s also more special is that it has a nice hierarchical queue structure. You can do things like set up a queue for each of the teams that you have and then within those you can say, well we’ve got a production queue, we have a queue for interactive work, we have a queue for batch jobs. Maybe the batch jobs get lower priority because you don’t care when they finish. The production queue gets high priority because it’s production. The interactive queue has less resources because you don’t need that many resources interactively, but maybe it’s optimized for spinning things up quickly so that the user’s not waiting. For scheduling on GPU hardware, things like the Kai scheduler are going to become more and more important. That might be too specific, but those are the two trends that I’m excited about right now.

Ross Katz: Not at all. I think it illustrates the point that we’ve made throughout this entire podcast that staying on top of the emerging AI landscape, the emerging MLOps and DevOps tools landscape and the emerging high performance compute landscape is impossible to stay up on any one of those landscapes, which is why you want the expertise of people like yourselves in the middle choosing the things that are most relevant and deploying them for you in a way that makes it easy for you to take advantage of them. One of the points that you made was basically if you’re provisioning on-demand GPUs on any of the hyperscalers, you’re lighting money on fire. Do you have a sense of the amount of money, relatively speaking, that you’re lighting on fire versus if you’re in Saturn Cloud it’s as simple as just hot swapping the workload over to Nebius, and it’s immediately less expensive. Would love to know the order of magnitude of that.

Hugo Shi: I think pricing depends on which cloud you’re talking about. Also AWS recently dropped some of their prices, but I think it’s on the order of 4x. Ilya, does that sound correct?

Ilya Burkov: Yeah, we typically find that it’s between 60 and 70% cheaper within that environment.

Hugo Shi: Right. That’s just the sticker price. The other part is that with AWS, GCP, Azure, you often can’t get on-demand. That means that instead of being able to try stuff out for a month, you’re committing to a year. Maybe you just wouldn’t do that, but on its face that’s just a 12x increase. It’s quite substantial. When you consider the fact that an 8x GPU, an 8x H200 instance is about $20,000 for a month, 70% is a lot of money.

Ross Katz: Hugo and Ilya, it’s been great to have you on the podcast. Briefly before we go, can each of you share where listeners can go to find out more about Saturn Cloud and Nebius?

Hugo Shi: Yeah, saturncloud.io is our website, you can go there.

Ross Katz: And to connect with you as well?

Hugo Shi: LinkedIn is probably the best. I think my LinkedIn username is hugo-shi.

Ilya Burkov: nebius.com, simple as that. Similar to Hugo, just find me on LinkedIn. We do have a life science and healthcare dedicated site on the Nebius website, so there’s lots of interesting case studies there and food for thought. Feel free to reach out, I’m happy to connect to have a discussion and really understand what we can do.

Ross Katz: Awesome. Hugo and Ilya, it’s been great to have you on the podcast, really appreciate the time and look forward to connecting down the line.

Ilya Burkov: Thank you.

Hugo Shi: Thank you so much.

Jason: And that’s it for this episode of Data in Biotech. If you enjoyed the episode please subscribe, rate, or leave a review in your podcast player of choice. See you next time.

Frequently Asked
Questions

How can my biotech company reduce the cost of GPU compute for our AI initiatives?
Consider specialized AI cloud platforms like Nebius, which report 60-70% cost reductions compared to general hyperscalers. These platforms offer dedicated GPU infrastructure and optimized pricing models designed for high-performance AI workloads. This allows for significant budget reallocation towards research and development rather than infrastructure overhead.
What steps can we take to ensure strict data security and compliance for sensitive biotech AI projects?
Look for AI infrastructure providers that build security and compliance into their core design, supporting frameworks like HIPAA, GDPR, and relevant ISO standards. Platforms like Saturn Cloud and Nebius focus on creating secure environments with integrated guardrails and auditability, allowing your teams to work with sensitive patient data and proprietary research confidently.
How can our data scientists effectively deploy and manage unique, often academic, machine learning models without extensive DevOps support?
Utilize platforms that offer open environments and abstracted infrastructure management, such as Saturn Cloud. These solutions support diverse runtime requirements for specialized academic code and simplify deployment. This approach minimizes the need for deep DevOps expertise within scientific teams, letting them focus on model development and analysis.

Need a data partner for life sciences?

CorrDyn helps biotech and pharma companies build the data infrastructure that accelerates research and operations.

Book an intro call