Listen on
Overview
Many biotech R&D organizations face a critical data problem: their fragmented infrastructure, reliant on spreadsheets and siloed legacy systems, actively hinders innovation and inflates costs. This “tool fatigue” and lack of interoperability make companies “truly non-scalable,” preventing them from building a proprietary data asset that accelerates drug discovery and increases enterprise value.
In this episode, host Ross Katz talks with Guru and Satya Singh, co-founders of biotech software company SciSpot, who explain why biotech companies need a “digital brain” to become AI-ready from day one. Guru, a former lab scientist, and Satya, a data veteran from Expedia, combine their experiences to diagnose the industry’s data challenges, offering a vision for a future where biotech labs learn from every experiment. They share how SciSpot’s API-first, middleware approach standardizes data models, connects disparate systems, and fosters real-time feedback loops.
The conversation explores how SciSpot’s platform transforms unstructured lab data into structured knowledge graphs, provides templatized workflows for rapid implementation, and bridges the communication gap between wet lab scientists and computational experts. Learn how moving beyond traditional ELN/LIMS approaches can help your biotech organization build a strong data moat and accelerate the journey to market.
Key Takeaways
Biotech R&D’s scalability challenge stems from disconnected data, not just lab work.
Traditional biotech operations, heavily reliant on manual data entry, spreadsheets, and siloed ELN/LIMS systems, make companies “truly non-scalable.” As R&D becomes more externalized and distributed across CROs and cloud labs, the inability of these tools to share and harmonize data creates significant bottlenecks. Organizations cannot build a long-term data asset when data is trapped in disparate systems and requires manual intervention for analysis.
A “digital brain” for biotech is built on AI-compatible data and real-time feedback.
For biotech companies to become smarter with every experiment, their data infrastructure must automatically transform unstructured lab notes into structured, labeled datasets. By vectorizing and storing this information in knowledge graphs, companies create an AI-ready foundation. This approach enables the training of proprietary AI models, establishing a direct feedback loop that guides future experiments and ultimately generates a valuable, unique data moat.
API-first middleware, rather than a single “source of truth,” accelerates data flow in biotech.
Instead of forcing all data into one monolithic system, biotech needs agnostic middleware that connects diverse tools—from legacy ELN/LIMS to instruments and external partners. An API-first platform with intelligent data transformation layers ensures data is ingested and harmonized correctly, even from systems without native APIs. This reduces “tool fatigue,” allows companies to use the best tools for specific tasks, and focuses efforts on competitive differentiation rather than system building.
Templatized data models and workflows eliminate months of implementation pain.
Historically, implementing new data systems in biotech could take six months or more. SciSpot addresses this by offering pre-built data models and integrations for over 20 distinct biotech company types. This templatization allows for “day one” implementation, accelerating time to value and enabling companies to quickly establish a structured data backbone tailored to their specific assays, instruments, and regulatory requirements.
Related: CorrDyn provides data engineering expertise and helps organizations develop strong AI strategies. We also specialize in the biotech and life sciences industry. Learn more about how biotech manufacturers gain data value for competitive advantage.
Full Transcript
Jason: Hi everyone, this is Jason, producer of Data in Biotech. Before we get started, I wanted to let you know about our latest white paper. It’s a comprehensive guide to implementing machine learning models in biotech manufacturing. It’s a complete overview of all the potential problems of ML adoption and, more importantly, how to solve them. To download it, simply visit connect.corrdyn.com/biotech-ml. We’ve also dropped the link in the show notes of this episode. Welcome to Data in Biotech, a podcast from CorrDyn where we explore how companies leverage data to drive innovation in life sciences. Every two weeks, we sit down with an expert from the world of biotechnology to understand how they are using data science to solve technical challenges, streamline operations, and further innovation in their business. This week, we sat down with Guru and Satya Singh, two brothers that founded the biotech software company Scispot. Scispot’s infrastructure platform helps biotech companies create their own digital brain and accelerate their R&D process. During the interview, we covered a lot of topics, including how to bridge the gap between wet lab scientists and computational experts in biotech organizations, why the future of biotech is externalized and interconnected with a focus on standardized data models and easy information exchange, and how to enable real-time feedback within biotech that helps companies become smarter with every experiment. Here we go.
Ross Katz: Guru and Satya Singh, welcome to the Data in Biotech podcast. Just to get started and maybe Guru you can go first, could you just give us a brief introduction to you and your background and what brought you to this point?
Guru Singh: I’m a former lab rat, proud lab rat — worked at the bench, but I was always entrepreneur at heart. Satya and I grew up in a very business family. My first startup was when I was 8, 9 years old. I used to sell aluminum waste from my dad’s shoe factory. Always very entrepreneurial, not very social, arguably. When I moved to the US over a decade ago after grad school, I worked for other biotech software companies. What I realized is there’s tool fatigue in the industry. Research is becoming more and more externalized, distributed, wet lab, dry lab, external partners, but most of the tools, software, hardware, don’t talk to each other and the data is not easily harmonized, aggregated, or usable. So teamed up with Satya, launched Scispot just to fix that problem.
Ross Katz: Satya?
Satya Singh: My background is all data related. I worked as a product and data head in Expedia, Hotels.com, Discovery Channel, Eurosport Player. Enjoyed all of them, but I realized I’m basically making people click more — for good reasons, like sports or travel to broaden the mind, but coming to biotech, I feel like it’s more fulfilling. Big pharma, biotech, there’s more meaningful results you can drive with data rather than just making people click more. I felt like this was something I could relate to. And Guru and I, we lost our mother in 2011, so we always wanted to do something together in life science in memory of her. The drug came too late for her — two years after she passed the drug was available in the market. That hit us really hard personally, too. That’s a little bit of background from my side.
Ross Katz: Yeah, that’s a really touching story. So what made you decide to focus on Scispot? How did the idea of Scispot come into being?
Guru Singh: Observing this industry for almost 15 years now, what I realized is when I was at the bench, the cost to generate data was really expensive — generating big data from labs was really expensive. For example, it cost over 10 million dollars to sequence the entire human genome when I was in the lab about 15 years ago. Now it costs a few hundred dollars. This is just one example, but the cost to generate big data is going down. At the same time, we are seeing substantial growth in AI. AI has been democratized to some extent and I don’t know where we’ll be in five years. The marriage between this organic intelligence — which is bio, that generates tons of big data — and inorganic intelligence, which is artificial intelligence, they’re made for each other. Very strong partners, because one generates a lot of data now relatively inexpensively, and the other, AI, can analyze that data and make sense out of it. If we want to make the biggest impact to accelerate R&D, the best thing we can do is take these bio entities with too many variables and tons of big data and help them become AI ready. They can learn from their data, accelerate their speed, and hopefully bring drugs to market two years faster, three years faster, maybe 10 years faster. Within a few months, we will start seeing new drugs in the market. Obviously, this is wishful thinking, but that’s our hope.
Ross Katz: Yeah, that makes a lot of sense. So how is Scispot structured in order to solve these particular problems of bringing the organic intelligence and inorganic intelligence together in that optimal feedback loop?
Guru Singh: When I look at a biocompany, I feel like they all in the future will have their own digital brain. Just like when you have a baby with a limited number of synapses and neurons — with more exposure, the brain becomes enriched, and it’s more plastic when it’s early stage. That’s why we focused on these startup-stage biocompanies, because their brain is still evolving, it’s very plastic, they are creating new nodes, new synapses. We see every biocompany having a digital brain that will learn from every experiment they run, every protocol deviation they make, every inventory item they use, every result they generate, every data transformation they do. We provide data infrastructure that vectorizes the data and creates a graph database for every company. It’s a new generation of lab solution because we think ELN and LIMS are a decade-old approach. This is a more data science focused platform to help you create your digital brain and become smarter as you grow. Data is your primary asset. Drug candidates are also valuable, but data is your primary asset.
Ross Katz: Maybe it would just be useful to break down the Scispot platform into the components of it that help to create that world that you’re talking about where companies are learning from every experiment that they run and every transformation that they do and you’re creating that knowledge graph, that knowledge base that companies can work from?
Guru Singh: The heart of the product is a data lake with built-in connectors, even for complementary or competing systems. On top of that, we have apps — alternatives to ELN, LIMS, SDMS — focused on data science, API-first apps that connect on top of it. When we go to smaller companies, they use our data infrastructure. They need ELN for experiment design and execution, but every experiment has a data model behind it, meaning you can directly open the knowledge graph per protocol, per experiment and see all the nodes connected. What materials you used, what protocol steps you ran, results file — everything is connected in a graph database. Similarly, we have instrument connectors, so you can programmatically bring data from other apps and instruments and send it back to your own cloud storage. The heart of the product is the data lake; on top of that we have multiple apps, which is optional. You can use Scispot native apps or you can keep using your own ELN and LIMS.
Ross Katz: Satya would you add anything?
Satya Singh: One key differentiator in the market: I come from a travel data background where the margins are so low that everything’s open. You have an API where you collaborate between Trivago and TripAdvisor, or Hotels.com says, “I’m going to focus on business travelers and I’ll let my competitor focus on family travelers, I’ll just take commission on the referral I’m making through this API.” Biotech is probably a decade behind that. I think we’re coming much closer now and I’m seeing a lot of progress in startups. Our philosophy right from the beginning was we don’t want to be another Microsoft for bio or one source of truth for bio. Without naming any competitors, we don’t want to be the “put everything here” platform. We want to be the middleware where we’re agnostic to what tools you’re using. With that philosophy in mind, it has really helped us differentiate. You want to use Scispot ELN? Great. You want to use a LIMS you’ve built in-house? You want to use Scispot integration because you’ve already invested millions of dollars in your current ELN? That’s fine too — here’s the pipeline that helps you transform that data into a graph database. That’s been our philosophy from the beginning and it has driven good results. It’s been challenging to build all of that, so we’ve spent a lot of time building this infrastructure and we’re continuously adding more features. But that’s been our mindset in terms of our product roadmap.
Ross Katz: I saw on your website that you have Alt-ELN, Alt-LIMS — you’re positioning yourselves as able to take on the capabilities of all of these traditional systems, but do it in a way that is data-first, API-first, embodying modern software and data principles. What kind of pushback or cultural barriers are you encountering when you try to sell into the market in scientific disciplines where people might be used to the tools they’re already using or maybe not as apt to try to do things a different way?
Guru Singh: If we can make them use the system, experience the system, there is an instant aha moment. But at the time of selling, it becomes a challenge because individual scientists are not necessarily thinking about future-proofing their company or making it AI-first. They’re thinking: I prepare multiwell plates, I manually send them to the instrument, and I have my calculations set up in Excel. It’s working for me, and if I can do this in ELN, that’s what I need. But that’s a very short-sighted approach and it’s very hard to convince them. We need to do a better job articulating why you want to use an API-first, data science focused system. That’s our biggest challenge — it’s more educational in a way, but as soon as we put the product in the hands of customers, there are no questions.
Satya Singh: Where it helps most is having diversity in the team. Scientists sometimes speak orthogonal languages depending on what background they come from — which is true in every industry. But I’m applying my principles from different industries into biotech, and I can see a lot of patterns. One of the things is: how do you align wet lab and computational folks into that Venn diagram? How do you make sure their principles are not orthogonal? Or even if they are orthogonal, how do you use different tooling that can integrate into similar patterns? Those are the things I’ve learned over the years working with these organizations.
Ross Katz: You talked about the aha moment that people have when they get into the platform. Can you help me understand what that aha moment looks like? What is it about Scispot that gives people that feeling of “this is the way it’s supposed to be — this is how I want to be doing my biological data science workflows”?
Guru Singh: The biggest aha moment is Scispot Glue — you can connect with all of your internal systems. We leverage AI to transform your data and automatically create your data model. The first aha moment is standardizing your data model. If you are a lab-grown meat company, an organ-on-chip company, or doing a proteomics workflow, your data model is slightly or significantly different. We have that learning. You connect with all the data sources you have, we design your data model. Aha number one. The second aha moment is when you connect all of your instruments with Scispot, because we have built-in not just an extraction layer but a transformation layer and analytics layer. You don’t have to manually receive CSV spreadsheets, upload them somewhere, set up your formulas, or send them to your bioinformatics team. That’s last decade’s flow. We want to change that. There are also long-term aha moments. Most companies, at least now, start working on their AI strategy after a year or two — in most cases after Series A. Some companies start as AI-first, but most prioritize it at later stage. If they are using the platform, because we are already cleaning the data, vectorizing it, storing it in graph database, labeling and classifying it with AI, they can train their own models. That’s really an aha moment — they don’t realize they were sitting on top of this gold mine. The more they use the system, the smarter their company becomes.
Ross Katz: One of the things that I’ve gathered, from looking at your website and from the conversations you’ve had, is that yes, you’re creating the explicit data model for companies based on their systems, but you’re also doing this knowledge graph work — implicit data analysis, entity extraction, and relationship building. You mentioned graph databases a little bit ago. Can you give us some insight into what that process looks like from Scispot’s angle and what value that brings that wouldn’t otherwise be there?
Satya Singh: It’s implicit, but there’s also a human in the loop because there are a lot of things humans want to tag for accuracy. When a user comes in — let’s say they start using Scispot, Alt-ELN and Alt-LIMS — and they already have a lot of experiments designed and executed with samples and inventory associated, their lab ops manager or someone from the lab comes in and says, “I want to see the chain of custody for this particular sample.” They can go to the knowledge graph and see the complete chain of custody: this sample was created five days ago, it was used across three methods or protocols for this experiment by this user. But the implicit tagging hasn’t captured how much inventory is left, how much needs to be reordered, or whether this experiment was successful. We take those inputs from the human. So there is implicit tagging of chain of custody, and there is a human who can provide additional feedback. We’re continuously helping the customer optimize that knowledge graph. It can be applied in two ways — on structured data and on unstructured data — because a lot of what scientists capture is very unstructured. In the back end we create embeddings, there’s function calling, and we store this data in a vector database that helps us achieve this feedback loop from having human in the loop.
Ross Katz: Can you give me an example of what that looks like from the user’s perspective — when you have data in unstructured format and insights are getting parsed out of it that I can gain without having to tag it myself?
Satya Singh: Within our LabSpace of the Alt-ELN — that’s a Scispot internal term — you can use a backslash to embed different elements from your chemistry, molbio, protocols, your actual data. You can embed things in this notebook. As you’re embedding, it is automatically contextualizing all of that data behind the scenes, implicitly tagging it and understanding what it is used for and how it’s being used.
Guru Singh: Even when you are doing note-taking at the time of experiment execution or in the experiment design planning stage, you’re using terms like — let’s say you are working on Parkinson’s disease, targeting the LRRK2 gene. The system has the capability to auto-label these things: this is your gene, this is a protein of interest, this is the strain you’re working on. What is an experiment, if you think about it? It’s a recipe, a set of instructions, eventually for a machine to do the work for you. Currently we don’t have a connected way to do that, so humans are involved doing pipetting and all the rest. But ultimately it’s like writing rules — like code — to tell the machine what to do. That’s what experiment documentation is. We turn these unstructured notebooks into a very structured dataset with the right labels and right context, create a knowledge graph out of it, stored in a graph database and a vector database. You are not just documenting for auditing or regulatory purposes. You are documenting for data science, to really create your data moat in the long term.
Ross Katz: So is the idea that I run one experiment with one protocol, then run another experiment with a similar but not identical protocol, and I can understand the relationships between those two experiments — areas of overlap — so that as a data scientist I can just grab all experiments that look like this without having to do a lot of the data extraction myself? Am I thinking about that right?
Satya Singh: Basically, yeah. You can derive those deviations from this model as well to say this was a journey, this method had a baseline recipe and then it followed this journey with these deviation at every point in the journey.
Ross Katz: I’ve heard that you offer templatized workflows as part of the Scispot platform as well. Can you share a little bit about what a templatized workflow means in the context of Scispot and how that works from the user perspective?
Guru Singh: We have identified over 20 different types of biocompanies in our sweet spot. We understand how their data model is structured, what type of experiments they run, what type of plates they use, what type of integrations they perform. We templatize all of that. When we onboard a customer, we don’t spend months or days on implementation. Implementation happens on day one. People think it’s just a marketing line, but because it’s been templatized and learned from past experience, we can set up all the data models, all the databases and how they’re connected, and all the integrations they need based on their workflow — whether it’s MassSpec heavy, HPLC heavy, or qPCR for diagnostics. When they bring in data, we already know what type of transformation, QC, and automation to apply. Cell passaging workflow? You have to do the cell count, so flow cytometry. We already have that learning. We have templatized accounts for over 20 different types of biocompany, which covers arguably 80% of companies in our market. Our goal is to cover 100% of companies so there’s no human involved in a six-month implementation. I see many systems that still take six months just to do the implementation — it’s very painful.
Ross Katz: How does the way that a biotech organization thinks about data or data science workflows change once they adopt the Scispot platform and the Scispot paradigm?
Guru Singh: Most startups, when they’re looking for a solution, have very rigid workflows or very specific needs. But as they start scaling, their workflows evolve. They started as a contract research organization, but now they have their own drug candidate — that requires a slightly different version, their requirements change. Or if their products enter late stage and bio-manufacturing is happening, more GxP rigid workflows are involved. These companies are evolving, but more traditional systems miss the main point by creating hardcoded solutions. Your company is growing, it’s a plastic brain which is growing. What we do differently is the platform is moldable and it can grow with you. That’s very important. Most companies are plastic in nature.
Ross Katz: Satya?
Satya Singh: The UI and user journey is also changing for these companies. With AI coming into the picture, we’ll see a lot of tools have a hybridized approach — not just a UI interface or API-first, but natural language embedded in between those two as well. We’re not there yet for everything to be natural language. Microsoft tried this with SharePoint where they made the whole thing natural language based and they got some good insights — some parts worked really well, some parts didn’t work at all. I don’t think I’m on the extremes of either spectrum; there’s a balance. But a lot of biotech companies are realizing that the user journey is changing — how they think about these tools from both a wet lab and a computational perspective.
Ross Katz: The question that comes to mind for me is: what is the vision for how a biotech organization is supposed to work 10 years from now, as Scispot sees it? Because it’s clear that you all have a different perspective on how the ecosystem is supposed to play out — the role that Scispot plays as a third-party vendor, helping create the boilerplate that all of these organizations are building on top of so that they can focus on the areas that are their competitive advantage. Can you share a little bit about what that vision looks like?
Guru Singh: The biotech of the future is very externalized — whether they are using robotics labs or more CROs. Especially at early stage, it doesn’t make sense to even rent a bench space and run experiments themselves. We’re seeing most companies focus, even on day one, on in silico workflows, then leveraging CROs to validate some of their assumptions. Cloud labs are still expensive, so we are not there yet, but I think it’s going to happen sooner or later. The wet lab portion is also diminishing compared to the dry lab portion. The future requires tools that support companies that are more externalized in nature. Some work is happening in-house, but you cannot just focus on intra-lab connectivity — you have to focus on inter-lab connectivity. Compound synthesis is happening outside, analysis through a CRO, but some assay binding work is happening in-house. You cannot have all the external work stored in spreadsheets, the internal work in your ELN, and be collaborating with partners through email or asking your CRO to upload things to an S3 bucket. That makes your company truly non-scalable. You raise enough money, you create your own tech team, and they spend a couple of years building a tech stack. We want to change that. We want to make sure on day one, you have a blueprint of your company — you choose your workflow, what’s happening in-house through your own instruments, and what’s happening through external academic labs, CROs, and CDMOs. You choose all of that, you get your standardized data model, an easy way to exchange information, an easy way to automate things. That’s the goal. We are moving in that direction; we’re not 100% there yet.
Ross Katz: It seems like that vision is the driver of why build an API-first platform — having that API-first platform creates structure for the way things are automated internally, but it also creates the interface for things to happen externally as you start interfacing more with third-party CROs or CDMOs. Am I thinking about that right?
Guru Singh: API-first is a first step in the right direction. But then you have to think about: when you’re connecting two tools, what transformation do you need to make sure the data is ingested in the right format? What we do differently is — yes, it’s API-first, but if you’re bringing data from qPCR, what do you have to do in between? That middleware, the transformation happening in between, so it magically goes into the different system as if that data belongs to that system. How can you easily use multiple tools without becoming too dependent on one system? I think that’s the direction, and that’s what the industry has desperately needed.
Ross Katz: In data and technology broadly, API-first is just how things are done, but in biotech there are a lot of systems for which that paradigm is not necessarily there. How do you handle the barriers and roadblocks you face in trying to integrate these biotech-specific systems into this API-first platform?
Satya Singh: When I was at Hotels.com, we had this debate: software engineers think of everything as a product-centric approach, and then data science and data engineering teams also started thinking like software engineers. They used to think differently, but as they started building more tools and big data came into the picture, that changed. Bio is also going through that transition. You have computational folks who understand API-first, API as a product, data as a product — and then wet lab people who are now also coming on board where they’re using natural language to do some of these things. They don’t need to learn how to code or understand the technology, but they’re basically hitting an API in some shape or form and realizing the importance of that. Their thinking is also shifting toward how computational folks in bio think. It’s a journey where everybody’s coming closer to software thinking, which is beneficial because ultimately tools are software, they are in the cloud, and everybody can benefit if they follow the same patterns. I’m seeing that transition, but you’re right — it’s still not straightforward, because there’s change management involved depending on the type of company and how many legacy tools they’ve adopted, because if there’s a legacy system they have to use, they’ve already invested in it and it becomes quite hard to then bring in an API layer where you might need to do something manually.
Ross Katz: It strikes me that this way of working in the biotech space asks your customers to be data competent and data native. Where do you see the gaps in data capabilities at the companies that onboard with you? I see an entire spectrum that varies according to the roles that exist in these different biotech organizations. The wet lab scientist tends to be focused mostly on how to accomplish their workflow, because most of their mental capacity is spent driving the biological insight they’re trying to drive. They’re not really concerned about the entire ecosystem that might make use of that insight. Then you have the IT team, mostly concerned about integrating the software into the broader ecosystem and making sure it meets security and compliance specifications. Then the computational team, trying to do scalable analysis and make sure all the metadata is AI ready, to the point that Guru made. And at the business leader level, it’s: I just want my summary operational metrics for how our business is running, what resources we’re consuming, and how close we are to our next milestone. Building a platform that serves all of those personas and solves all of those problems is really hard — which is why there are multiple players in the space with different visions. But I find your vision of serving as the middleware between all these systems, while also creating workflows and processes behind the scenes that bootstrap these companies to drive insights and do AI faster, to be an interesting paradigm.
Satya Singh: It depends on the persona. We satisfy both wet lab and computational personas because we have a graphical user interface and the API toolset. The gap is reducing, but there’s still a significant gap between the two. Wet lab optimizes for: in this UI, I should be able to search and discover my data, plan experiments, execute experiments, and provide results or some sort of feedback loop to computational folks. The priorities for computational folks are slightly different — they want to programmatically create 96-well plates, connect with robots, connect with instruments. Most of them don’t want to go through a GUI to get there eventually. It’s a very hard problem to solve because it requires a lot of development. That’s why we front-loaded all of these problems so we can scale faster — we built a lot of this orchestration right from the beginning. But we try to modularize both pieces depending on who we’re onboarding. If we’re onboarding a computational person, we start with the API for Scispot’s LabSheets — LabSheets are scientific databases — and off they go. The onboarding entry point is from there. If we’re onboarding a wet lab scientist, someone who is really focused on the graphical user interface, the onboarding and user journey is slightly different. I think that will continue to evolve as we get smarter with natural language. We might bridge the gap, but there’s always going to be some differentiation between the two.
Ross Katz: I want to look toward the future for companies building on Scispot. Can you give some examples of where you think companies that adopt Scispot are going to be able to achieve success in ML or AI in ways they wouldn’t otherwise be enabled to achieve?
Guru Singh: It’s real-time feedback, a feedback loop. Every sample you process, every test you run at a molecular diagnostic company — your company should become smarter. Otherwise you are just a testing company. You’ll always be valued at some 10x multiple depending on which market you’re in. But a company’s value should be 1000x if they really have their digital brain. How we make companies different is they have an instant feedback loop. Every time they run an experiment, every time they run a test, the company becomes smarter. That’s the goal.
Ross Katz: Is the idea that, for an experimentally driven company, Scispot is going to be recommending the next experiment you need to run in order to drive the next insights?
Guru Singh: That’s short-term and we’re almost there. Proactive AI assistance — we are almost there, we are running some tests for that. But I’m thinking even further than that. You have this intelligent brain which is growing every time you’re running an experiment, connected with your own proprietary models. You have this feedback loop in place. But yes, guiding them in designing the right experiment, figuring out next steps, or placing an order for the right antibodies — these things are already almost there as we are learning and improving the process. The goal is: how can we help companies build their own proprietary models connected with their data, where 100% of the data is AI compatible? Most of the data currently AI cannot even touch. That’s the problem.
Ross Katz: Interesting. Satya any thoughts on that?
Satya Singh: This reminds me of what the Nvidia CEO mentioned a few months ago when he talked about how AI will touch biology and how biology has so many variables — and that’s where LLMs would have the biggest impact. Companies would be able to build their own smaller foundational models, not necessarily large language models, but smaller models specific to their digital brain, as Guru talked about earlier. My North Star is: as companies use Scispot, they can build that proprietary database that belongs to them, and that becomes their moat to grow the company. As they learn and have the feedback loop from their experiments, samples, inventory, and dry and wet lab connections, can they form their own graph database — whether that’s in Snowflake, Databricks, or Scispot, agnostic to where it lives? Having that graph database they can say, “I can turn this into a smaller language model, not a large language model, but my company’s LLM — and this is my moat for longer-term sustainability.”
Ross Katz: As we draw to a close, are there any parting thoughts you would leave our audience with? What are the situations where people should come and check you out and how can they do that?
Guru Singh: My number one recommendation: if you are starting a biocompany, think of it as a data company. You’re thinking about running your first set of experiments, creating a prototype — but your company will be bought not because of your assets. It’s going to be bought because you have some proprietary model or proprietary dataset. If that’s the focus, consider reaching out to us. We’re happy to share the insights we’ve gathered. You can go directly to scispot.com/demo to book a demo. If you want me and Satya on the call, just add a note — we love speaking with companies. Otherwise someone from our team will show up and help you prepare your data infrastructure.
Satya Singh: My push would be: don’t start by thinking about what you need from an ELN, LIMS, or SDMS. A lot of companies start with “I need a lab information management system” or “I need an electronic lab notebook.” I would rather flip that question — like how Amazon thinks about this — what are the outcomes, and then work backwards from there. They might want to start with LIMS, that’s great, but what are the outcomes that LIMS can help them drive? A lot of that is sometimes missing. I think a lot of modern biotech companies are realizing: my outcome is I should have a data model I can share with my computational team. Now let’s work backwards to what tools can help me do it. Some ELN might do half of it or 80% of it. Starting from outcomes and working backwards would be great, because traditionally companies have built these very meditated buckets of ELN, LIMS, and SDMS, and I feel like we should break away from that. Those are industry terms, so everybody uses them as a common language, but it’s a good time for a paradigm shift from those three terms.
Ross Katz: Guru and Satya, it’s been a pleasure to talk with you and learn more about Scispot. Really appreciate the time and look forward to connecting down the line.
Jason: And that’s it for this episode of Data in Biotech. If you enjoyed the episode, please subscribe, rate, or leave a review in your podcast platform of choice. See you next time.





