Skip to content
Data in BiotechEpisode 17

Electronic Lab Notebooks with SciNote

Brendan McCorkle, CEO of SciNote, discusses how ELNs balance structure and flexibility to replace paper and Excel in the lab.

43:18Full transcript below
BM

Brendan McCorkle

CEO at SciNote

Overview


Biotech R&D cycles often hit a wall: the clash between a scientist’s need for experimental flexibility and the organization’s demand for structured, analyzable data. This tension isn’t just a technical hurdle; it directly impacts innovation velocity, data quality, and the ability to derive meaningful insights for competitive advantage. Companies frequently find themselves choosing between enabling researchers with unbound freedom or imposing rigid data capture methods that stifle creativity, leading to either lost data context or stalled development pipelines.

Brendan McCorkle, CEO of SciNote, understands this dynamic deeply. With a “by scientists for scientists” product thesis, he challenges the industry’s approach to Electronic Lab Notebooks (ELNs), advocating for systems that bridge the gap between bench science and business objectives. His perspective is grounded in years of building developer, clinician, and teacher-centric tools, now applied to scientific research.

In this episode, Brendan and Ross Katz discuss how to manage the philosophical divide in the ELN market, the distinct needs of various stakeholders from R&D scientists to executive leadership, and the critical importance of bidirectional data flow. They explore integration challenges, the pathway from paper to enterprise-grade data infrastructure, and the strategic groundwork required for future machine learning and AI adoption in the lab.

Key Takeaways

ELNs often create “air gaps” that hinder FAIR data, rather than enabling it.

Many ELN vendors inadvertently make data less Findable, Accessible, Interoperable, and Reusable (FAIR) by prioritizing extreme flexibility or creating proprietary walled gardens. The critical challenge is building systems that provide enough structure to normalize data for downstream analysis while preserving the scientific creativity required in early R&D. This balance prevents data loss and ensures context flows from bench to business.

Most “R&D” ELNs underserve the “Development” phase, slowing scale-up.

Traditional ELN approaches cater heavily to pure research, where maximum flexibility is prized. However, this often leaves a critical gap in the “development” phase—the transition from novel experiment to repeatable process. Systems that fail to integrate structured data capture and lineage tracking at this stage impede the ability to scale, replicate, and ultimately commercialize scientific breakthroughs, leading to significant rework and lost time.

Data analysis without scientific context is inefficient; feedback loops are essential.

Isolating data analysis from the original scientific context diminishes its value. If data teams receive raw output without metadata, experimental conditions, or researcher intent, significant time is wasted inferring meaning or correcting inaccuracies. Establishing bidirectional feedback loops between scientists and data analysts—where context flows out and insights flow back—is crucial for making faster, more informed decisions and preventing the repetition of flawed experiments.

After digitizing paper, the next critical step for lab data is phasing out Excel.

While many labs focus on digitizing paper notebooks, the hidden data problem lies in widespread Excel use. Excel spreadsheets lack reliable version control, security, and scalability for modern biotech datasets, making data difficult to audit, integrate, and analyze at scale. Moving to database-backed systems with proper interfaces is the “second phase” of digital transformation, essential for securing intellectual property and enabling sophisticated analytics.

Related: CorrDyn helps biotech organizations build strong data foundations, with services ranging from data engineering to ensure seamless data flow, to data quality initiatives that prevent costly errors. Learn more about how biotech manufacturers gain data value for competitive advantage.

Full Transcript

Jason: Hi everyone, this is Jason, producer of Data in Biotech. Before we get started, I wanted to let you know about our latest white paper. It’s a comprehensive guide to implementing machine learning models in biotech manufacturing. It’s a complete overview of all the potential problems of ML adoption and more importantly, how to solve them. To download it, simply visit connect.corrdyn.com/biotech-ml. We’ve also dropped the link in the show notes of this episode. Okay, let’s get into it. Welcome to Data in Biotech, a podcast from CorrDyn where we explore how companies leverage data to drive innovation in life sciences. Every two weeks we sit down with an expert from the world of biotech to understand how they’re using data science to solve technical challenges, streamline operations and further innovation in their business. This week we sat down with Brendan McCorkle, CEO of SciNote, a cloud-based ELN with lab inventory compliance and team management tools. During the interview, we discussed the importance of balancing flexibility and structure in the ELN space to accommodate the varying needs and workflows of different stakeholders, such as R&D scientists, process engineers, and data analysts. We also discussed the challenges and importance of data sharing and collaboration in R&D biotech organizations, striking the balance between flexibility and constraints in implementing ELNs, as well as the future of ELNs and the incorporation of machine learning and AI technologies. Here we go.

Ross Katz: Brendan McCorkle, welcome to the Data in Biotech podcast.

Brendan McCorkle: Thanks, Ross. Thanks for having me.

Ross Katz: Well just to kick us off, can you give us a brief overview of your career to date and what brought you here?

Brendan McCorkle: I came to entrepreneurship the scenic route, which is part of why it’s hard. Like a lot of my peers, I tried about everything else imaginable before ending here. I went to school to play baseball, not to do school. I ended up with a psych degree, was a paramedic for a while, and it took me a minute to realize that the thread between all of these — including being a terrible employee at a lot of these roles — was that I was deeply unsatisfied by the status quo, thought I was smarter than everyone else around me, and thought I could do anything about it as an entry-level employee. In hindsight, it took me a minute to put together what this shape really was, but I think this screams entrepreneur. I went to grad school to join my friends who were all software engineers from a former athlete psych degree. This was actually tricky to work for big tech and got a degree in engineering management. This coincided with the first financial crisis and I found myself educated the right way, no longer owing my employer at the time a few years afterwards, and no one was hiring. If I wasn’t going to have a job, I might as well start one. I started a company. The rest is history, but that was the first of a few. All of my companies have had a little bit of what I call a by-us-for-us product thesis. We built developer tooling by developers for developers. I helped spin out a healthcare company from Penn that was by clinicians for clinicians, and most recently before SciNote turned around a consultancy with a small EdTech product that was by teachers for teachers. Of course now SciNote is by scientists for scientists. I seem to have a thing, and this by-us-for-us foundation is near and dear. Now that I’ve done it a few times, I guess that’s what I do.

Ross Katz: Tell us about SciNote. What’s the company’s mission and how does it go about accomplishing it?

Brendan McCorkle: SciNote was started to improve science, which is wonderfully ambitious. Specifically, to take a piece of the scientific process which hasn’t changed since Marie Curie — taking notes on a piece of paper is still the magic backbone of the lab hundreds of years later, which is staggering when other parts of the world have moved into technology. The goal is to stop the data loss from losing a notebook or having it get damaged, and to facilitate collaboration between peers in a lab. It’s hard to do that on a piece of paper. It’s really an individual unit of work. Science has gotten bigger, but the technology powering science hasn’t. SciNote’s mission is to preserve the knowledge of science by eliminating places where it gets lost first, and then facilitating new knowledge generation.

Ross Katz: Just hearing you talk about it, I have a biased view because I come from a data perspective, but all of these problems sound like problems of having all the data available in an integrated way — where you have metadata about the experiments being run and the insights being generated, and all of that connects to the metadata that exists about the inventory you’re using.

Brendan McCorkle: You just mentioned two of the four FAIR letters in FAIR standards. The I is interoperable and very few vendors have an API. The A is accessible, and the F and the A kind of blend together. I just gave a talk where we discussed the air gap — cheekily saying that the A, I, and R of FAIR are at risk, because it’s very easy for vendors to actually make all three of those worse instead of better. As an industry trying to move towards FAIR data, we have to be cognizant of the choices we make in our product and distribution. The easiest findable, the easiest accessible is just plain text on the internet — which is horrifying when it comes to data security. One extreme is plain text on the internet, no way. The other extreme is a walled garden where it’s at least secure and findable if you use theirs, and it’s really tricky to balance in the middle. Other folks in the market have looked at this and gone, wow, it’s really hard to balance these two, and then they’ve chosen to leave that difficulty. We’ve been stubborn enough to say the market is telling us we can’t leave this point of difficulty. Our users are telling us you have to support our difficulty and we love them. Some of this balancing is really tricky. We’ve held onto that balance and said, too bad that it’s hard, we’re going to still support this level of difficulty because we have to.

Ross Katz: I want to spend a little time on the user perspective. One of the challenges is that in this integrated data ecosystem, you’ve also got stakeholder groups across the organization with different needs from the data these systems generate and different questions they want to ask about what’s being captured in a platform like SciNote. Could you walk us through what those stakeholder groups are and what each group typically cares about?

Brendan McCorkle: The easiest unit to talk about is the scientist at the bench. This is where we started and it’s the beginning of this journey. The introduction of tension starts here, because especially in R&D — we are disproportionately powering R&D scientists — there is a direct correlation between increasing flexibility and better science, and between constriction and the reduction of space to be creative. There’s a push towards being as flexible as possible. The first philosophical divide in the market is whether you say: because of novel science, testing new things, this is the hardest thing to structure because we don’t know yet — let’s let the scientist have as much flexibility as possible. The ultimate pure ELN is literally a blank piece of paper that’s digital. Paper on glass. Totally unbound. That’s the stuff of nightmares for the data team or the informaticists, but for the R&D scientist, it might be the promised land. Somewhere between maximum flexibility and closing it all the way down — which is what you want for manufacturing, where you actually want it to be the same way every time — you lose some creativity. We have decided that ultimate flexibility isn’t worth the pain it causes further down the line. We try very hard to be as flexible as possible, as opposed to as flexible as you could be. These are nearby ideas but they’re not the same. We have peers who’ve chosen ultimate flexibility and peers who’ve chosen no flexibility, and those self-sort into the R of R&D and manufacturing. This is a big philosophical divide. We are trying to balance these two, which lets us tell a story about R and D — and the D, development, is underserved in our market. When people say this is a tool for R&D, they really mean R. They don’t mean R and D. That’s where we’re getting pulled into a new part of the world where we have stakeholders who need something different.

A great example is someone at a big pharma company responsible for the transition from R to manufacturing. They’re testing whether they can do something at scale. There’s some structure now — it’s important to see if it works and track dependencies — but they’re still looping, so they can’t do fully structured yet. Also it’s biology, so we care about lineage. Mini soapbox here: inventories are not inventories, and this is one of the dangers in the conversation between ELN and LIMS. If you’re doing cell-based bio work, you care about lineage in your inventory. It’s really hard to do non-linear stuff, and lineage is a good example of non-linear stuff that doesn’t work in a table. We have to support somebody who’s starting to care about structure — that’s a process engineer moving towards the D of R&D — or a big enough organization that has a data scientist or informaticist who wants to do analysis. Here’s where it gets interesting. When we don’t pass the context of the science to the person doing the analysis and then loop the analysis back to the scientist, we’ve abandoned what industries like Toyota learned in auto manufacturing: allow people on the floor to pull the cord and say it’s off today. That’s an explicit endorsement that context at the floor matters. But if we don’t pass that context from the bench into the business unit, we’re saying Toyota’s wrong. We’re passing data out without context, and telling the person doing the analysis that the data collected by the person with context doesn’t matter. And if we don’t pass results back, we’re telling the person at the bench that the analysis isn’t important. Someone at the bench might know it’s off today, that there’s waste, that an experiment line should be shut down — or that this is working better than expected, we’re onto something. Or in the bioprocess, maybe we’ve found a more efficient way to do this. Both are great for the business of science, and we cut this line of collaboration by not allowing data to flow in both directions. Not willing to lean into this problem space means saying a whole cluster of terrible things: no, you can’t have that information; no, we don’t think your opinion matters. These are places where we try hard to make sure there’s a feedback cycle — scientist plans an experiment, runs it, someone does analysis, that feeds back into the next one. This is how science is taught. This is how science is done in every academic and federal lab. We must bring a solution that maps to this world. When people don’t support non-linear workflows, when they don’t support lineage in inventories — that’s not serving science. If you aren’t serving that shape, you’re making science worse instead of better. That should lead to lower revenue. We should be intentional about the choices we make in this problem space.

Ross Katz: I love what you’re talking about in terms of moving from people who think of R&D as research only into development — and that an ELN platform integrating across the different aspects of a biotech data ecosystem needs to lead in that direction. Can you provide a more concrete example of what the interfaces look like for the research scientist versus people further up the line in bioprocessing or development, and how they interact with the system differently to accomplish their different goals?

Brendan McCorkle: You can slice the first one a bunch of different ways, and this is part of the challenge as a software vendor. A biologist and a chemist may want different workflows. Even biologist in lab A and biologist in lab B may have different workflows. There’s a balancing act where, on top of everything we’ve been talking about, you have huge variance intra-company and inter-company — between different locations or a collaboration between public and private companies. You have to cast the net wide enough to reach the “invisible software” goal, where when somebody goes to reach for the tool it’s just there. I always picture the auto mechanic sliding under the car and reaching out — the tool they need next is there and they can’t even see it. One piece of this is allowing people to enter data as many ways as possible. Someone in a clean room can’t write anything down — maybe they need to talk into the system. Maybe you have a really old scale that’s not digital. Maybe you’re taking a picture with your phone of the scale output. Or there’s someone on your team who refuses to give up paper — you might be taking a picture of a paper notebook, but at least you’re capturing a record. You can tag it, have metadata at all. You know where it was, who it was, when it was. All of that moves towards some data truth. There’s a bit of post-processing too: from the picture into the system of truth, knowing who, when, where, what — can you OCR something that’s a picture so it ends up in the data fields? That’s the vendor’s responsibility. If you entered it manually in an open text field, someone else uploaded a Word document, someone else took a picture — if those can all land as the same thing in the data set, you’re not offending the people downstream. You may be normalizing differently depending on the input vector, but the goal is to let data entry be as open as possible while trending the same things to the same values. That way when we pass it to someone doing analysis, we don’t have the equivalent of an Excel file with 19 different headers. T-test with both capital Ts, T-test with a lowercase t, a space before the hyphen, no hyphen — these are all a T-test. The data scientist has significant work to do if they get 19 versions instead of one. This level of constriction does not change the scientist’s ability to do science. It’s not going to change whether they use different amounts of acetone. It’s still a T-test. The software can know this if it runs regularly — it’s not AI, it’s low-level technology lift to normalize a data set so someone else doesn’t have to. That’s where this bridge happens. A photo with automatically applied metadata, a prompt saying “you ran this protocol before, would you like to again?” — that saves five clicks and three data entry fields. If one person wants to work off a checklist and another wants a static list or a picture, and we support both, we have two happy users. It’s a juggling act, but the goal is flexible UI elements that mirror flexible workflows so that when people are moving through the workflow, they’re not thinking about the steps — they’re doing science. We have to support a bunch of different entry points and exit points.

Ross Katz: It strikes me that there are myriad problems to solve across myriad stakeholder groups, and you mentioned upfront not constraining the creativity of the science — you’re trying to enable the scientist — but at an organizational level, at a laboratory management level, you want to apply some constraints. How does that work?

Brendan McCorkle: ELN vendors are making it worse unless we do something differently. That’s the rub. If we go to ultimate flexibility, we are making data less FAIR. We’re upsetting people downstream and up the hierarchy too. It’s cliché, but software vendors talk about giving a dashboard to management as though that’s the whole thing. If you don’t have really good data fidelity and interactivity between these pieces, you might not even be able to populate a dashboard because it’s that much of a mess. The scientist wants to buy this software to do their job, but they have to convince their boss or procurement. Part of the enterprise sales motion is: what’s the view for the boss? The head of research operations, the informaticist — a very important and hard to employ individual in our space right now — they should be happy with it too. What do they see? Is it tucking nicely into their workflows? Those workflows might be completely different. They want to look across the whole portfolio of R&D. They don’t care about the protocol inside that experiment. They want machine utilization, scientist utilization, is anybody in the lab when they said they would be? Why did we use three times as much acetone as we used to for this? Are we doing a new experiment or did something move that we need to look into? This is somebody’s job and they need to take that zoomed-out view. It’s tough to see that fully stitched-together view of data in the lab. The first step is realizing the consequence: all of these fronts are places for data to get lost or for context to get lost, which makes integration harder. At least two of the letters in FAIR data get harder fast without thinking about this problem space.

Ross Katz: How do executive stakeholders at companies who work with SciNote evaluate whether the system is doing what they need it to do?

Brendan McCorkle: I think we want this answer to be different than it is. The warm and fuzzy version is that executives really want their scientists to be happy — scientist, use whichever tool you’d like. Unfortunately, that’s not always how it works. There are certainly employers with very high eNPS scores partly because they do that. There are also places more interested in the velocity of things through their R&D pipeline, and they’re aware of the cost of loss. The business case is a blend of efficiency and loss reduction. It’s a strong proxy to healthcare, where better handoffs between doctors on shift change lead to better patient outcomes and fewer adverse effects. The business case is actually the lowering of the bad stuff. Fewer false positives in the R&D pipeline. The longer a false positive goes forward, the more painful it’s going to be to remove. Faster things through the pipeline, less time spent on non-science by scientists. The scientist might say I love using this, but what they’re saying also includes: we’ve seen people spend nine hours less a week on non-science. Some of these places have a thousand scientists. 9,000 hours of non-science — that’s roughly 250 FTEs of work that’s now freed up. They can do more experiments, the team could be staffed at 80% of what it used to be with the same output. There’s a lot you can do with 250 FTEs worth of time every week in pharma or biotech. The scientist says “I love this,” and the boss says “I love the impact of this reduction” — and that’s just time. Not the reduction of replicated studies. Not higher fidelity data into the analysis machine. Not faster pipeline velocity with fewer false positives. All of those get more science done, better science done, less wasted time. To the business, that’s less wasted money and faster velocity. They’re in a competitive space. If you get your stuff through 10% faster than your peer, maybe you show up first. That’s a really big deal.

Ross Katz: That makes a lot of sense. From a data team perspective, what are the workflows you see in practice in terms of how data teams integrate with SciNote and how does SciNote support those teams?

Brendan McCorkle: This is a little bit of a work in progress, because a lot of times we don’t see this workflow — the systems aren’t connected well. I can give some anecdotes. We have a customer, Gatehouse Bio. They’re uncommon historically but becoming more common, in that they have biologists, microbiologists, engineers, and data people. They’re a proper modern organization. They have internal resources who are frustrated with each other because it’s a relatively small company — not a thousand scientists, more like 10 or 100. Small enough that you know everyone on the team, which maybe makes it more acute when you’re doing things that bother your teammates. The T-test problem I mentioned comes from their data team: we need an ELN because our scientists need to be aligned, we need to collaborate, and can we put some bounds on what goes in there? The analysis gets even more complicated because they do high-throughput, high-resolution microscopy, and they have a machine generating petabytes of data every day. Most companies talk about big data but really just have bigger-than-Excel problems. These folks have proper petabytes. That’s hard to move over the internet. Their vendor has a dedicated software system for it — and part of the solution is acknowledging that the ELN is the wrong place for petabytes of microscopy data. It’s putting a square peg into a round hole. But now there’s another player in the shape. People at the company need to access those images quickly, the data has to live somewhere, and you can’t trivially send that much data over the wire every day. The answer might just be a low-res snapshot for a log, or a link to the other vendor’s site. That’s a good example of playing nicely with other parts of the workflow. We don’t need to own the workflow of that data vendor — their use case has already said, the system of record is over there for this part of the lab. We just need to make sure you’re aware it happened. Very light. Not petabytes — maybe kilobytes. It doesn’t even need to be a picture, just the line “images were run today.” That might be enough for the rest of the team. Part of the magic is we have to support these workflows even when we don’t fully know what they are. That changes our answer from having perfect context to preparing for those workflows. The state of the market right now is that so much of the conversation about AI, ML, and informatics is really about preparing the data for that work before actually doing it. We’re catching you at peak “it’s coming” — the gap between the bench workflow and the data workflow is the problem, and we’re collectively working on closing it. This work is typically done somewhere else and not looped back into the ELN. It’s only just starting.

Ross Katz: That’s great. You mentioned integrating with the microscopy platform to get URLs out of that system into SciNote. What are some of the other integration points you view as crucial for bringing all of that R&D work into the SciNote platform?

Brendan McCorkle: Earlier I mentioned a fictitious Pfizer with 15 LIMS and 7 ELNs. But we can talk about real Agilent — they have their own ELN, they built it themselves, and then they also brought us in. There is no world where there’s only one system. If I step back, there are a few categories. ELNs are one of them — we are human-entered, human-generated data. There’s also system-generated data and machine-generated data. When I talk about other ELNs and LIMS, I’m talking about other software. Machine-generated would be the scale or the microscopy. There are a lot of machines in the lab and some of the more modern ones have the ability to be data-enabled. This was part of our origin story — to try to combine at least two of these three. All three pull into a place and that place needs to talk to legacy data and other vendors. The largest buckets are human-generated data, system-generated data, and machine-generated data. Then there’s the downstream — towards the data scientist, towards the informaticist, who probably has a BI tool and is running data pipelines in another system. They need to ingest data into that system and export results back. Even within human-entered data, some people work in Excel or Google Docs and want to pull that in. Take the scale picture example. The scale is almost trivially small — not the cost or headache of a petabyte microscopy machine. But we could prevent them from entering a negative weight, which is never going to happen in reality. That’s like the T-test example: a small change that messes up downstream data badly. You start running an average on weights and you let a negative number in — no more confidence in anything that comes out, and someone has to go back and chase that all the way to the source. Then when you get into the business systems — scheduling machines, utilization analysis — the closer you get to manufacturing, QA and QC start showing up. They have systems. Maybe they’re already using a company like Ganymede, they have an SDMS, a data management platform. We’re also a data management platform, so where’s the boundary? It’s fine — we can use both — but we have to talk about using both. If everybody is friendly, that’s a beautiful thing. We have integrations into other software, into machinery, into business systems. Maybe there’s an accounting system for tracking acetone inventory. Maybe there’s a quality system. Maybe it’s a fully validated lab where they have to track IQs and OQs and validate the system itself. Did anything I just mentioned change? Now we have to document that it changed and re-validate the whole thing. In validated labs, you now have another axis: we can’t do anything quickly because all the processes are validated too. We have to operate not like a modern software company but like NASA — you can’t push a change without knowing the downstream consequences, you do one release a year because of validation. We have to respect the validation constraints of our customers. This is a process integration, a people integration — if the lab is validated, we have to act like it. If we push a change without telling them, everybody’s world is pain.

Ross Katz: What advice do you have for R&D organizations or biotech companies more broadly who are considering implementing ELNs as a category or SciNote in particular in their workflows?

Brendan McCorkle: It’s important that someone making this decision hears from someone on my side of the table: do not accept the answer that it can be all us. This is hypocrisy at best, hubris at worst. It’s not true. There is a profound economic motivation for a software vendor to say this as part of their sales cycle, but there’s a responsibility for us to say publicly, no. If you’re shopping, this idea of meeting where you are — the responsibility is to let us show you how we connect to where you are. Don’t let go of where you are. It’s your way, not our way. Don’t be shy about the fact that you probably already have a system of record. What we need to know from our side as you’re evaluating us: what are you already considering as your system of record? What are the stakeholders that need to know? And I mean people, machines, software, and process — validated labs included. What are those four pieces in their current state and where are you trying to get them to go? Look for movement from where you are to where you want to go. It’s probably more than one piece of software and you probably already have some of them. Don’t succumb to sales people telling you otherwise. Push on connection points: how do I make sure I’m navigating with vendors who understand where we are right now, who are friendly with their peers and our own legacy software? Look for companies with APIs. Look for companies with partner ecosystems. These are signals. If they have a partner ecosystem and an API, you have confidence before even talking to any of us that they understand what I just said — that they can’t go it alone and can’t change everything. Stay strong, because you know more about your stuff than you think you do. Look for us to tell you how we can help, not how we can change you.

Ross Katz: What new needs do you see emerging for scientists or data teams, or biotech organizations more broadly, in terms of their expectations from an ELN?

Brendan McCorkle: I heard a talk last week where we spent a lot of time on the elimination of paper in the lab. There was a gentleman running informatics at a large pharma company saying we should start talking about the second phase of the journey towards digitization. Phase one is eliminate the paper. I agree — step two is the elimination of Excel. What we introduce in successfully eliminating paper is bench scientists learning how software works. That’s exciting, but also a little scary, because Excel is actually not the right tool for the job. It’s the tool they’ve been using. But just like paper, once upon a time paper was the way — and if the movement away from paper is right, then paper is wrong in view of modern technology. Excel is really tricky to do securely. You step out of the machine, do some work, bring it back in. What changed? Who changed it? You asked me earlier about what the bosses care about. When Bob leaves the company and maybe doesn’t give his laptop back, what data are we losing? Do we have access to all of his research? It was easy when all the lab notebooks were locked in a room down the hallway — at least we knew he didn’t leave with one of them. Now it’s digital, but what’s the modern equivalent? Can somebody else pick up his work? Do we have access? Can we keep the machine running? Same problem with Excel as data sets get bigger. The more we’re moving towards AI and real data analysis, the data sets are getting too big for Excel. You now need an interface into a database. Excel is one of those, but it’s a very light one. People don’t appreciate how light of a layer it is. Three labs, 50 people each, a million rows, a dozen columns — it’s not Excel time anymore. Science hasn’t internalized what the development side of the world already knows: LOL Excel. The next thing to walk people through is that it’s okay to do computation in a table that’s not Excel or Google Sheets. It’s okay to have an interface into a database. First we get to “okay,” then we get to “better.” Excel is coming into the crosshairs. That’s coming.

Ross Katz: As a last question, what are your perspectives on incorporating machine learning and AI into the SciNote platform? Are there any future features or integrations planned?

Brendan McCorkle: Much of what’s happening now is really more ML than AI. It’s not as sexy to say that because AI is the new hotness. But there’s some really powerful stuff that looks more like search suggestions — “hey, do you want to run that protocol again?” That’s really impactful as a user, but it’s ML not AI. There is some cool stuff happening in industry right now and most of it’s actually ML. We do some of this. For us, it’s “advanced search” — how do you look at a lab that has 5,000 protocols built up over years? A new scientist who comes on looks at 5,000 and goes, what am I supposed to be doing today? ML is the antidote to that, not AI. My AI answer is that we’re preparing for AI and a lot of the AI work is likely going to happen inside their business, not inside ours. With one exception: we actually did launch an AI tool a couple of years ago. We launched a feature called Manuscript Writer three years ago. It wasn’t using ChatGPT — that was still a glimmer in someone’s eye. We were taking all of the context we had about your work and populating materials and methods, the summary, linking to the research you were doing. And connecting this to advanced search — suggesting other research similar to what you’re doing rather than ranking by popularity. Web ranking is based on popularity, not relevance. Having the context of your actual work driving the results is really interesting. That’s what AI and ML look like in an ELN: taking automated tasks off your plate because of the context of the other pieces you’re doing. The reason this didn’t fully work at the time is that companies have 20 or 30 years of research in PDFs or actual notebooks locked down the hallway. That data has to get into the LLM. That LLM also probably needs to not be ChatGPT because of trade secrets. Both of those were enough tension to stop this a couple of years ago. Looking forward, there will be — probably more than one, because that’s how markets work — private-friendly LLMs for biotech and life sciences where a Pfizer or Novartis would look at it and say, I could actually accept using this one. Not the current shape of ChatGPT. Maybe run locally or an industry-specific one tailor-made for the security and trust constraints of biotech. That’s nearby — someone’s probably working on it right now. And ML and AI are also powering OCR solutions. As those get more accurate and cheaper, the cost of getting data into a system at all gets lower. If both of those pieces play nicely, we revive Manuscript Writer — accurate enough, auto-tagging metadata, trusted. That day is soon, not today, but soon.

Ross Katz: Brendan, this has been a fascinating conversation. I really appreciate you making the time. Thanks for coming on and we’ll plan to see you down the line.

Brendan McCorkle: Awesome. Thanks, Ross. Super appreciate your time and the opportunity to share my soapbox.

Jason: And that’s it for this episode of Data in Biotech. If you enjoyed the episode, please subscribe, rate or leave a review in your podcast platform of choice. See you next time.

Frequently Asked
Questions

How does a properly implemented ELN improve R&D efficiency and pipeline velocity?
An effective ELN significantly reduces time scientists spend on non-science tasks, such as manual data entry or searching for information. By centralizing data and facilitating collaboration, it can save thousands of hours annually, allowing scientists to conduct more experiments, reduce replicated studies, and accelerate breakthroughs. This directly translates to faster movement of discoveries through the R&D pipeline and reduced operational waste.
What are the key data integration challenges when deploying an ELN in a complex lab environment?
The main challenges involve integrating human-generated data (from scientists), system-generated data (from other software), and machine-generated data (from lab instruments) while maintaining context. SciNote addresses this by offering robust APIs and actively working with existing legacy systems, rather than demanding a single "all-us" solution. This approach ensures data from diverse sources can flow into downstream analytics tools without significant manual normalization.
How can an ELN bridge the data structure needs between early-stage R&D and eventual manufacturing scale-up?
ELNs must balance R&D flexibility with the increasing need for structure in development. SciNote supports this by offering customizable interfaces that maintain context while normalizing data, allowing for lineage tracking crucial in biology-based work. For validated labs, it also requires process integration and rigorous documentation to ensure compliance, treating changes like a "NASA operation" to avoid disruptions.

Need a data partner for life sciences?

CorrDyn helps biotech and pharma companies build the data infrastructure that accelerates research and operations.

Book an intro call