Skip to content
Data in BiotechEpisode 34

Unlocking the Power of AI in Microscopy

Ilya Goldberg (CSO) and Reese Findley (AI Data Scientist) of ViQi on applying machine learning to automated brightfield microscopy for life sciences.

47:32Full transcript below
IG

Ilya Goldberg

Chief Science Officer at ViQi

RF

Reese Findley

AI Data Scientist at ViQi

Overview


Phenotypic screening has been stuck on a manual bottleneck for decades: a biologist staring down a microscope, counting plaques or annotating cells one field at a time. Live-cell dyes and faster automated imagers widened the signal net, but analysis never scaled with acquisition. Dose-response studies that layer thousands of compounds across multiple cell lines and time points generate more images than any lab can interpret by hand. AI trained on controls plus compound-by-dose variation now clusters those phenotypes quantitatively, and brightfield microscopy lets the same plate be imaged repeatedly without killing the cells.

Ilya Goldberg (Chief Science Officer, ViQi; previously ran a hybrid computational–life science lab at the NIH for nearly 20 years) and Reese Findley (AI Data Scientist, ViQi; PhD in neuroscience studying olfactory search in mice) join Ross Katz to walk through how ViQi builds image-based assays. They cover AutoHCS, their automated high content screening tool that trains AIs on controls, compounds, and doses in a few hours. They explain AVIA, a viral infectivity assay that detects cells producing virus long before cell death and works on any virus once trained. They unpack why training is per-lab and per-imager, how ensembles of stochastic AIs outperform averages, and why a collaboration with Saguaro Biosciences on non-toxic live cell dyes opened up time-course analysis that was previously prohibitive.

Key Takeaways

Brightfield plus AI replaces endpoint assays with time courses

Fluorescent probes require fixing cells, which means one measurement per experiment. Brightfield imaging of live cells with non-toxic dyes lets the same plate be imaged at the start, middle, and end of an experiment on a robotic stage, so kinetics come effectively for free. Reese Findley described this as the biggest unlock in AutoHCS: drug efficacy over time is visible in ways a single endpoint cannot capture.

Train per-lab, per-virus AIs in hours by layering on controls

ViQi abandoned the idea of a universal infected-cell classifier once it was clear that different viruses, cell lines, and even two microscopes of the same model produce images an AI can distinguish. AVIA and AutoHCS instead train fresh models against in-plate negative and positive controls, using transfer learning from pre-trained CNNs like EfficientNet for AVIA and feature-classifier ensembles for AutoHCS. Training that once took weeks now finishes in a few hours per time point, even across a 1500-compound screen.

Automated phenotypic profiling still needs a biologist to validate clusters

Clustering is stochastic — retraining the same AI on the same data produces a different model every run because GPU race conditions randomize the training path. ViQi uses voting ensembles to stabilize answers and then hands results to a bioinformaticist who cross-checks clusters against open-source pharmacological data. Compounds that land in the wrong cluster are the interesting ones: likely off-target effects, and the reason drug development fails most often.

Full Transcript

Jason: Welcome to Data in Biotech, a podcast from CorrDyn where we explore how companies leverage data to drive innovation in life sciences. Every two weeks we sit down with an expert from the world of biotechnology to understand how they’re using data science to solve technical challenges, streamline operations, and further innovation in their business. In this episode, we sit down with Ilya Goldberg and Reese Findley from ViQi, a company at the forefront of high-end imaging and AI solutions for biotech. They share the story behind ViQi’s beginnings, dive deep into their groundbreaking imaging technologies, and how machine learning is revolutionizing the analysis of cellular images. We cover advancements in live cell imaging, lab automation, and how AI is being trained for viral infectivity assays and drug discovery. Ilya and Reese also talk about the future of AI in life sciences and its potential to transform lab workflows and accelerate scientific breakthroughs. Here we go.

Ross Katz: Ilya Goldberg and Reese Findley, welcome to the Data in Biotech podcast.

Ilya Goldberg: Hi.

Reese Findley: Hi.

Ross Katz: To kick us off, starting with you Ilya and then to you Reese, would you give us a brief introduction to your background and what brought you here today?

Ilya Goldberg: My background is computational and life science. I have a PhD in life science. I ran a research lab that was a hybrid at the NIH for almost 20 years. I joined ViQi about six years ago. Before ViQi, I did a startup in medical devices that used AIs to predict cancer risk in lung nodules from CT scans.

Ross Katz: Reese?

Reese Findley: My background is in neuroscience. I did my PhD in olfactory search in mice, studying the neural correlates. I now work for ViQi as an AI data scientist and I do drug discovery and phenotypic profiling using automated AI tools.

Ross Katz: Since you’re both at ViQi and we’re here to talk about ViQi today, would you all give an introduction to ViQi and the kind of work that you do?

Ilya Goldberg: ViQi started in the late 90s as a open source project at UC Santa Barbara, the computer science department, as a collaboration with several biology departments, to address high-end imaging needs. Multidimensional microscopy like 3D, multi-channel, large sets of images. It was a project to build infrastructure to support this and analysis.

Ross Katz: It sounds like a lot of what ViQi does is related to imaging and using imaging and machine learning for different types of problems that you’re trying to solve. Can you give us an overview of the way that imaging and machine learning or AI are applied?

Ilya Goldberg: Imaging is really central to the company since its beginning. A lot of us started using AIs to solve imaging problems in the early 2000s. We started building our first AIs to do this. It was readily apparent early on, even with very primitive AIs — we were using perceptrons which are not deep neural nets and other conventional classifiers — we could outstrip experts in imaging. People who interpret images professionally every day, even that primitive of an AI could surpass what they could do very routinely. That told us, holy cow, you can do a lot with this. Some of us started way back then exploring this space and moving into not using stains, so using brightfield microscopy so you’re using a much wider net to measure things about images.

Ross Katz: You’ve been dealing in both imaging and AI for a long time now. I’m wondering, where do you see the combination of imaging and AI as having made the greatest contributions to biology and to the types of work that you do?

Reese Findley: In the broader scheme of biology, technological development has been booming in several areas. We’re seeing live cell dyes so you can capture time courses. We’re seeing better resolution automated imagers, faster automated imagers, and then these AI developments. I think the biggest question being answered right now is what can be automated and what can’t. That’s what ViQi does really well, is what is appropriate to automate and what is very difficult and really needs that subjective expert science opinion behind it. That’s where the most advances have happened in the field in general, taking those areas where you can automate and shortening that pipeline down, reducing massive bottlenecks that have been done manually for decades down to hours that an automated imager can do with an automated analysis system.

Ross Katz: That’s a really interesting point and I would love to hear from you, what have you learned about the boundaries between what can be automated and what can’t be automated in terms of the type of work that you do?

Reese Findley: In my work specifically, I work on our assay, automated high content screening. One thing we do is a type of phenotypic profiling where we train an AI on a bunch of different compounds that have been applied to cells, and we use the distances to create this distance-based dendrogram. All of that works really well. We usually get really cool biological meaning out of it. However, it demands a biologist to look at the final product. The reason I say that is because this clustering is subtle and stochastic, and anybody who’s worked in clustering — I used to cluster neurons — anybody who’s done that knows that different people can get very different results depending on the analysis methods they’re using. It’s a huge problem, especially in sensory neuroscience, so that requires that expert eye, some bioinformatics validation of our clusters. There’s a lot of advanced science that goes into interpreting the end result of our automated analysis.

Ross Katz: I’d love to dive into AutoHCS in just a moment, but before I do, can you help orient me to how ViQi partners with clients or customers and what’s the type of work that you do in support of the biotech ecosystem?

Ilya Goldberg: We solve two different problems. One of them is related to your last question about automation. One of the big things we have to deal with that’s not automated is the interface. You have an imager in the lab, and then you have all these high-end tools that are online on the cloud to analyze this. Every time you move this stuff around — your experiment down to an imager, an imager’s images up to the cloud — there’s an interpretation. What does all this mean? Adding metadata to your images, specifically dealing with image formats. Everybody’s got their own image format in life science. There’s over a hundred of them. Dealing with that is a lot of where automation breaks. For instance, describing your experiment. The imager knows how it took the image, so you can query the imager what’s the exposure time, etc., but what compound is in each well of a plate, that’s up to reading it out of a spreadsheet or sometimes out of a screenshot of a spreadsheet. All of the various ways of breaking that have been done and that’s where we spend a lot of time.

Ross Katz: I should think of this as assay as a service, that a life science practitioner or a biotech organization recognizes the need for an assay and they believe, and I guess…

Ilya Goldberg: We call it that. It has an unfortunate acronym but we do call it that. Yes, it does. But I think the question that comes to mind for me is how does someone know that the type of assays that ViQi provides are the assays that they’re looking for versus the other assays that they might develop in-house or the assays that they might be able to contract with other people to create?

Reese Findley: Anybody who’s worked in labs for a long time will tell you there’s so many steps that are just aggravating and frustrating and make you want to pound your head against the wall and some of us pride ourselves in doing those. That’s a scientist-like thing. Mine was calibrated airborne olfactory stimuli. Super hard. But anyway, there’s something that’s absolutely aggravating that takes a super long time and, to be honest, I don’t think it crosses most people’s minds to try to automate those steps. What they want to automate is their image analysis afterwards and that’s what they’re talking to their automation teams about. What ViQi’s doing all the time is trying to identify those hair-pulling steps that are taking too long and causing too many issues and find a creative solution to automate that out so people can do the fun part of science. When you ask what are people coming to ViQi for, it’s those creative solutions because we’re spending all of our time thinking about those as opposed to trying to do it while executing bench experiments and meeting deadlines.

Ross Katz: That makes sense. Before we dive into an automated viral infectivity assay, can you talk about what are the uses for a viral infectivity assay broadly?

Ilya Goldberg: It’s a surprisingly broad use, but anytime you have a research lab or manufacturing or vaccine development, you need to measure how much virus do I have. How much infectious functional virus do I have? In vaccines, for instance, you’re looking at a population that’s been immunized with a vaccine and new variants come up and you want to know, this population, are they resistant to these new variants? You need to grow up your new variant, hopefully before it becomes pandemic level, and collect blood samples from your patients and measure the inactivating antibodies they have in their serum that inactivate the virus. You’re measuring how much virus do I have left after I treat my starting virus with a vaccinated person’s serum. That’s vaccines. Of course, to get to vaccines you have to manufacture a lot of virus and during the manufacturing process, you’re growing up cells, normally viruses are grown in cells. You need to measure how much did I get, because it’s all biology so it’s different every day. Things fail, you need to measure that as early as possible because it’s a weeks-long process to grow this stuff up in vats and purify it so that you can inject it into people. Huge number of steps all along the way from R&D to normal manufacturing process, you’re always running these assays. In R&D, you’re studying viruses, you need to know how much virus there is. It’s an absolutely ubiquitous assay in virology.

Ross Katz: That makes sense. What are the traditional methods for conducting this type of assay?

Ilya Goldberg: Normally those methods rely on enough cell death that you can detect it easily. Viruses normally kill cells. If you have cultured cells living on a lawn on a plastic dish, this is a plaque assay — basically you dilute your virus enough so that you have a scattering of individual viral particles that are infectious land on individual cells. Then you wait long enough for the neighboring cells to become infected so that you get a little ring of cell death, a hole in your continuous lawn of cells. That’s a plaque and then you sit there and count them. What I mean is why am I sitting here counting plaques, shouldn’t I take a picture of this and get some software to count them for me. You’re solving the very last step of a problem that you could have solved a lot earlier. Why are you waiting all this time for multiple rounds of infection? Why rely on cell death for your output?

Ross Katz: Interesting. In the traditional plaque-based assay, the images come at the end and then there’s a counting that occurs whereas with AVIA, if I understand correctly, one of the things that you’re doing is you’re taking images at the beginning, the middle, and the end and…

Ilya Goldberg: We’re taking them at the beginning. Instead of detecting cells when they’re good and dead, we’re going to detect cells that are producing virus. Much earlier in the cycle. Cells have to be healthy in order to produce virus, otherwise they’re not going to do it. They’re very much alive and healthy and cranking this virus out, and viruses have evolved to hijack cellular process to make copies of themselves instead of the normal business of the cell. That changes the way the cell looks. Those changes are very subtle. There’s a lot of variation in how cells look, so it’s a perfect problem for an AI to solve because there’s a huge amount of variation — a noise filter is what you’re relying on it to do. There’s a lot of information, a lot of it’s irrelevant, you need to focus the AI in: this is what infected cells look like, this is what uninfected cells look like. Now go and count through these images and count areas that are infected versus not. That’s what the assay’s based on. The other thing that people normally do with microscopy of cells is they use antibodies or other molecular probes that highlight very specific molecules in these cells with fluorescent dyes so that you can detect infected cells because you’re detecting viral proteins that normally don’t exist in cells unless they’re infected. That’s another kind of assay, but you don’t have to do that with AVIA because there’s no probes, no specific detection of specific proteins, you’re broadly just asking cells that are infected look different from cells that aren’t. This is something an AI can tell very easily and you can look at these images and if you can see anything, you certainly are not going to be able to quantify this and turn it into a count.

Ross Katz: Going back to what Reese was mentioning earlier, it strikes me that the methodology that you’re approaching, yes, you have the images that you’re taking and then you have the AI that’s consuming the images that can detect very early these very subtle changes in the cells that indicate that virus is going to be produced, but then there’s this automation and design of experiments component that allows you to feed sufficient information to the AI model that it’s able to detect and understand the level of virus that’s happening inside the cell based on it. I would love for you to explain the methodology of how that comes together.

Ilya Goldberg: By design, this is tied into high content screening, so it’s a good segue into that also, because it all relies on equipment automation. People have developed automated imagers since the mid-90s or so. We were there early enough where we went to Olympus and Nikon and all these microscope manufacturers and said, hey, what you should do is have a robot essentially move this plate around, do autofocus and image cells. And people were like, the hell are you talking about? People use coverslips and sit there and meticulously manually image these things. The first microscopes for high content screening were made by hand to put robotics X, Y, Z positioners on these plates. These plates are standard form factored. They’re this big, 3x5 inches…

Ross Katz: 96 wells typically, 384 wells, exactly…

Ilya Goldberg: Exactly. That’s the same format except now you’re growing cells in the bottom of each well and putting drugs on them or a genetic manipulation library. That’s the base of the experiment. Now you need to acquire millions of images and potentially a stack of plates that you want a robot arm feeding into your microscope and taking them out of the incubator and putting them back. Autofocus is a key component of this, somewhat precise stage positioning to get to the right places in all these wells. That’s the automation. Upstream from that, if you have a chemical library with a million compounds, you have robots already that deal with pulling compounds out of the freezer, arranging them on plates, plating cells, all of this. That’s already existing. Part of the development of AVIA was very specifically to plug it into a HCS infrastructure. You don’t have steps that are not easily done by robots — a plaque assay requires an overlay so you have to pour a gel over this thing. That’s really hard to handle because it has a narrow temperature range. Anytime you add a step — you need to stain the cells or wash them — when you’re dealing with dozens or hundreds of plates, it’s a huge multiplier for how much time, effort, reagents you need. All of that was a process of what can we get rid of to increase automation? Because automation is really the key to this.

Ross Katz: Because of the speed and efficiency with which you can get the results. You can get the result in a much shorter period of time and at much lower expense than you could otherwise.

Ilya Goldberg: Not only that. Those are really important properties of automation. Another thing that a lot of people skip is precision and reproducibility. Which you get from automation as opposed to people manually pipetting and adding reagents. Pipetting mostly. If you can get robots to do the pipetting, you get better precision. Because our tools are so precise, we can routinely detect differences in people’s pipetting techniques and recognize that this person is typically low readout on the assay whereas this one is consistently high or they have different levels of variation because people are different and with robots you can eliminate a lot of those differences.

Ross Katz: My understanding is that you have this automation that’s taking images of each of these wells and what you’re putting in each of the wells is designed to provide the controls and provide the degree of infectivity that characterizes the distribution that the model needs to pick up on.

Ilya Goldberg: It’s important to think about two stages. The first stage is training. You have to train the AI. You don’t need to train it in every experiment, but if you train an AI that can be used across experiments, in life science, every time you run an experiment things are slightly different. You have to train your AI across these differences, not just cell to cell differences but experiment to experiment differences. You can do this and that’s a reproducible AI. Now that you have it, you can feed it an experiment. Very different modes of operation for training and for processing. It’s important to recognize that because you’re doing things very differently. In training, you need a lot of repetitions, a lot more data, because you need the AI to be exposed to a lot of examples. When you’re processing, you need a lot less data because the AI will give you an answer for each individual little piece of an image or each individual cell.

Ross Katz: Makes sense. My understanding is that one of the benefits of AVIA is that you can apply it to any number of viruses that might require a viral infectivity assay and so it’s broadly applicable. When you talk about the difference between training and inference, what’s coming to mind for me is are you having to go through a separate training process for each type of virus that you’re running through the assay?

Ilya Goldberg: Yes. Exactly. That has to do with automation of AI training, because instead of trying to come up with a universal AI that recognizes infected cells, different viruses and different cells have very different interactions and they produce very different phenotypes. We’ve dropped that whole idea and said no, we’re just going to train an AI for each lab. Even if we work with this kind of virus in this kind of cell line, we’re going to train in your lab because there are enough differences lab to lab, certainly different people use different imagers. The optics are different enough from even the same model imager that you can detect — if you’re asking can I train an AI to detect the same model imager in the same lab, just two different models, of course you can. An AI can be trained quite easily to tell you which imager an image came from. Those differences are definitely baked into the data and if you’re trying to universalize your AI, you have to train it across the variation you’re going to see. Otherwise it’ll focus on the wrong things. That’s one of the caveats of dealing with AIs — you can very easily get yourself into a situation where the AI is trained on absolutely the wrong thing. We typically refer to that broadly as bias and that fits with how the general public perceives bias in AI. The source of the problem is the same.

Ross Katz: My understanding is that the primary modeling methodology that you’re using for the assay that you’re talking about here are convolutional neural networks or CNNs. Is that correct? And what is the application of CNNs to the way that you do this work?

Ilya Goldberg: We’re AI agnostic. We use different CNNs, but we also use more traditional AI techniques that are called machine learning and classifiers before they were neural based. It’s kind of split just because of how things started. AutoHCS tends to be based on features and classifiers, image features — we’ll get into what image features are — but it’s more how people approached this in the 2000s before CNNs became very practical. AVIA, on the other hand, uses modern pre-trained CNNs trained on other imaging problems unrelated to cells or viruses. While you’re training we are all training these AIs by clicking on images of stoplights and bicycles and cats and dogs.

Ross Katz: Are there any particular model families of CNNs that you’ve found transfer learning is really good for biological images?

Ilya Goldberg: We tried a fairly broad array of these. We tend to use EfficientNet in various versions, trained initially on ImageNet data and then transfer learning. We typically do direct transfer learning to infectivity problem from weights loaded from image discrimination kind of problem.

Reese Findley: It’s also important to note that we’re AI agnostic and we’re pretty willing to try anything, but we do optimize the AI we use based on the experiment we’re running. We’ll get into it, but with AVIA, we tend to want binary answers about little tiny tiles, and that works really well with CNNs. Whereas with AutoHCS, we’re actually using the AI’s confusion to give us information about the biological status of an image. That’s why we’re using feature classifiers — they’re more willing to be confused, whereas CNNs are snappy, they have very specific answers they want to give and they’re confidently correct or confidently wrong. We’re always doing development, always testing how different AIs respond to our experiments, but we find that there is optimization to be had depending on what type of question you’re asking.

Ross Katz: Interesting. How do you think through the selection of the players in the ensemble when developing an assay like that?

Ilya Goldberg: That’s an important feature of AIs that people don’t appreciate enough, is that they are stochastic. If you train an AI with the same data set and then ask it the same question, you’re going to get different answers. Keeping everything exactly fixed. That is because during AI training there are processes that are randomized. In a GPU, for instance, all of these things are happening at the same time, that sets up zillions of race conditions. How the race is resolved is essentially random. You can do this in a completely deterministic way, but you’re essentially serializing the process. It makes AI training impossible. Every time you train an AI with the same data, you’re going to get a different AI. That’s one reason why you would have essentially a voting kind of model for an answer that is subtle enough, because a lot of times AIs just get the right answer and why would you get another take on it, it’s pointless. But there are enough problems where the answer is not well-determined or the AIs are just not able to do this well enough. If you have a group of AIs that vote, even if it’s the same AI, you’ll get a better answer because the stochasticity is just statistics, it helps you get a better answer. The other reason to use it is because different AI models are going to be biased in different ways and give you different answers. We typically don’t average the results from this ensemble. When we do this, we always see that the AI result is always better than the average.

Ross Katz: That makes sense. You have an ensemble of models, should I think of that as multiple CNNs trained with different hyperparameters or the same models, different model types? It sounds like you also bring some traditional computer vision type of approaches to bear, are you incorporating those into the ensemble and then on top of that, you’ve got potentially a decision tree model or something that sits on top of it that uses the information from the different models to make a decision? Am I thinking about that right?

Ilya Goldberg: With CNNs, typically the input is pixels, the input layer of neurons is really pixels, raw pixels or normalized pixels, and then you train the AI model from that. But you’re not limited to that because you can put tabular data into the input layers in addition to the pixels, because it doesn’t know what’s a pixel or what’s a cell in a table. You can combine results from conventional image analysis, which typically comes out as tabular data, together with your pixel data. We have done this more as conceptually easier with feature-based machine learning and classifiers because the input is a table. You’re going to say, some of this table comes from algorithms that process pixel data and turn it into a table of numbers, and other stuff is genetics or pharmacokinetics. Physiology, what’s your blood pressure — that all becomes tables of numbers and then it churns through the AI. Similar idea, we call this multimodal analysis because part of your data is image-based and part of it is not.

Ross Katz: That makes sense. And this is how a lot of the biological foundation models are trained right now too, with different multimodal data sets concatenated together and then brought together in order to get the signals that you’re looking for. Reese, can you give us an introduction to high content screens and how they’re used broadly and then what is AutoHCS and what do you bring to the table in that?

Reese Findley: Our automated high content screening tool is used to analyze high content screens, which are screens usually — not always but usually — collected on automated imagers, and they are high content because they generally have either brightfield images or several channels and you can pull a lot of information out of them. Specifically if you have AIs you can pull a lot of information out of them. Our automated high content screening tool has expanded quite a bit, but it started out to analyze these drug discovery assays where you apply a lot of different compounds at a lot of different doses to a series of cells and you see what they do. We train a series of AIs. This layers on top of AVIA, uses controls versus infected cells to train an AI to determine the difference between the two. We use controls with AutoHCS too, except in this case we are looking at target phenotypes. We have a negative control versus — let’s say you’re interested in apoptosis, you have a certain dose of staurosporine as your positive control. Now you have an identified target phenotype, you know what it does, when it does it, and you can compare other test drugs and see if they do the same thing. You can also train AIs on dose and get dose-response curves and those are extremely valuable because some compounds have a very binary response, they’re either on or off; most actually induce multiple phenotypes depending on what dose they’re at. Usually at the highest dose it’s some sort of extreme stress or death, so we expect to see several clusters with those types of compounds when we train on dose. Finally, we train what we call the all compound AI, which is what I was discussing earlier, an AI trained on every compound at a specific dose. We can do either the only dose that’s on the plate, we can do a highest dose. We also have a mechanism of training AIs to determine the lowest effective dose according to the AI. That’s the first time the AI thinks it looks different than the negative control, that’s the lowest dose that anything happens, according to the AI. I actually changed the name to lowest detectable dose since effective has some connotation to it. That’s the basic framework of AutoHCS and that’s what we’ve used to analyze screens that are manually collected with three compounds up to 1500 compounds. We have expanded this toolkit to also do experiments that are a little more complicated. For example, we’ve done background screening where you’re looking at neuroprotection and in this case there’s a compound applied that induces peripheral neuropathy, let’s say on neurons, and then there are protective compounds applied or at least test protective compounds and you want to see how much does this now look like the negative control where I didn’t apply anything. We’ve had pretty good success with identifying test compounds that are protective in that case. We’ve also looked at stem cell differentiation where stem cells are put into two different growth medias and then where do they differentiate, what are the stages of differentiation, and our AIs have been able to determine differentiation across a time course at the key stages we would expect. That leads me to one of the great benefits of AutoHCS: we have the capacity to analyze across time courses. Something we talk a lot about is analyzing in brightfield. The brightfield images actually contain quite a bit of information and people are realizing that now, starting to analyze in brightfield. The benefit is that you can now take a time course, you’re not killing your cells. There’s also live cell dyes, one of our big collaborators is Saguaro Biosciences, they developed live cell dyes and we do analysis of their test conditions over time courses. Because we have such an optimized system we’re able to train AIs very quickly. I’ve been able to train AIs at each time point across a massive screen and have it take a couple hours. That’s becoming the key benefit of AutoHCS: our capacity to rapidly and accurately train AIs and then analyze the data. There’s a lot to talk about in the toolkit and we can focus on any individual aspect but that’s a broad overview of everything we’ve done with it so far.

Ross Katz: I love it. And you’re right, there’s so many places to touch on. Where I’d like to go from here is understanding from the time that you receive imagery and you have a problem defined, how do you think through designing the way that you collect your experimental data or the way that you structure the modeling and the assay process to get the signal that you’re looking for?

Reese Findley: Primarily with AutoHCS, I tend to use feature classifiers. We have our own version of image feature extraction where we’re extracting about 2000 to 3000 features per image, actually per image tile, because we tile them down. With feature classifiers, we don’t need quite as much information or data as a CNN, especially since we are again using that AI confusion to give us information about how close or far apart in phenotype two compounds induce. We can do this analysis on smaller plates. The benefit is that we can increase it to a very large size, but we actually can do this analysis on — our absolute minimum is recommended 12 images per condition, although we have done less just to see what happens. We recommend about 12 images per condition and we design plates with the scientist usually. We can receive plates that already have data on them and we figure it out, but it’s nice to sit and design the experiment with the scientist because you want to consider: are you interested in dose response? Then you do need to save a fair amount of the plate for dose response so that we can get enough replicates of each dose to give you good information about it. Are you not interested in dose response? Okay, let’s leave that to the side, let’s maybe put two doses on each plate and focus on having replicates of each of your compounds. Do you have an unlimited budget? Do it all and we’ll pick what we want to analyze. That’s the initial conversation I’ll have with collaborators: what is your experimental question? What are you interested in answering? A big benefit of ViQi is my background is not in computer science formally. I’m a biologist. I can sit and have a very high-level conversation with our science teams and then go and do the technical work.

Ross Katz: I loved where you were going in terms of how the dose response question goes. Do your clients generally know in advance what are the phenotypes that they’re looking to detect or are you helping them to determine what are the different positive and negative controls that you want to introduce?

Reese Findley: It really depends on the client. Some people want to detect a certain type of cell death or cell stress and they know exactly what compound induces that and that’s going to be their positive control. That happens pretty frequently. They’re always going to know more about the compound space than we do because they’re in the lab, they’ve been studying this, they know exactly the phenotype they’re interested in. There’s also the larger and more general phenotypic profiling where basically you’re given a very large screen and they want to know which compounds look similar. I don’t actually have a target phenotype. I’m just interested in how all these compounds compare to each other. That’s where we get into that space where you need to make a decision of what to automate. Our automated analysis will cluster all of those compounds into interesting clusters. But then — she’s not here today but we have a bioinformaticist who will go and do a pharmacological validation and look at open source data sets, look at the pharmacological background of each of these compounds, what’s known, and start doing ground-truth validation of each of our clusters and making sure they make sense and to what degree do they not make sense. Those could be compounds that are inducing off-target effects which is something people are very interested in. If there’s a compound that really doesn’t fit in the cluster that it’s in, that’s great. That’s exciting. That’s the one we point to and say you’re interested in this. You should go do some other work to look into this compound. It really depends on the client or collaborator’s needs. Generally they know best what they need and I’m able to say, okay, if you are super interested in apoptosis, let’s identify three positive controls for it and make sure that we’re doing our due diligence there. Let’s also find a positive control for a different type of cell stress or cell death so that we delineate between the two. Those are the kind of questions I ask to make sure we’re doing our due diligence, but I would say generally, with a little bit of guidance on how the AI works, biologists know what they want and they have a good understanding of what they need on the plate.

Ilya Goldberg: Probably the most challenging kind of question is, is this compound doing anything weird? Most failures in drug development have to do with toxicity. The compounds are doing something inappropriate. That’s the off-target effects. You might have a positive phenotype kind of outcome that I want compounds to do, but you want to compare them to other compounds that do other kinds of things and see if at certain doses or under certain conditions or in certain cell lines your compound of interest is actually falling into a different cluster. In the worst case, a cluster of compounds that are known to cause problems. This phenotypic mapping is the new area that is offered by AI because you’re quantitatively comparing similarities of phenotype. That’s a very hand-wavy thing, but AIs can quantify that.

Ross Katz: It strikes me that because of the image-based data sets that you’re using there’s so much information content inside of the images that can answer both the questions that you’re trying to answer with the assay but also leaves information on the table for the other follow-up questions that you might want to ask such as Ilya what you’re talking about with toxicity and with the ability to cluster things together. You can then take a different angle on the question and compare the images to different things that you weren’t expecting to have to compare them to but because you’re starting from that place of working with microscopy you have the information content available to do that. Am I thinking about that right?

Reese Findley: Along those lines, something I’ve thought a lot about in this space is the phenotypic space that the AI is exposed to. If you have two compounds that do really similar things and you train a binary AI to distinguish these two compounds, if it’s successful (which it isn’t always, it might be confused), now in that phenotypic space those two are going to look really far apart. But now you add a third compound that does something very different and all of a sudden they get very close together. You’re giving the AI a relative amount of phenotypic information. I’m literally running it right now — I have a series of time points and a series of compounds and I started out by combining all the time points and compounds and running a really big AI on it. Then I look at what I think is an interesting cluster and now I train an AI on just that sub-cluster so that I can get more space between those compounds and have a better understanding of how they separate in that smaller phenotypic space. Exactly like you’re saying, there’s so much information in the image content and depending on how much information you give an AI about that, it’s going to make determinations based on that. Having a good understanding of how AIs train and what they’re exposed to is going to determine how they structure the clusters you get back.

Ross Katz: Are there any case studies that you can share from your client work? Projects that you felt really proud of the outcome that you delivered?

Reese Findley: Personally, I really like our time courses. I think they’re novel, I think they’re interesting, I think that it demonstrates a drug’s efficacy over time in a way that these single time points really don’t. Maybe it was naive of me coming from a neuroscience behavior background, but that’s something that I found very striking and really was a very fast result for me to look at and say, oh wow, we should be doing this the whole time. We should be doing time courses with every experiment because the kinetics provide a lot of information. Being able to develop that and having it run so quickly, so consistently, I’m proud of it.

Ilya Goldberg: That specific example of time courses has traditionally been prohibitive because if you’re doing your imaging on fixed cells, which you need to do if you’re doing fluorescent probes and stuff like that, the cells are dead. There is no second time point. If you want to get another time point, you have to start the whole experiment again. If you’re doing this in brightfield on live cells that haven’t been stained, you get the other time points essentially for free because it’s literally a robot that’s doing the imaging at the other time point. It’s a huge area that opens up studying things in a time course. The only reason it wasn’t done before — people know this obviously, the only reason that it hasn’t been done is because it’s prohibitive, expensive. Being able to process things in brightfield in these broad non-toxic dyes is what opens this whole thing up. I think it’s a really good point because it opens up a whole territory that’s been unavailable by other means.

Ross Katz: This has been really fascinating and as we head towards the end, I have a couple wrap-up questions. What excites you most about the future of microscopic imaging and AI in life sciences?

Ilya Goldberg: For me it’s the fact that it’s becoming more common. Being at the technology development forefront, a lot of times you wait a long time for the technology to actually become widely used. To me, the future is the breadth of applications and there’s more brains and eyes involved. Everybody’s now thinking how can I use AI. The breadth of applications now is really exciting. It’s really taking off on this logarithmic expansion, and there was a couple of decades where people were ticking things over and you’re like oh my god this is going to be really great, and a lot of times in technology development you’re like okay when is it going to be really great. That time is now. To me the most exciting part is how broadly used it’s going to be very soon. And now even.

Ross Katz: Reese?

Reese Findley: For me I bring it full circle, I am excited about alleviating frustrations for bench scientists. That is such a cool way to contribute to science because so much of their energy and attention goes towards these frustrating tasks where it could be going towards analysis and innovation. I love that that’s what our tools do, I look forward to continuing to do that with different assays and it’s really exciting for me when I can come to somebody and say, oh yeah, that thing that took four weeks for you and made you want to pull your hair out, let me do that for you and I can do it in a fraction of the amount of time.

Ilya Goldberg: That applies to medicine as well where you have imaging like radiologists and pathologists who would be like aren’t I out of a job because my job is basically staring at images and saying what’s going on there. The much truer and better answer is that radiologists and pathologists now get to do far more interesting work because most of what they do is the trained monkey work and AIs are great at this. They’re not great at the corner cases and edge cases of medicine; these are the interesting cases. This is where a lot of knowledge of medicine can be brought in and become relevant as opposed to screening through thousands and thousands of samples where the answer is very clear-cut. If you drop that out of their workflow, it actually makes their jobs a lot more interesting than they are now because they are doing stuff that AI should be doing.

Ross Katz: That makes a lot of sense. As we wrap up, for listeners who are intrigued by ViQi or microscopic imaging-based AI in general, where would you suggest they go to learn more?

Ilya Goldberg: Our website is a good place to start. We’re at viqiai.com. That’s VIQIAI.

Ross Katz: That’s V-I-Q-I-A-I.com, okay.

Reese Findley: Our website has a lot of our previous publications and if you go to our LinkedIn you’ll see announcements of the conferences that we’ll be at and there’s contact information on both.

Ross Katz: Ilya and Reese, it’s been a pleasure talking with you today. Thanks so much for joining and look forward to connecting down the line.

Ilya Goldberg: Great. Thank you.

Reese Findley: Thank you. This was fun.

Ilya Goldberg: Nice to meet you.

Jason: And that’s it for this episode of Data in Biotech. If you enjoyed the episode, please subscribe, rate, or leave a review in your podcast platform of choice. See you next time.

Frequently Asked
Questions

Why train a separate AI for each lab and each virus instead of a universal model?
Different viruses produce different phenotypes in different cell lines, and even two microscopes of the same model produce images an AI can tell apart. ViQi trains per-lab, per-assay AIs in hours by layering training on top of in-plate controls, which avoids the bias that comes from forcing one model across every variation.
What does brightfield microscopy with AI unlock that fluorescence-based assays cannot?
Brightfield keeps cells alive across a time course instead of killing them at a single endpoint, and it pulls signal from the full image rather than a handful of stained proteins. Combined with live cell dyes from collaborators like Saguaro Biosciences, teams can run time-course phenotypic profiling that was previously prohibitive on cost and cell survival.
Where does AI-based phenotypic profiling still need a biologist in the loop?
Clustering output from a screen of hundreds to 1500 compounds is stochastic, and different analysis choices produce different clusters. ViQi pairs automated AutoHCS results with bioinformatics validation against open-source pharmacological data — the expert eye confirms which clusters carry biological meaning and flags off-target effects worth chasing.

Need a data partner for life sciences?

CorrDyn helps biotech and pharma companies build the data infrastructure that accelerates research and operations.

Book an intro call