The Scientist in the Loop: Why Biopharma R&D Still Needs Human Judgment

Jul 29, 2026 | Health Tech

Image Source: Authors Own
Independent Contributor
Written by: Matt Hasan and Farah Hasan
On behalf of: aiRESULTS

Two years into the biopharma industry’s most sustained experiment with artificial intelligence, the honest answer to “what has changed” is not the one most conference keynotes suggest. It is not that biopharma AI now runs the laboratory.

It is that AI has become very good at a narrow set of tasks with clean, verifiable inputs, and still depends on a human expert to decide which of its suggestions are worth pursuing. That distinction, between automation of information and automation of judgment, is where the real story of this period sits.

The 2026 Biotech AI Report from Benchling, drawn from one hundred biotech and biopharma organizations actively using AI, offers the clearest picture yet of where this line falls. A handful of use cases have moved past the pilot stage into daily practice: literature review, now adopted by seventy six percent of surveyed organizations, protein structure prediction at seventy one percent, scientific reporting at sixty six percent, and target identification at fifty eight percent (Benchling, 2026).

What these four have in common is not glamour but tractability. They succeed because they run on data that is clean and easy to verify against an external standard, whether that standard is a published paper, a crystal structure, or a regulatory template. Generative design and biomarker analysis, by contrast, have not broken out of pilot mode, and the report attributes this directly to data that is scattered, incomplete, and hard to validate (Benchling, 2026).

This is a useful corrective to the more sweeping claims that circulate about AI in drug discovery. It is not that the harder problems are less amenable to modeling in principle. It is that the data infrastructure underneath them has not caught up, and no model, however capable, can validate a hypothesis against data that does not exist in usable form.

Where the human remains the load bearing element is in the judgment that follows a model’s output. Reporting on the same Benchling data, Drug Discovery News noted that sixty six percent of scientists report increased confidence in large language model outputs over the past year, but that this trust is maturing into what the industry has started calling a “trust but verify” posture, in which scientists still rely on domain expertise to decide when an AI generated hypothesis is worth the cost of testing it (Drug Discovery News, 2026).

This is worth dwelling on, because it names precisely the kind of expertise that is hardest to formalize into a benchmark. A model can rank candidate targets by predicted binding affinity or novelty score. It cannot yet weigh that ranking against the tacit sense, built over years at the bench, of which results tend to replicate and which tend to evaporate under a slightly different assay condition. That sense is not documented anywhere a model could learn it from. It lives in people.

The same report offers a second, quieter signal about where trust actually sits inside organizations: the split between buying biopharma AI capability and building it in house. Roughly sixty percent of biotech teams buy proven commercial components, while fifty five percent build or fine tune models internally where their proprietary biology is genuinely unique (Drug Discovery News, 2026).

Read against the adoption figures above, this split looks less like a procurement preference and more like a risk map. Commodity tasks, the ones with external verification standards, get outsourced to vendors. Tasks that touch a company’s genuinely differentiating science stay in house, under closer human supervision, because the cost of an undetected error there is much higher.

The clearest illustration of where the loop currently closes, and where it does not, comes from the multiagent discovery systems now generating real results. FutureHouse’s Robin, described in a Nature paper published in May 2026, integrated literature review, hypothesis generation, and data analysis to propose repurposing the glaucoma drug ripasudil for dry age related macular degeneration, through a mechanism involving enhanced phagocytosis by the retinal pigment epithelium, and separately flagged the circadian clock modulator KL001 as a promising candidate (Ghareeb et al., 2026).

Both findings were then evaluated in vitro using primary human retinal pigment epithelium cells. That final clause is the part worth underlining. Even the most autonomous system currently in the literature still terminates its reasoning cycle at the door of a human run wet lab. The agent proposes. The bench disposes.

Whether that boundary holds is an open question the industry is actively trying to erase. Pharma industry commentary on “lab in the loop” architecture describes a near future in which AI systems generate experiment plans that robotic systems execute directly, with results analyzed and fed back into the model’s reasoning cycle without a human in the intermediate step (Pharmaphorum, 2026).

If that architecture matures, the wet lab checkpoint that currently catches a flawed hypothesis before it consumes real reagents and real time may itself become automated, which raises the stakes on getting the verification step right before removing the person who currently performs it.

This is also where the governance conversation earns its keep, and where this piece connects to the regulatory questions we addressed in an earlier installment on AI governance in biopharma. The same biopharma AI capabilities that accelerate literature search, protein design, and candidate identification carry recognized dual use risk, a concern serious enough that computer scientists, biologists, and policymakers have called for coordinated safety benchmarks specific to biological applications (ABC Bench, 2026).

Separately, on the more mundane but no less important question of operational error, some newer laboratory information systems now give AI agents the same permissions as human users while logging every action in an audit trail built to satisfy 21 CFR Part 11 requirements (Genemod, 2026). That audit trail is not a glamorous feature. It is, at present, one of the few concrete mechanisms by which a mistaken automated action in a regulated lab can be traced back and corrected.

The stakes of getting this right compound in a way specific to this industry. Drug development typically runs ten to twelve years from target to approval, and reporting on this AI transition notes that because of this timeline, improvements or errors introduced at the earliest discovery stage compound over the full length of the program (Drug Discovery News, 2026).

An error a model introduces in year one is not a year one problem. It is a decade long problem that surfaces, if it surfaces at all, only after enormous downstream investment has already been made on the strength of an unexamined assumption.

None of this argues against the technology. The literature review and target identification numbers alone represent a genuine and durable gain in how quickly a scientist can get from question to evidence.

But the industry’s own data is telling a more disciplined story about biopharma AI than its marketing usually does. The scientist is still in the loop, not as a ceremonial safeguard, but because the judgment required to decide what is worth testing has not yet been reduced to a benchmark. Whether that remains true in two years is probably the single most consequential open question in biopharma R&D right now.

 

Author Bio

Matt Hasan, PhD, is CEO of aiRESULTS and founder of the AI Humanist Movement. He advises Fortune 50 organizations on AI strategy and governance and previously held leadership roles at Deloitte, IBM Global Business Services, and Capgemini.

Farah Hasan, PhD, is a scientist in immunotherapy development with experience at BioNTech, NYU Langone Health, and the Icahn School of Medicine at Mount Sinai.

    References: ABC Bench: An Agentic Bio Capabilities Benchmark for Biosecurity. Proceedings of the 43rd International Conference on Machine Learning, PMLR 306, 2026. Benchling. 2026 Biotech AI Report. Benchling, 2026. Drug Discovery News. "The 2026 AI Power Shift." February 24, 2026. Genemod. "Top 10 AI LIMS Platforms in 2026: A Buyer's Guide for Biotech and Pharma Labs." 2026. Ghareeb, Ali Essam, et al. "A Multi Agent System for Automating Scientific Discovery." Nature, May 19, 2026. McKinsey & Company. "From Linear Gates to Learning Loops: Rewiring Biopharma R&D with AI." 2026. Pharmaphorum. "AI Scientists and the Robotic Labs of Tomorrow." February 24, 2026.
    All content is published for informational purposes only and does not constitute medical, legal, or investment advice. For more information, see our Terms and Conditions

    Articles that may be of interest

    Safety, Toxicology, and Risk Assessment of Essential Oils

    Safety, Toxicology, and Risk Assessment of Essential Oils

    Essential oils are complex mixtures of volatile phytochemicals that have long been incorporated into traditional medicine, personal care products, aromatherapy, and topical therapeutic formulations. Their diverse biological activities—including antimicrobial,...

    read more
    The False Economy of Returning to Manual PGD Processes

    The False Economy of Returning to Manual PGD Processes

    Across healthcare, financial pressure is forcing difficult decisions. Pharmacy teams are being asked to do more with less. Service expansion, dealing with NHS patients, workforce shortages remain acute, and every investment is being scrutinised. In this environment,...

    read more

    Articles that may be of interest

    Safety, Toxicology, and Risk Assessment of Essential Oils

    Safety, Toxicology, and Risk Assessment of Essential Oils

    Essential oils are complex mixtures of volatile phytochemicals that have long been incorporated into traditional medicine, personal care products, aromatherapy, and topical therapeutic formulations. Their diverse biological activities—including antimicrobial,...

    read more
    The False Economy of Returning to Manual PGD Processes

    The False Economy of Returning to Manual PGD Processes

    Across healthcare, financial pressure is forcing difficult decisions. Pharmacy teams are being asked to do more with less. Service expansion, dealing with NHS patients, workforce shortages remain acute, and every investment is being scrutinised. In this environment,...

    read more