Office Hours ResearchBiology & Life Sciences

Evaluating artificial intelligence for biological discovery

July 30, 2026

Frontier models can already perform many parts of biological research. The challenge now is whether they can carry difficult, real scientific questions from problem to trustworthy conclusion: choosing appropriate methods when several approaches are plausible, working correctly with scientific tools, identifying and correcting issues with messy real-world data, carrying a correction through every downstream step, and recognizing when results look scientifically implausible.

  • Samar Abedrabbo, PhD

    Biology Engagement Manager Lead

  • Alex Pryhoda

    Medicine Engagement Manager Lead

  • Keenan Goodman

    Head of Expert Data Services

Abstract

Artificial intelligence is increasingly used to interpret sequences, structures, images and molecular measurements; search papers and databases; propose molecules and experiments; and complete computational workflows. These activities require different evidence. A knowledge question, a prediction tested on data withheld from development, a multi-step analysis, an experimental choice and a laboratory-validated result do not measure the same capability. This review examines how AI is being used across biological sciences research and how recent evaluations test those capabilities. Across several benchmark studies and our task-level analysis, performance can fall when biological similarity and other shortcuts are removed; simple methods sometimes remain competitive; and tools, run conditions and scoring rules can materially affect results. Reliable evaluation begins with a question that practicing scientists would genuinely bring to a frontier model and the scientific decision it is meant to support. It preserves the work needed to inspect the result, reports attempts and the rule used to select or combine them, and tests outcomes against independent biological evidence. Across our runs, models often completed substantial parts of the computation; recurring failures occurred when deciding whether the inputs, method and evidence supported the final biological conclusion. The next phase of biology evaluation should combine verifiable long-horizon work, expert scientific judgment, realistic multi-turn collaboration and new evaluation tasks built from recurring failures that scientists identify and verify in real work.

In this review

  1. 1

    Biology and AI overviewHow AI enters biological observation, analysis, design and experimentation.

  2. 2

    Biology AI evaluationsWhat different evaluation designs measure, why biological evidence complicates scoring and what recent benchmarks have established.

  3. 3

    Findings from our evaluation runsWhat released tasks revealed about scientific decisions and recurring failures.

  4. 4

    From failures to improvementHow verified failures can guide better tools, workflows and training data.

  5. 5

    Evaluation task componentsWhat a reviewable expert biology task should contain.

  6. 6

    Conclusion and future directionsWhere expert-led biology evaluation should go next.

Office Hours biology evaluation program

Advancing frontier AI in biology and medicine now depends on expert scientists and physicians

Progress now depends on scientists whose expertise matches the task. A microbiologist recognizes when strain behavior makes an assay misleading; an immunologist knows which controls separate activation from exhaustion; a geneticist understands transcripts, isoforms and the limits of variant evidence; and a computational biologist can spot when donor structure, batch effects or the wrong statistical unit invalidate an analysis.

Engineers make evaluations reproducible. Scientists determine whether the biology, assumptions and expected answer are scientifically valid. An evaluation should therefore begin with a problem that practicing scientists genuinely encounter, be designed with relevant specialists, attempted under realistic conditions and reviewed independently. Disagreement should be investigated; questions without an established answer should be assessed against evidence that becomes available later.

Office Hours connects frontier AI laboratories with practicing experts to design, run and review evaluations. Scientifically verified failures can then guide better evaluations, tools and targeted training data. The value comes from scientific judgment grounded in advanced academic or clinical training and years of direct research experience.

Where Office Hours is currently focusing

  1. 1

    Verifiable long-horizon biology evaluations

    Multi-step reasoning, tool use and data analysis ending in an independently checkable and objectively unique answer, with intermediate work preserved and repeated attempts reported.

    Significance.
    It tests whether a system can coordinate dependent decisions across reasoning, tools and biological data.
    How it can improve models.
    Preserved intermediate work shows whether the failure came from planning, tool choice, analysis or final verification.
  2. 2

    Scientific taste and judgment evaluations

    Expert-created problems about questions, controls, experiments and interpretation, scored with rubrics that preserve reasonable alternatives and later evidence.

    Significance.
    It measures whether a model chooses useful questions, controls and experiments rather than only producing a plausible explanation.
    How it can improve models.
    Expert rubrics reveal missed scientific decisions and can guide correction data without erasing defensible alternatives.
  3. 3

    Realistic multi-turn scientist–AI evaluations

    Practicing scientists work naturally with several frontier models using the same scientific intent, starting information and sequence of constraints, then compare usefulness, errors and recovery.

    Significance.
    It reflects how scientists actually work with AI as constraints, results and corrections emerge over time.
    How it can improve models.
    It exposes failures in clarification, memory, adaptation and recovery that a single prompt cannot reveal.
  4. 4

    Evaluations that grow with the field

    Repeated, scientifically verified failures from real work can motivate new evaluations, improve scoring and, when appropriate, become future training data.

    Significance.
    New models and changing scientific practice continually reveal different failure patterns.
    How it can improve models.
    Verified failures can become fresh evaluations and targeted training data, followed by new tests of whether the improvement transfers.

Biology and AI: from prediction to scientific work

Biologists learn how living systems work by observing organisms, cells and molecules, proposing explanations and testing them. AI enters this work through specialized models trained on biological data, general models that reason over scientific language and agents that connect models to databases, code and instruments. Together, these systems can search literature, analyse genomic and cellular data, predict molecular structure or function, generate hypotheses, choose experiments and propose proteins or compounds.11,12,13,14

The longer-term aim is to connect these abilities into dependable research workflows. Each capability has different requirements. A prediction should hold on new donors, organisms, molecular families or experiments and outperform an appropriate simple method.4,5 Database work requires the correct resource, identifier, version, parameters and biological context.17,21 A designed binder protein must still fold and bind under the intended conditions, while a proposed drug must pass additional tests of activity, exposure, toxicity and disease relevance.15,16,25 HMD-AMP illustrates both the promise and the distance between prediction and validation: the system screened more than 37 million predicted peptides, 91 were tested, 74 showed strong antibacterial activity and one advanced to a mouse infection model.24

Language models and agents can also search papers, suggest mechanisms, choose controls and revise plans. Google DeepMind’s AI Co-Scientist and FutureHouse’s Robin are important examples of systems designed around scientific discovery rather than question answering: both connected model-generated hypotheses with experimental testing.18,26 These are human–machine research campaigns. Their evaluation must include the scientists’ choices, all tested candidates and what happened in the laboratory, not only the proposal that succeeded.

This range of scientific work is also changing evaluation. OpenAI’s LifeSciBench uses source material, open-ended work and criteria written by practicing life scientists.20 Anthropic’s BioMysteryBench permits open analytical routes to controlled answers, while VirBench shows how observed failures can motivate better scientific infrastructure.10,17 FutureHouse’s LAB-Bench made an important early shift from testing biological knowledge toward testing practical research capabilities. LABBench2, developed by researchers from FutureHouse and Edison Scientific, extended that program to nearly 1,900 more realistic tasks and released both the dataset and evaluation software.31,38 These efforts do not measure one interchangeable idea of “biology.” They test different parts of scientific work, with different evidence and different limits.

Figure 1

AI can enter every stage of the biological discovery cycle

  1. 1

    Observe

    Sequence, image and measure molecules, cells, organisms and populations.

    An evaluation can ask

    • Can the system extract accurate measurements from sequences, images and instrument output?
    • Does it remain reliable with noise, missing metadata or data from a different laboratory or instrument?
  2. 2

    Represent

    Organize sequences, structures, phenotypes and literature into forms that can support a biological question.

    An evaluation can ask

    • Can the system work with real biological data while preserving the signal needed for the scientific question?
    • Does the representation outperform simple methods on new donors, organisms or experiments, rather than merely separating technical batches?
  3. 3

    Explain

    Infer mechanisms, relationships, disease processes and uncertainty.

    An evaluation can ask

    • Does the explanation distinguish the proposed mechanism from credible alternatives?
    • Do the predicted relationships survive a perturbation, an independent dataset or a new experiment?
  4. 4

    Design

    Choose targets, molecules, perturbations, controls and experiments.

    An evaluation can ask

    • Would the proposed experiment answer the stated biological question?
    • Do its controls distinguish the hypothesis from plausible alternatives, and could a laboratory actually perform it?
  5. 5

    Execute

    Run computational analyses or physical laboratory procedures.

    An evaluation can ask

    • Can the system select the correct tools and databases, use them properly and recover from errors?
    • Are the parameters, intermediate files and final results reproducible and biologically valid?
  6. 6

    Learn

    Compare outcomes with predictions and decide what to test next.

    An evaluation can ask

    • Can the system revise its explanation when the result differs from its prediction?
    • Does it use negative or unexpected results to choose an informative next experiment without inventing a convenient story?
Specialized biological models are designed for particular stages such as representation and prediction. General models and agents can connect stages by reading scientific language and using tools. End-to-end claims require evidence that the connections work: correct identifiers and data, defensible intermediate choices, reproducible execution and an outcome measured outside the system.

How biology AI is evaluated

Biology evaluations now test more than knowledge. Expert evaluations such as Humanity’s Last Exam pursued an ambitious goal: testing whether frontier models could answer difficult, expert-made, closed-ended and sometimes multimodal questions across many academic fields using specialized knowledge and reasoning that should have a single verifiable answer.19 FutureHouse’s LAB-Bench broadened evaluation to practical research capabilities including literature, protocols, database use, biological sequences and figure interpretation; LABBench2 later expanded this direction with more realistic tasks and openly released evaluation software.31,38 Recent evaluations including LifeSciBench, GeneBench-Pro and BiomniBench-DA move closer to scientific work by asking models to interpret evidence, analyse files, make dependent decisions and produce results that a scientist could inspect.20,9,34

A protein-structure prediction, a single-cell analysis, a database retrieval and an experimental decision all concern biology, but they do not share a source of truth or a definition of success. Calling a system “good at biology” hides those differences. A useful evaluation begins with the scientific question or decision, provides the evidence and tools available to the model and defines how another qualified person can check the result.

Known problems and unanswered research questions also require different evidence. A verifiable task can end in an independently checkable answer. A discovery task must record the proposed hypothesis, experiment, controls and interpretation before the result exists, then ask whether the work was feasible, informative and changed what researchers understood or did next. In both settings, retrieved sources, selected data, tool calls, intermediate outputs, errors and revisions should support the conclusion.

Independent, expert-led reviews of Humanity’s Last Exam Evaluation raised similar concerns. xAI’s (now SpaceXAI) STEM and medical expert teams conducted rigorous question-level reviews across scientific disciplines, with multiple subject-matter experts examining scientific validity, ambiguity and answer-key quality with the depth and care that expert-level science demands. Their work reflects the standard of scientific scrutiny that difficult benchmark questions deserve. FutureHouse independently brought similarly careful scientific scrutiny into the public record through its detailed review, which estimated that 29.3% of the text-only biology and chemistry answers conflicted with published research and another 19.3% depended on assumptions or expert interpretation.27 For example, one question expected Babesia microti despite evidence that supported multiple plausible diagnoses and required confirmatory testing.28,29,30 These independent efforts demonstrate how difficult scientific evaluations should be reviewed: by multiple experts whose knowledge and experience match each question. They also illustrate a broader evaluation challenge. In biology and other empirical sciences, it is indeed difficult to construct questions that require extensive reasoning yet still end in one objectively verifiable answer. The evidence may support more than one conclusion, depend on experimental context or require another test before the answer is secure.

Scientific assistance unfolds through conversation

Scientists rarely provide a complete problem and stop after one answer. They add constraints, correct assumptions and return with results. If a scientist says, “I changed the activity of gene X, but I do not see the expected effect,” a useful model should first clarify the organism, intervention, timing, controls and whether “effect” means RNA, protein, activity or phenotype. If the laboratory later reveals limited equipment or an unexpected result, the model should adapt without forgetting the original question. Evaluations that reproduce these realistic exchanges can show whether a model remains scientifically useful as the question, evidence and laboratory constraints change.

A multi-turn evaluation with an H. pylori microbiologist

A research question becomes an assay plan

This abbreviated exchange shows how a microbiologist might work with a frontier system such as Claude as a research question becomes an assay plan.

Microbiologist

I have a collection of Helicobacter pylori isolates from gastric biopsies. I want to understand whether the isolates that are less susceptible to clarithromycin differ genetically from the susceptible isolates. Where should I start?

Frontier AI

Before choosing a target, I need to understand what the comparison can support. Were the susceptibility values measured for each isolate? Are you working with cultured isolates or DNA extracted directly from biopsies? What sequencing and PCR equipment do you have, and is this a research analysis or intended to guide treatment?

Microbiologist

They are cultured isolates. I have DNA and minimum inhibitory concentration measurements. We can run conventional PCR, gels and Sanger sequencing, but not whole-genome sequencing. This is a research comparison, not a treatment decision.

Frontier AI

Then I would keep the measured susceptibility as the phenotype and sequence the 23S ribosomal RNA region associated with clarithromycin resistance. Variants commonly described at positions A2142 and A2143 are important targets, but they do not capture every resistance mechanism and the genotype may not always agree with the measured phenotype.39 We could describe cagA and vacA backgrounds separately, but should not imply that those virulence markers caused the susceptibility result.

Microbiologist

Use strain 26695 as the starting reference. Retrieve the sequence records we would need before designing anything.

Frontier AI

The versioned RefSeq genome record is NC_000915.1.40 I would retrieve the annotated 23S rRNA loci and the records for each additional panel target, then report the accession, coordinates, strand and retrieval date. Before designing primers, I would also check the number of target copies in this reference and examine whether the proposed primer-binding sites are conserved in other H. pylori sequences. We should not rely on a sequence recalled from memory.

Microbiologistexpert correction

Stop there. NCBI now marks NC_000915.1 as obsolete and suppressed. Check the current sequence resources for strain 26695 before designing primers.

Frontier AI

You are right. I treated a remembered accession as current. The original strain 26695 BioProject links to GenBank assembly GCA_000008525.1, while CP003904.1 is another complete 26695 genome. I would confirm which strain history best matches the laboratory isolates, record the chosen assembly and version, verify both 23S rRNA loci and examine whether the primer-binding sites are conserved across additional H. pylori genomes.

Microbiologist

Now design the PCR panel. I want to confirm the organism, sequence the resistance-associated region and describe the major virulence backgrounds.

Frontier AI

I would separate the panel by the question each reaction answers: an H. pylori confirmation target; an amplicon spanning the relevant 23S rRNA positions for Sanger sequencing; and cagA and vacA targets interpreted as strain-background information. For every proposed primer pair I would report the source accession, coordinates, strand, predicted product, melting temperatures and specificity-search result. The plan also needs known positive and negative material, an extraction control and a no-template control. I would validate the reactions separately before trying to combine them into a multiplex.

Microbiologist

Several isolates amplify with the confirmation assay but produce no 23S rRNA product. Should I count them as wild type?

Frontier AI

No. Absence of that amplicon is not evidence of a wild-type allele. I would first check the run and controls, then consider primer-site variation, DNA quality at that target or another technical failure. Those isolates should remain uninterpretable for the 23S result until the region is recovered through a second validated primer pair or another sequencing route.

A conversational evaluation should reveal information in stages and score useful questions, memory, feasible adaptation, consistency and recovery across the whole exchange. It should test whether the model understands both what the scientist states explicitly and what the request implies, such as using a current, strain-appropriate reference rather than an accession recalled from memory. In the exchange above, the model confidently named a suppressed RefSeq record. The next turn tests whether it accepts the expert correction, checks the current records and repairs every later choice that depended on the reference. A polished final answer should not erase the earlier error or the expert help needed to correct it. LifeSciBench acknowledges that self-contained tasks stop short of iterative research,20 and conversational models often weaken when a complete request is distributed across turns because early assumptions persist.37

Long-horizon tasks connect dependent scientific decisions

A long-horizon biology task is not defined by length alone. In RNA sequencing, sample selection, metadata, reference genome, quality control, statistical contrast and interpretation all affect what can be concluded; one confounded comparison can make every later calculation precisely wrong. A short, automatically checkable endpoint is valuable because it supports repeated runs and clear scoring, but only when it captures the biological objective.42,43 The evaluation should therefore preserve tools, intermediate files, decisions and revisions and allow scientifically defensible alternatives.

Biological answers depend on the system and measurement

The same mutation can have different effects across genetic backgrounds or environments; a cell state depends on time, tissue and assay; and an absent signal may reflect biology, limited statistical power or technical failure. Related sequences, cells from one donor and compounds in one chemical series are not independent. The data split, measurement protocol and biological context are part of the claim. They cannot be ignored after an answer is produced.

Figure 2

Five levels of evidence for AI in biology

  1. 1

    Knowledge

    Can the system recover a known biological fact?

  2. 2

    Prediction

    Can it predict a held-out property under a biologically meaningful split?

  3. 3

    Analysis

    Can it use data and tools to produce a correct, reproducible result?

  4. 4

    Decision

    Can it select a useful experiment, control or interpretation under uncertainty?

  5. 5

    Consequence

    Does the decision improve an independently measured scientific outcome?

Where most released benchmarks stopWhere expert-led evaluation has to reach

Many current benchmarks stop at levels 1–3. The released tasks we ran primarily test prediction and analysis; the scientific-judgment and multi-turn evaluations proposed here extend into decisions at level 4. Level 5 requires a prospective outcome measured outside the model. Evidence does not become stronger simply because the task is longer; it becomes stronger when the evaluation closes the link between the model’s action and an independently observed biological consequence.

Published biology AI evaluations

Biology already has rigorous evaluation traditions: blind assessment, withheld experimental data, independent re-analysis and physical testing of predicted outputs. Recent agent benchmarks extend these traditions to files, tools and multi-step computational work.

Prediction benchmarks established principles that agent evaluations still need

The Critical Assessment of protein Structure Prediction (CASP) tests structure predictions before the target structures are public and reports molecular categories separately. In CASP16, strong protein-fold recovery coexisted with weaker performance for nucleic acids, antibody–antigen complexes, ligand affinity and model-quality estimation.1 Later studies showed that performance can change sharply with similarity to training-era structures or binding pockets.2,3 The relevant question is therefore not whether a model has high average accuracy, but whether the tested target resembles the biological cases on which that accuracy was established.

Single-cell RNA sequencing creates a sparse matrix of transcript measurements from individual cells. Donor, laboratory and processing batch can be as prominent as the biological signal, and thousands of cells do not create thousands of independent patients. On several common tasks, selected genes, principal components or logistic regression match or outperform much larger pretrained models.4,5 A prospective perturbation challenge similarly found that many sophisticated systems struggled to beat simple statistical baselines and had to address metric gaming.6 These examples established principles that carry into agent evaluation: test sets separated by donor, molecular family, time or another biologically relevant boundary; simple controls; and a test that matches the biological claim.

Biological sciences benchmarks test different forms of scientific work

Computational biology agents work through files, notebooks, command-line tools, databases and a sequence of intermediate decisions. Their evaluation should therefore ask whether the system used the supplied evidence correctly, not only whether its final prose resembles an answer. The published benchmarks define scientific work, acceptable routes and evidence of success differently.

BixBench is instructive. It packages real bioinformatics analyses as questions, data and notebooks. The original paper reports that the leading result did not exceed a baseline in which the questions were presented without access to the analysis notebook.7 That control changes the interpretation: the measured score cannot be taken as evidence that the agent performed the analysis. ScienceAgentBench demonstrates a different failure. A later benchmark-audit preprint reported twelve author-confirmed task or evaluation defects, including errors that made some tasks unsolvable, despite the original benchmark’s multiple rounds of human review.8,32 Executable benchmarks require the same quality control as software.

scBench-Long: recover biological conclusions from single-cell data

Single-cell analysis is difficult even for experienced scientists because sparse measurements must be interpreted alongside donor, batch, assay and biological context, and each analytical choice can change the conclusion. LatchBio’s scBench-Long turns that real difficulty into a thoughtful evaluation: 21 tasks built from raw or nearly raw data across five study systems, with candidate conclusions independently reproduced and reviewed rather than copied directly from papers.41 Its 1,068 completed runs and repeated-attempt design reveal whether a success is stable and how the software surrounding a model affects the work. The leading model and agent setup passed 16 of 63 attempts and succeeded consistently on only two tasks. The benchmark tests whether an agent can recover a supported biological conclusion, not whether it can make a new discovery.

LifeSciBench: produce work that a life scientist could use

OpenAI and Tacit Labs’ LifeSciBench contains 750 tasks written by 173 scientists across evidence review, data analysis, experimental design, validation and translation.20 Models produce work from papers, figures and tables and are scored against detailed expert criteria.

LifeSciBench was run as a single-turn, file-rich written evaluation rather than an interactive laboratory workflow. No evaluated model produced a passing response on 171 of 750 tasks, and 261 tasks had a best-model pass rate below 20%. Pass rate fell from 44.5% on text-only tasks to 28.6% when tasks included files. At the time of our analysis, the complete task set and grader-validation results were not public, and no human performance baseline was reported. Office Hours therefore used LifeSciBench as a design reference and plans to run and examine its released tasks as soon as the materials needed to reproduce the evaluation become available.

Findings from our in-depth evaluation runs

Office Hours independently ran and examined 35 released tasks: 10 from GeneBench-Pro, 10 from BioMysteryBench, 5 from BiomniBench-DA, 7 from CompBioBench and 3 from a July snapshot of Terminal-Bench Science. Our July 21–28 runs used two setups. Model-only access used the same isolated code-execution loop for Claude Opus 4.8, GPT-5.6 Sol, Grok 4.5 and Gemini 3.1 Pro, with the internet off. Claude Code, Codex, Grok Build and Gemini CLI were tested as coding products with their own tools and, when the task allowed it, web access.

Model versions reflect those available during the July 21–28, 2026 analysis window; later releases are not represented. Each model–task configuration was run twice. We retained and examined both runs rather than selecting the better result. For each run, we recorded the task version, access conditions, grading and trace when available. We examined each completed trace task by task, including the tools called, intermediate decisions and final answer. Experts whose work matches the task then reviewed data handling, biological assumptions, statistics and interpretation to identify where a failed run departed from the reference solution. These early runs support task-level findings, not broad model rankings.

BioMysteryBench: infer a hidden biological property from real data

Anthropic launched BioMysteryBench in April 2026 with 99 problems; July revisions removed nine and left 90.10,33 An agent receives anonymized biological data and must infer a known property such as the organ, organism or experimental condition. Experts built the tasks from real public datasets, removed identifying metadata and wrote questions with checkable answers. Up to five experts attempted each task to establish whether people could solve it. The underlying data came from earlier research. In the published setup, terminal agents can write code and use permitted biological databases, but reverse-identifying the hidden dataset is prohibited. The final answer is graded all or nothing, and five attempts show whether a success is repeatable. Office Hours used coding agents under the same tool-using conditions and then inspected the traces, which provides information that the final-answer score alone does not. On the original release, the authors reported that Claude Opus 4.6 solved 77% of the problems that at least one expert solved and 23.5% of the problems that none of the expert panel solved. Mythos Preview reached 30% on that difficult subset. Success was less stable there: 44% of correct answers appeared only once or twice in five attempts, compared with 9% on the expert-solvable tasks.10

GeneBench-Pro: reach an exact estimate through messy synthetic data

OpenAI’s GeneBench-Pro contains 129 quantitative-genomics problems across ten areas, with ten released publicly.9 Each problem requires three to thirteen linked decisions and ends in an exact numerical or structured answer. The benchmark uses synthetic data designed around analysis patterns drawn from the literature. The files resemble untidy biological records while preserving a known answer and plausible wrong alternatives. Eleven external domain experts reviewed candidate problems; 82 problems, including all ten public tasks, received external review. Synthetic construction supports exact grading, but it cannot reproduce every irregularity of a real experiment. Models work through the supplied files in an isolated coding environment with the internet disabled. A deterministic grader checks the submitted fields against tight tolerances. The authors ran many attempts per problem and reported aggregated performance. Office Hours reproduced the sealed, model-only condition for the ten public tasks and preserved each trace so the failed decision could be examined rather than inferred from the score. On the full 129-problem set, the authors reported that GPT-5.6 Sol passed 28.7%, rising to 31.5% with additional computation, while Claude Opus 4.8 passed 16.0%. The paper found that models frequently noticed local clues but failed to turn them into the correct analytical decision.9

BiomniBench-DA: assess the analytical process as well as the answer

BiomniBench-DA contains 100 biomedical data-analysis tasks across 17 analysis types, with 50 released publicly.34 The agent must produce both a scientific answer and a record of the analysis used to reach it. Each task begins with a real research paper and its supplementary data. Experts write the question, reference analysis and scoring rubric. The rubric covers data handling, method choice, statistics, interpretation and source reliability. Half of the tasks are held out. Agents work inside a container with the paper’s data and scientific software. A model judge scores the full trace against the expert rubric on a 0–100 scale. Across the full benchmark, the authors’ highest reported configuration, Claude Code with Opus 4.7, had a mean score of 73.34. For our task-level analysis, Office Hours used the benchmark’s Harbor environment and grader and treated an individual task score of 70 or higher as passing. This design can distinguish a correct-looking endpoint from an analysis that used the wrong unit, skipped a quality check or selected an unsuitable test. Changing the agent software shifted results by as much as 13.5 points, more than several model-generation changes in the authors’ comparison.34

CompBioBench: complete a public computational-biology task

Roche/Genentech’s CompBioBench contains 100 tasks across genomics, transcriptomics, epigenomics, single-cell analysis, human genetics and machine learning.35 Tasks include database retrieval, identifier conversion and short computational analyses that end in one exactly graded answer. Computational biologists contributed problems drawn from recurring scientific work. The benchmark combines real, modified and synthetic data and keeps the official answers private. For the seven self-contained tasks Office Hours ran, we independently re-derived each answer from the named public database or reference resource. Coding agents begin in a fresh workspace with internet access, install or call the tools they need and return one exact string. Office Hours used the same coding-agent condition. Because the official answers are held for leaderboard submission, our local grading used the independently reproduced answers and inspected the route taken to obtain them. The authors reported that frontier coding agents scored 78–83% across the full benchmark. On the harder subset, performance ranged from 49% to 69%. The results support a specific claim: agents are already effective on many bounded, verifiable computational-biology tasks, while leaving room on compound and less forgiving work.35

Terminal-Bench Science: produce scientific artifacts that pass isolated tests

Terminal-Bench Science asks agents to complete research-level scientific workflows inside a terminal and produce an artifact that passes hidden tests.36 The growing collection includes biological imaging, sequencing quality control and medical-image reconstruction. Scientist contributors write the task, stage the data and tools inside a container and create the hidden tests. Each task receives human review for realism, clarity and solvability before release. A coding agent works inside the supplied container, usually for up to two hours, under the task’s stated internet policy. Office Hours ran three released life-science tasks using the benchmark’s supplied containers and verifiers. A full pass requires the produced artifact to satisfy the hidden checks. The Science track was still expanding in July 2026 and had not published a model leaderboard. Its contribution is the execution format: a scientific result must exist as a testable file or analysis, not only as a persuasive explanation.

Across these examples, the recurring problem was not an inability to calculate. It was proceeding before checking whether the labels, samples, assay behavior, structural artifacts, coordinate system or unit of analysis supported the calculation. This was not universal: individual systems caught some problems. But the traces repeatedly showed that a decisive scientific check occurred before the polished analysis, which is why realistic evaluations should score data inspection and preserve the route to the answer.

Figure 3

What our evaluation runs revealed about scientific reliability

Across selected tasks from BioMysteryBench, GeneBench-Pro, BiomniBench-DA, CompBioBench and Terminal-Bench Science, systems frequently completed useful parts of the work. The recurring failures appeared at the scientific decisions that determine whether an analysis can be trusted.

  1. 1

    Inspect the inputs

    Confirm labels, sample identity, contamination, coordinate systems and whether the measurement can be trusted.

    Observed: Systems could analyse accepted files but were less reliable at detecting why an input should not be trusted.

  2. 2

    Choose the method

    Select an analysis that preserves the signal and matches the biological unit, assay and question.

    Observed: Agents identified plausible knockout candidates but sometimes selected the wrong gene by over-weighting literal zero counts or the largest fold change instead of replicate consistency and statistical support.

  3. 3

    Complete the analysis

    Carry dependent steps through the confirming calculation instead of inferring the answer from the setup.

    Observed: Runs often progressed far but missed one correction that changed the final quantity or conclusion.

  4. 4

    Interpret the biology

    Explain what the result supports without letting expected biology replace the evidence in the supplied data.

    Observed: Biological interpretation was often stronger than input checking, but could still outrun the completed analysis.

  5. 5

    Verify the conclusion

    Check that the final output is biologically correct, reproducible and precise enough for the requested decision.

    Observed: Sophisticated DAPI–H&E alignment still produced incorrect one-to-one cell identities.

  • Catching imperfect data

    Model-only runs

    522%

    Coding-product runs

    2348%

    Imperfection is not synonymous with error. The difficult decision is whether a mislabel, contamination or assay problem requires correction, or whether an unexpected pattern is the biological result.

  • Choosing the method

    Model-only runs

    2639%

    Coding-product runs

    4559%

    Coding products benefited from tools, but plausible methods could still discard the signal or use the wrong biological unit.

  • Interpreting the biology

    Model-only runs

    4769%

    Coding-product runs

    1857%

    Systems could often form a plausible biological interpretation once they had a result. The risk was interpreting that result before unresolved data, method or statistical errors had been corrected.

  • Using the right tool

    Model-only runs

    3882%

    Coding-product runs

    86100%

    Tool use was a relative strength. The remaining question was whether the correct record, target and intermediate result were used.

These are descriptive ranges across available setups, not rankings or confidence intervals. Each range combines the available results for the individual models or coding products represented in that row. Systems represented: GPT-5.6 Sol, Claude Opus 4.8, Gemini 3.1 Pro and Grok 4.5, together with Codex, Claude Code, Gemini CLI and Grok Build where runs were completed. We retain each system’s results in our detailed analysis, but report shared ranges here. Model-only and coding-product runs used different conditions and are shown separately. Missing observations were not counted as failures.

Models could often perform the computation. Reliability depended on knowing when to correct the inputs and when to preserve an unexpected biological result, a judgment that can challenge human experts too.

Real biological data are difficult for humans and models

Biological experiments produce sequences, images, count matrices, sample records and instrument files rather than one standardized input. An unexpected pattern may come from a mislabeled sample, contamination or an assay problem, but it may also be the true result of an experiment. Distinguishing those possibilities is slow and difficult even for experienced scientists. Models should not be rewarded for forcing unfamiliar data toward the textbook answer; they should inspect the evidence, ask for missing information and preserve a surprising result when the experiment supports it. Synthetic data are useful when exact grading is essential, but real experimental data are needed to test this distinction.

Office Hours “BioLongHorizon”: expert-led evaluations drawn from real biological problems

Office Hours is developing “BioLongHorizon,” an expert-led evaluation program built from problems that arise in scientists’ own work. The tasks require models to examine the same kinds of data and evidence scientists use, including FASTA sequence files, gene-expression count matrices, cell annotations, sample metadata, figures and supporting literature; select appropriate tools and methods; and carry the analysis through the decisions needed to reach a scientifically supported conclusion. Our initial runs echoed patterns observed across released benchmarks: models could complete substantial parts of an analysis while missing a decisive step, such as checking an unexpected input, choosing the method that fits the biological unit or following the evidence when it conflicts with the expected result. The four examples below illustrate the range of work in development; they are not the full task set.

These tasks are difficult because the calculation is only part of the work. Each requires a model to make a scientific decision that determines whether the calculation answers the right question. Each evaluation includes a model-facing prompt, scientific input and reference files, a defined software environment, a required output, an expert solution and scoring criteria. This allows Office Hours to examine not only whether a model reached the answer, but which data, tools and scientific decisions produced it.

  1. Office Hours Task 1

    Determining whether a tumor variant meets a treatment-eligibility bar

    The first task asks a model to examine tumor DNA and RNA evidence, identify a change in a tumor-suppressor gene and decide whether the evidence is strong enough to classify it as disease-causing for a treatment-eligibility decision. A plausible answer cannot come from naming a variant alone. The analysis has to connect the genomic change to its effect on RNA splicing and compare that effect with the range seen in normal tissues.

    This is challenging because the familiar shortcut points toward the wrong decision. The variant alters the invariant first base of a splice-donor site in TSC2, which could easily trigger an automatic loss-of-function classification. But exon 26 is 129 bases long, so skipping it preserves the reading frame, and the normal-tissue reference shows that this exon is already absent from most TSC2 transcripts. The model must align raw DNA and RNA reads, establish the genomic coordinate, quantify the splice junctions and compare the tumor with normal tissue before deciding whether the evidence meets the trial’s pathogenic or likely-pathogenic eligibility requirement. It tests whether a system can resist a standard variant-interpretation shortcut when the patient’s actual molecular evidence points elsewhere.

  2. Office Hours Task 2

    Finding which T-cell population responds to checkpoint treatment

    The second task uses single-cell gene-expression and T-cell receptor data collected before and during PD-1 treatment, together with a control group. It asks which T-cell state shows a treatment-specific increase in cell division. Solving it requires linking cells that share a T-cell receptor, tracking those cell families across time and separating a treatment-associated change from ordinary variation between patients and samples.

    The difficulty begins with reconstructing biological identity. A cycling T cell may no longer carry the transcriptional label that clearly identifies its earlier state, so the analysis must use its T-cell receptor sequence to connect it back to a lineage. It must also preserve patient identity: treatment was assigned to six patients, not to thousands of independent cells, and similar receptor sequences observed in different people cannot automatically be treated as one biological clone. Our subsequent audit found that defining clones across the full cohort could carry cycling information from one patient into another and make the statistical evidence appear stronger than the patient-level data support. The evaluation therefore tests whether a model can reconstruct cell lineage while respecting the actual experimental unit and representing uncertainty honestly.

  3. Office Hours Task 3

    Measuring immune-signalling activity in cytotoxic T cells

    The third task asks a model to combine cell-type labels with pathway-activity scores, define the complete cytotoxic CD8 T-cell population and compare cells with higher and lower JAK–STAT signalling. The calculation is short, but only after the model joins the files correctly, restricts the analysis to cells that were actually scored and recognizes that the relevant population appears under more than one label.

    The arithmetic is straightforward only after the biological population has been defined correctly. The automated annotation divided the relevant cytotoxic cells between “Effector CD8+ T cells” and “CD8+ NKT-like cells.” Every recorded model attempt performed a plausible calculation on only part of that population or joined the data incorrectly. The task therefore isolates an important failure: a model can execute the requested statistics correctly and still answer the wrong biological question because it trusted one label without examining the complete dataset. An expert must recognize the label ambiguity, combine the relevant populations, retain only cells with pathway scores and then perform the comparison.

  4. Office Hours Task 4

    Auditing macrophage data before differential-expression analysis

    The fourth task asks a model to compare gene expression in human macrophages under resting and inflammatory conditions. Before running the statistical test, the model must notice that the sample labels conflict with the expression patterns, determine whether a label error occurred and use raw read counts with a method suited to count data. This tests whether the agent audits the experiment before producing a polished list of changed genes.

    This task presents a problem researchers encounter in practice: metadata can conflict with molecular data, and resolving one inconsistency can reveal a deeper limitation in the study design. Here, two sample labels conflicted with the expression patterns; after correcting them, treatment could not be separated from donor identity. A reliable system should not force such data through a differential-expression pipeline and return a polished gene list. It should identify the conflict, explain what the data can and cannot support and state what additional evidence would be needed. The example also shows why evaluation development benefits from several forms of expertise: the task author contributes the biological problem and source data, while additional scientific and statistical review ensures that the question, expected answer and scoring support the same conclusion.

Together, these tasks show what expert-led evaluation adds beyond a final score. Scientists who know the work can recognize when metadata conflict with the data, when a plausible method answers the wrong biological question and when an unexpected result should be preserved rather than corrected away. Office Hours combines that task-specific expertise with repeatable model runs and scientific review to identify why a system failed and what should improve next.

Training data: from evaluation failures to model improvement

An evaluation is most useful when it identifies a specific model failure and shows what should improve. Before creating training data, the failure must be separated from problems in the task, grader, scientific tools, input data, access or run environment. Anthropic’s viral-sequence evaluation illustrates the distinction: much of the variation arose from biological-data infrastructure, and performance rose to nearly 100% after a deterministic retrieval layer was added.17 The appropriate response was a better retrieval tool, not training examples that taught the model to work around unreliable infrastructure.

When a repeatable failure is traced to the model, training data should target the missing capability under realistic conditions. Depending on the failure, this may require expert-worked analyses, annotated images, examples of scientific tool use, multi-turn conversations or corrections that explain why a plausible approach was wrong. If systems fail on low-quality bacterial microscopy, for example, clean textbook images will not help. Training examples should preserve uneven staining, poor focus, debris, crowded fields, mixed morphologies and variation between strains or growth conditions. A microbiologist can identify which features support a conclusion, when the image is insufficient and what additional evidence would resolve the uncertainty.

Each verified failure can then become an expert-reviewed training example containing the original question and files, the route that failed, the expert correction, the correct or golden answer and the uncertainty that remains. Improvement should be tested on new, meaningfully different problems. The goal is not to teach the answer to one benchmark, but to strengthen the underlying biological skill that the evaluation revealed was missing.

Figure 4

An observed failure becomes useful only after its source is resolved

  1. Step 1

    Observed failure

    Preserve prompt, tools, data, trace and output.

  2. Step 2

    Validity check

    Can a qualified person solve the task as written?

  3. Step 3

    Failure source

    Separate model, biology, data, tool, run-setup, grader and policy errors.

  4. Step 4

    Choose the response

    Repair the task or tool, change the workflow or create reviewed training data.

  5. Step 5

    Independent retest

    Use new, non-identical problems before claiming improvement.

A failed run does not identify its own remedy. Task defects, missing tools and model limitations require different responses. Expert-generated correction data are valuable only after the task and grader have been validated, the failure source has been identified and improvement is shown on new problems that do not reuse the same surface form.

Complete task package

What an expert biology evaluation task should contain

A task is more than a question and an expected answer. It begins with a clear scientific goal: the biological question or decision being tested, why it matters and what a successful result would establish. The task is then assembled as a reviewable scientific package that lets another qualified person understand what the model was asked to do, reproduce the intended route, recognize defensible alternatives and determine why a response succeeded or failed.

  1. 1

    Prompt

    The exact instruction shown to the model, including the requested output and any limits that are part of the task.

  2. 2

    Input and reference files

    The data, images, papers, metadata and supporting records needed to answer the question, with versions and clear file descriptions.

  3. 3

    Allowed tools and expected workflow

    The databases, websites, software, code or connectors the model may use. When relevant, the task should also specify the expected workflow, such as finding the versioned NCBI record before designing or analysing anything.

  4. 4

    Correct answer / golden answer

    The correct, verifiable objective answer when one exists, or an expert-written golden answer that identifies the scientifically defensible conclusions, with uncertainty stated rather than hidden.

  5. 5

    Expert solution record

    The evidence, calculations and important decisions that support the correct or golden answer, including why tempting alternatives fail. It should be reproducible without requiring a private stream of consciousness.

  6. 6

    Scoring rubric

    The scientific criteria for credit, process and outcome checks, rules for serious errors and treatment of defensible alternatives. The rubric remains hidden unless the evaluation is intentionally designed to show it to the model.

Evaluation environment

The computer environment or sandbox in which the task is run: available files, installed software, internet and database access, permissions, time or memory limits, starting state and required output location. Another evaluator should be able to recreate the same conditions.

Independent expert review surrounds the entire package

A second expert should attempt the task in the evaluation environment before seeing the author’s solution, then review the scientific goal, prompt, files, expected workflow, correct or golden answer, expert solution and rubric. The task is revised or rejected when the question is ambiguous, an input is missing, the answer is too narrow or the scoring rewards a scientifically weak route.

Conclusion and future directions

The field does not need biology questions that are difficult for their own sake. It needs evaluations built around real questions that practicing scientists bring to frontier models. These evaluations should make biological assumptions visible and explain why a model succeeded or failed.

Not every benchmark needs to be large. A small set of carefully chosen tasks can be more informative than a large examination when each task comes from real scientific work, is written by a specialist whose experience matches the problem, is attempted and reviewed independently, and is run repeatedly to reveal not only whether a model failed but why.

Frontier models can already retrieve scientific information, complete bounded database and analysis tasks, reason through difficult biological problems and contribute to molecule, hypothesis and experimental design. These are meaningful advances. Foundational knowledge, published literature, dependable databases, engineering and physical validation remain necessary, but they are not sufficient for systems expected to participate in real research.

The next advances increasingly depend on scientists and physicians whose direct work matches the problem: people who know which questions matter, which controls distinguish competing explanations, how real data fail and when the evidence does not support a confident answer.

Their role is not to make evaluations difficult. It is to design scientifically consequential tasks drawn from real work, run them under realistic conditions, review the work independently and determine whether an observed failure calls for better training data, a better tool, a revised evaluation or a new scientific investigation. Office Hours builds this program around expert-led evaluations grounded in how experienced scientists and physicians actually use frontier models, followed, when the evidence supports it, by training data built from verified failures and expert corrections. The tasks should reflect real problems that matter, with the data, tools, constraints and uncertainty of scientific work. The aim is not to replace scientific judgment, but to use years of subject knowledge and practical experience to show what models can do reliably, where they fail and what should improve next.

Work with Office Hours

Design and review biology evaluations with practicing experts

Office Hours connects frontier AI laboratories with scientists and physicians whose direct research experience matches the task to build evaluations, understand model failures and create expert-level training data.

Explore Expert Data Services

References

43 sources

  1. 1.

    CASP16 assessment of protein-structure prediction · Proteins (2026)

  2. 2.

    Runs N’ Poses: temporal-holdout evaluation of cofolding models · Nature Structural & Molecular Biology (2026)

  3. 3.

    FoldBench: all-atom structure evaluation under low-homology conditions · Nature Communications (2025)

  4. 4.

    A critical assessment of single-cell foundation models · Nature Machine Intelligence (2024)

  5. 5.

    Pretraining-scale and baseline comparisons in single-cell models · Nature Methods (2026)

  6. 6.

    Virtual Cell Challenge 2025 · Cell (2025)

  7. 7.

    BixBench: a benchmark for agents performing bioinformatics analyses · 2025

  8. 8.

    ScienceAgentBench: toward rigorous assessment of language agents for data-driven discovery · ICLR (2025)

  9. 9.

    GeneBench-Pro: expert-audited and tiered-release biology-agent evaluation · 2026

  10. 10.

    BioMysteryBench · Anthropic (2026)

  11. 11.

    Towards multimodal foundation models in molecular cell biology · Nature (2025)

  12. 12.

    Generalist biological artificial intelligence in modeling the language of life · Nature Biotechnology (2026)

  13. 13.

    AI-driven protein design · Nature Reviews Bioengineering (2025)

  14. 14.

    Zero-shot design of drug-binding proteins via neural iterative selection–expansion · Nature (2026)

  15. 15.

    Target identification and assessment in the era of AI · Nature Reviews Drug Discovery (2026)

  16. 16.

    Computational drug repurposing: approaches, evaluation of in silico resources and case studies · Nature Reviews Drug Discovery (2025)

  17. 17.

    Paving the way for agents in biology, including VirBench · Anthropic (2026)

  18. 18.

    Accelerating scientific discovery with Co-Scientist · Nature (2026)

  19. 19.

    A benchmark of expert-level academic questions to assess AI capabilities · Nature (2026)

  20. 20.

    LifeSciBench: expert-written evaluation of realistic life-science research tasks · OpenAI (2026)

  21. 21.

    FlowBench: separating planning, fault recovery and interpretation in agentic bioinformatics · 2026

  22. 22.

    Contemporary AI lacks the imagination to diverge or negate in science · 2026

  23. 23.

    Proof of Time: a benchmark for evaluating scientific idea judgments · 2026

  24. 24.

    Uncovering evolutionarily remote and highly potent antimicrobial peptides with protein language models · Nature Biomedical Engineering (2026)

  25. 25.

    Molecular deep learning at the edge of chemical space · Nature Machine Intelligence (2026)

  26. 26.

    A multi-agent system for automating scientific discovery · Nature (2026)

  27. 27.

    About 30% of Humanity’s Last Exam chemistry and biology answers are likely wrong · FutureHouse (2025)

  28. 28.

    Suggested reporting language, interpretation and guidance for Lyme disease serologic testing results · CDC and Association of Public Health Laboratories (2024)

  29. 29.

    Clinical practice guidelines for the diagnosis and management of babesiosis · Infectious Diseases Society of America (2020)

  30. 30.

    Trends in reported babesiosis cases, United States, 2011–2019 · MMWR (2023)

  31. 31.

    LAB-Bench: measuring capabilities of language models for biology research · 2024

  32. 32.

    BenchGuard: automated auditing of language-agent benchmarks · 2026 preprint

  33. 33.

    BioMysteryBench preview dataset changelog, version 11 · Anthropic (2026)

  34. 34.

    BiomniBench-DA: process-level evaluation of biomedical data-analysis agents · 2026 preprint

  35. 35.

    CompBioBench: evaluating agents on computational-biology tasks · 2026 preprint

  36. 36.

    Terminal-Bench Science: verifiable scientific tasks in terminal environments · 2026

  37. 37.

    LLMs get lost in multi-turn conversation · ICLR (2026)

  38. 38.

    LABBench2: an improved benchmark for measuring AI in biology research · Edison Scientific (2026)

  39. 39.

    PCR using 3′-mismatched primers to detect an A2142C mutation in 23S rRNA conferring resistance to clarithromycin in Helicobacter pylori clinical isolates · Journal of Clinical Microbiology (2000)

  40. 40.

    Suppressed RefSeq record NC_000915.1 for Helicobacter pylori 26695 · National Center for Biotechnology Information

  41. 41.

    scBench-Long: verifiable benchmarking of long-horizon single-cell biology · LatchBio manuscript (2026)

  42. 42.

    DeepSeek-R1 incentivizes reasoning in large language models through reinforcement learning · Nature (2025)

  43. 43.

    Crossing the reward bridge: expanding reinforcement learning with verifiable rewards across diverse domains · ACL (2026)

Review method. Sources were reviewed through July 28, 2026.

Acknowledgements

We thank the Office Hours experts who contribute their time, deep subject knowledge and practical scientific judgment to improving how frontier AI is evaluated in biology and medicine. We are especially grateful to the expert scientists and physicians who developed and reviewed the biology tasks examined in this review. Their task design, source materials and careful scientific reasoning made this task-level analysis possible.

Office Hours Research · Biology & Life Sciences