Benchmarking LLM Agents on Scientific Tasks

Toward a rigorous, extensible framework for evaluating AI as a scientific collaborator

Overview

As large language models (LLMs) rapidly advance, their potential to assist or even independently conduct scientific research has captured wide attention. From summarizing literature to generating code, LLMs already perform parts of the research process. Yet their credibility, validity, and replicability as scientific contributors remain largely untested.

The Benchmarking LLM Agents on Scientific Tasks project, led by the Center for Open Science (COS) with support from Coefficient Giving, is a multi-year, multi-team initiative to systematically evaluate how well LLM agents can perform and reason through the scientific research lifecycle.

A Multi-Persona Framework for AI in Science

The project defines a structured framework that benchmarks LLM agents according to three core researcher personas, each representing distinct capabilities and roles within the research ecosystem:

  • Replicate (Replicator): The agent’s ability to replicate published studies using shared data, code, and new data to test whether original results hold.
  • Evaluate (Peer Reviewer): The agent’s ability to critically evaluate completed research, identifying methodological strengths and weaknesses, assessing claims, and estimating replicability.
  • Generate (Discovery Scientist): The agent’s ability to propose novel research questions, design studies, conduct analyses, and interpret findings with scientific rigor.

Each persona is mapped onto stages of the research lifecycle—design, conduct, analyze, and interpret—producing a scalable matrix of tasks that can be used to benchmark progress in increasingly complex and autonomous forms of scientific reasoning.

The framework is designed to be expandable: as model capabilities evolve, new task layers can test higher-order reasoning, methodological generalization, and the ability to autonomously connect multiple stages of research into an integrated workflow.

Milestones

  • Project Launch: At COS, our work is about making research more transparent, rigorous, and verifiable. As AI tools enter the research workflow, we need evidence about what they actually contribute to scientific credibility. This project will evaluate LLM-based systems on real-world scientific tasks and compare their judgments to human baselines. Read more on our blog.

  • ReplicatorBench: The first active phase of this initiative has produced ReplicatorBench, a benchmark for evaluating LLM agents on research replication in the social and behavioral sciences.

    What sets ReplicatorBench apart from prior work is its focus on the full replication process, not just the final outcome. ReplicatorBench requires agents to first preregister their replication plan, then execute their analysis according to that plan, mirroring the standards of rigorous human research. It assesses agents across three stages: extracting and retrieving the right information, designing and executing the analysis, and interpreting results against the original claims. Read more on our blog.

 

Current Focus: Robustness Benchmark

Building on ReplicatorBench, the team is now developing a Robustness benchmark. Replication asks whether a result holds when tested with new data but a separate and equally important question is whether results hold when the same data are analysed in different, equally justifiable ways. Researchers routinely face genuine degrees of freedom in how they operationalize variables, preprocess data, and select analytical approaches, and reasonable choices at each of these decision points can lead to meaningfully different conclusions.

The Robustness benchmark evaluates whether LLM agents navigate this analytical landscape in ways that are consistent, justified, and transparent or whether their outputs are as variable as those of human analysts working on the same task. Results and protocols will be released openly as the benchmark matures.

Future Directions

Subsequent phases will extend the framework to the Evaluate and Generate benchmarks:

  • Evaluate: Assessing LLM agents’ ability to function as reviewers within open peer-review ecosystems, including alignment with expert judgment.
  • Generate: Testing LLMs’ capacity to autonomously design and conduct novel studies, grounded in preregistered planning and ethical guidelines.

Together, these benchmarks will create a continuous, community-driven infrastructure for testing the limits of AI as a scientific collaborator, identifying both capabilities and failure modes, and guiding responsible deployment in research.

Participation and Collaboration

The project is actively expanding. Opportunities for collaboration, including contributing annotations, evaluation rubrics, or replication studies, are welcome. Researchers and organizations interested in learning more or contributing in the future can stay updated through the COS website and project communications.

Please contact Tim Errington (tim@cos.io) or Shakhlo Nematova (shakhlo@cos.io) for more information.