Toward a rigorous, extensible framework for evaluating AI as a scientific collaborator
As large language models (LLMs) rapidly advance, their potential to assist or even independently conduct scientific research has captured wide attention. From summarizing literature to generating code, LLMs already perform parts of the research process. Yet their credibility, validity, and replicability as scientific contributors remain largely untested.
The Benchmarking LLM Agents on Scientific Tasks project, led by the Center for Open Science (COS) with support from Coefficient Giving, is a multi-year, multi-team initiative to systematically evaluate how well LLM agents can perform and reason through the scientific research lifecycle.
The project defines a structured framework that benchmarks LLM agents according to three core researcher personas, each representing distinct capabilities and roles within the research ecosystem:
Each persona is mapped onto stages of the research lifecycle—design, conduct, analyze, and interpret—producing a scalable matrix of tasks that can be used to benchmark progress in increasingly complex and autonomous forms of scientific reasoning.
The framework is designed to be expandable: as model capabilities evolve, new task layers can test higher-order reasoning, methodological generalization, and the ability to autonomously connect multiple stages of research into an integrated workflow.
ReplicatorBench: The first active phase of this initiative has produced ReplicatorBench, a benchmark for evaluating LLM agents on research replication in the social and behavioral sciences.
What sets ReplicatorBench apart from prior work is its focus on the full replication process, not just the final outcome. ReplicatorBench requires agents to first preregister their replication plan, then execute their analysis according to that plan, mirroring the standards of rigorous human research. It assesses agents across three stages: extracting and retrieving the right information, designing and executing the analysis, and interpreting results against the original claims. Read more on our blog.
Building on ReplicatorBench, the team is now developing a Robustness benchmark. Replication asks whether a result holds when tested with new data but a separate and equally important question is whether results hold when the same data are analysed in different, equally justifiable ways. Researchers routinely face genuine degrees of freedom in how they operationalize variables, preprocess data, and select analytical approaches, and reasonable choices at each of these decision points can lead to meaningfully different conclusions.
The Robustness benchmark evaluates whether LLM agents navigate this analytical landscape in ways that are consistent, justified, and transparent or whether their outputs are as variable as those of human analysts working on the same task. Results and protocols will be released openly as the benchmark matures.
Subsequent phases will extend the framework to the Evaluate and Generate benchmarks:
Together, these benchmarks will create a continuous, community-driven infrastructure for testing the limits of AI as a scientific collaborator, identifying both capabilities and failure modes, and guiding responsible deployment in research.
The project is actively expanding. Opportunities for collaboration, including contributing annotations, evaluation rubrics, or replication studies, are welcome. Researchers and organizations interested in learning more or contributing in the future can stay updated through the COS website and project communications.
Please contact Tim Errington (tim@cos.io) or Shakhlo Nematova (shakhlo@cos.io) for more information.

6218 Georgia Avenue NW, Suite #1, Unit 3189
Washington, DC 20011
Email: contact@cos.io

Unless otherwise noted, this site is licensed under a Creative Commons Attribution 4.0 International (CC BY 4.0) License.
Responsible stewards of your support
COS has earned top recognition from Charity Navigator and Candid (formerly GuideStar) for our financial transparency and accountability to our mission. COS and the OSF were also awarded SOC2 accreditation in 2023 after an independent assessment of our security and procedures by the American Institute of CPAs (AICPA).
We invite all of our sponsors, partners, and members of the community to learn more about how our organization operates, our impact, our financial performance, and our nonprofit status.