Skip to content
SingularityOBSERVATORY
Menu
TRACKER/RESEARCH-AUTOMATION/BENCHMARKS

DECISION VIEW / AI R&D

What can research agents actually do?

Three complementary benchmark perspectives: implementing research extensions, proposing and testing new methods, and completing an open research workflow.

Research benchmark results

Independent metrics · no combined score

RExBench

12 tasks

Best result with human-written hints

<44 %Research extensions

MLRC-Bench

7 tasks

Share of the gap to top human performance closed

9.3%New ML methods

MLR-Bench

201 tasks / 4 stages

Example of invalid results in one agent setup

Experiment validity risk80 %Invalid experiments

Research workflow coverage

Select a stage
Problem selection

MLR-Bench includes the ideation stage, but within a bounded benchmark framework.

Stage evidence ↗

The sources were published on different dates, but measure different tasks, evaluators and denominators. The module therefore has no defensible longitudinal performance trend yet; the next comparable benchmark release must be added to the same track before drawing a trend.

Analysis summary Assessed 2025-06-27
OBSERVATION

9.3%

Best MLRC-Bench gap closed

Objectively measured, but on only seven competition tasks.

Use the map to distinguish impressive partial results from reliable, objectively verified end-to-end research.

PUBLIC EVIDENCE STATUSEARLY

Three benchmark families open to peer scrutiny illuminate different parts of the R&D loop, but cannot be aggregated or read as one time series.

Claim evidence · research-automation/claim/evidence-state ↓
LATEST OBSERVATION
2025-06-27
LAST SOURCE CHECK
2026-08-21
CADENCE
editorial
COMPATIBLE OBSERVATIONS
0longitudinal points within the same benchmark version
TREND STATUS
NOT ESTABLISHEDRExBench, MLR-Bench and MLRC-Bench use different tasks, evaluators and denominators.
Download machine-readable evidence ↓
IMPLEMENTATIONLIMITATION

Can agents extend existing research?

RExBench v3 reports below 44% even with human-written hints on 12 realistic extensions.

OBSERVATIONbelow 44%

RExBench: best result with hints · Research extensions; not a general R&D rate.

INTERPRETATION12 tasks

RExBench coverage · Bounded benchmark.

NEW METHODSLIMITATION

Can agents compete with top humans on new ML problems?

The best MLRC-Bench agent closed 9.3% of the gap on seven objectively evaluated competition tasks.

OBSERVATION9.3%

MLRC-Bench: gap closed · Objective performance measurement against a baseline and top human performance.

INTERPRETATION7 tasks

MLRC-Bench coverage · Dynamic ML competition tasks.

WORKFLOWSIGNAL

Is a complete research process covered?

MLR-Bench covers 201 tasks and four stages, but reports serious problems with experimental validity.

INTERPRETATION201 tasks / 4 stages

MLR-Bench coverage · End-to-end benchmark design; not a single comparable score.

OBSERVATIONoften 80%

MLR-Bench: experimental errors · Example figure from a specific coding-agent setup.

COMMON SIGNALLIMITATION

What limits reliable autonomous R&D?

Across different evaluation approaches, implementation and experimental control remain clear weaknesses.

OBSERVATIONbelow 44%

RExBench: best result with hints · Research extensions; not a general R&D rate.

OBSERVATION9.3%

MLRC-Bench: gap closed · Objective performance measurement against a baseline and top human performance.

OBSERVATIONoften 80%

MLR-Bench: experimental errors · Example figure from a specific coding-agent setup.

MEASUREMENT GAPMEASUREMENT GAP

What share of R&D is automated?

There is no defensible common denominator for an overall automation percentage.

MEASUREMENT GAPnot measured

Overall AI R&D automation · Aggregate deliberately left empty.

SOURCES / PROVENANCE4 source records
benchmark-resultEdwards et al.RExBench: Can coding agents autonomously implement AI research extensions?
Retrieved
2026-08-12
Data
2025-06-27
Licence
Cited public paper; no paper text redistributed

Twelve research-extension tasks and reported agent evaluation; do not generalize to all research work.

Open original source ↗
benchmark-resultChen et al.MLR-Bench: Evaluating AI Agents on Open-Ended Machine Learning Research
Retrieved
2026-08-12
Data
2025-05-26
Licence
Cited public paper; no paper text redistributed

Open-ended ML research benchmark; the reported coding-agent failure example is not a cross-benchmark rate.

Open original source ↗
benchmark-resultZhang et al.MLRC-Bench: Can Language Agents Solve Machine Learning Research Challenges?
Retrieved
2026-08-12
Data
2025-04-13
Licence
Cited public paper; no paper text redistributed

Seven objectively scored ML research competition tasks; kept separate from RExBench and MLR-Bench.

Open original source ↗
reviewed-research-sourceEdwards et al.RExBench paper
Retrieved
2026-08-13
Data
2026-04-21
Licence
citation-only-review-required

The current arXiv v3 abstract, revised 2026-04-21, reports about 33% autonomous success and says the best hinted result remains below 44%. This supersedes the legacy candidate whose hash represented the identical raw page bytes rather than the normalized visible-content fingerprint required by source admission.

Open original source ↗
VERIFIABLE EVIDENCE / 12 CLAIMSOpen evidence +

Claims, calculations, sources and known limitations for this view. Source data and dashboard interpretation are kept separate.

Observation

RExBench v3 reports that the best result, even with human-written hints, is below 44% on realistic research extensions.

Edwards et al.: RExBench paper ↗1 measurement1 source facts
  • Depends on the setup and hints.
research-automation/claim/rexbench-capability
Calculation

Together, the three benchmarks point to implementation and experimental validity as a clear limitation, but they cannot be combined into one automation share.

  • Qualitative triangulation, not a meta-analysis or common score.
research-automation/claim/experimental-bottleneck
Limitation

The module publishes no overall R&D automation percentage because the benchmarks do not measure the same task, agent or success criterion.

  • A measurement gap, not evidence of zero automation.
research-automation/claim/no-aggregate
Limitation

The current benchmarks do not establish a complete, independent validation and reproduction loop after experimentation and interpretation.

  • A coverage gap in the selected benchmarks, not evidence that such validation never takes place.
research-automation/claim/validation-reproduction-gap
Method

The hypothesis stage has partial benchmark coverage: MLR-Bench includes proposals in its workflow, while RExBench defines twelve research extensions as concrete tasks.

  • The sources use different tasks and evaluators and do not establish a common success rate.
research-automation/claim/hypothesis-stage-evidence
Method

Research Automation has three sourced benchmark families, but zero longitudinal points within the same benchmark version. RExBench, MLR-Bench and MLRC-Bench have different tasks, evaluators and denominators and therefore establish no common trend.

  • The benchmark families cannot be combined into one automation share.
  • Maturity describes evidence coverage, not general research autonomy.
research-automation/claim/evidence-state
TRACEABLE FACT LOG / 9 SOURCE FACTSOpen log +
research-automation/fact/rexbench/best-with-hints
below 44%

Best RExBench performance with human-written hints

Edwards et al. · v3 abstract revised 2026-04-21
research-automation/fact/rexbench/tasks
12

RExBench realistic research extension tasks

Edwards et al. · Abstract
research-automation/fact/mlr/tasks
201

MLR-Bench open-ended ML research tasks

Chen et al. · Abstract
research-automation/fact/mlr/stages
4: idea, proposal, experimentation, paper writing

MLR-Agent research workflow stages

Chen et al. · Abstract and workflow description
research-automation/fact/mlr/invalid-results
frequently, e.g. 80% of cases

Coding-agent fabricated or invalidated experiment results

Chen et al. · Abstract
research-automation/fact/mlrc/tasks
7

MLRC-Bench competition tasks

Zhang et al. · Abstract
research-automation/fact/mlrc/gap-closed
9.3%

Best agent share of baseline-to-top-human gap closed

Zhang et al. · Abstract
research-automation/fact/coverage/no-aggregate
not published

Cross-benchmark aggregate automation score

Edwards et al. · Module methodology; intentional non-derivation
research-automation/fact/coverage/validation-reproduction-gap
not established by current admitted benchmarks

Independent validation and reproduction coverage

Chen et al. · Module coverage audit against the published four-stage workflow
INTERPRETATION LIMIT

The results are not compatible as a time series.

Different evaluators introduce different sources of error.

None of the benchmarks observes secret or commercial laboratory practice.

MODULE STATUSINCOMPLETE

Three separate benchmark perspectives; experimental validity is the clearest bottleneck

VARDARK OBSERVATORY / RESEARCH-AUTOMATION / BENCHMARKS