RExBench
12 tasksBest result with human-written hints
DECISION VIEW / AI R&D
Three complementary benchmark perspectives: implementing research extensions, proposing and testing new methods, and completing an open research workflow.
Best result with human-written hints
Share of the gap to top human performance closed
Example of invalid results in one agent setup
MLR-Bench includes the ideation stage, but within a bounded benchmark framework.
Stage evidence ↗The sources were published on different dates, but measure different tasks, evaluators and denominators. The module therefore has no defensible longitudinal performance trend yet; the next comparable benchmark release must be added to the same track before drawing a trend.
Objectively measured, but on only seven competition tasks.
Use the map to distinguish impressive partial results from reliable, objectively verified end-to-end research.
Three benchmark families open to peer scrutiny illuminate different parts of the R&D loop, but cannot be aggregated or read as one time series.
Claim evidence · research-automation/claim/evidence-state ↓RExBench v3 reports below 44% even with human-written hints on 12 realistic extensions.
RExBench: best result with hints · Research extensions; not a general R&D rate.
RExBench coverage · Bounded benchmark.
The best MLRC-Bench agent closed 9.3% of the gap on seven objectively evaluated competition tasks.
MLRC-Bench: gap closed · Objective performance measurement against a baseline and top human performance.
MLRC-Bench coverage · Dynamic ML competition tasks.
MLR-Bench covers 201 tasks and four stages, but reports serious problems with experimental validity.
MLR-Bench coverage · End-to-end benchmark design; not a single comparable score.
MLR-Bench: experimental errors · Example figure from a specific coding-agent setup.
Across different evaluation approaches, implementation and experimental control remain clear weaknesses.
RExBench: best result with hints · Research extensions; not a general R&D rate.
MLRC-Bench: gap closed · Objective performance measurement against a baseline and top human performance.
MLR-Bench: experimental errors · Example figure from a specific coding-agent setup.
There is no defensible common denominator for an overall automation percentage.
Overall AI R&D automation · Aggregate deliberately left empty.
Twelve research-extension tasks and reported agent evaluation; do not generalize to all research work.
Open original source ↗Open-ended ML research benchmark; the reported coding-agent failure example is not a cross-benchmark rate.
Open original source ↗Seven objectively scored ML research competition tasks; kept separate from RExBench and MLR-Bench.
Open original source ↗The current arXiv v3 abstract, revised 2026-04-21, reports about 33% autonomous success and says the best hinted result remains below 44%. This supersedes the legacy candidate whose hash represented the identical raw page bytes rather than the normalized visible-content fingerprint required by source admission.
Open original source ↗Claims, calculations, sources and known limitations for this view. Source data and dashboard interpretation are kept separate.
RExBench v3 reports that the best result, even with human-written hints, is below 44% on realistic research extensions.
research-automation/claim/rexbench-capabilityRExBench contains 12 research implementation tasks with automated evaluation.
research-automation/claim/rexbench-scopeMLR-Bench describes frequent fabricated or invalidated experimental results, for example 80%, for the tested coding-agent setup.
research-automation/claim/mlr-reliabilityMLR-Bench covers 201 tasks and a four-stage workflow from ideation to paper writing.
research-automation/claim/mlr-scopeIn MLRC-Bench, the best tested agent closed 9.3% of the gap between the baseline and top human performance.
research-automation/claim/mlrc-capabilityMLRC-Bench uses seven dynamic ML competition tasks with objective performance measures.
research-automation/claim/mlrc-scopeTogether, the three benchmarks point to implementation and experimental validity as a clear limitation, but they cannot be combined into one automation share.
research-automation/claim/experimental-bottleneckThe module publishes no overall R&D automation percentage because the benchmarks do not measure the same task, agent or success criterion.
research-automation/claim/no-aggregateThe current benchmarks do not establish a complete, independent validation and reproduction loop after experimentation and interpretation.
research-automation/claim/validation-reproduction-gapThe hypothesis stage has partial benchmark coverage: MLR-Bench includes proposals in its workflow, while RExBench defines twelve research extensions as concrete tasks.
research-automation/claim/hypothesis-stage-evidenceMLR-Bench includes experimentation in its workflow, but also reports seriously invalidated results for one tested agent setup.
research-automation/claim/experiments-stage-evidenceResearch Automation has three sourced benchmark families, but zero longitudinal points within the same benchmark version. RExBench, MLR-Bench and MLRC-Bench have different tasks, evaluators and denominators and therefore establish no common trend.
research-automation/claim/evidence-stateresearch-automation/fact/rexbench/best-with-hintsBest RExBench performance with human-written hints
Edwards et al. · v3 abstract revised 2026-04-21research-automation/fact/rexbench/tasksRExBench realistic research extension tasks
Edwards et al. · Abstractresearch-automation/fact/mlr/tasksMLR-Bench open-ended ML research tasks
Chen et al. · Abstractresearch-automation/fact/mlr/stagesMLR-Agent research workflow stages
Chen et al. · Abstract and workflow descriptionresearch-automation/fact/mlr/invalid-resultsCoding-agent fabricated or invalidated experiment results
Chen et al. · Abstractresearch-automation/fact/mlrc/tasksMLRC-Bench competition tasks
Zhang et al. · Abstractresearch-automation/fact/mlrc/gap-closedBest agent share of baseline-to-top-human gap closed
Zhang et al. · Abstractresearch-automation/fact/coverage/no-aggregateCross-benchmark aggregate automation score
Edwards et al. · Module methodology; intentional non-derivationresearch-automation/fact/coverage/validation-reproduction-gapIndependent validation and reproduction coverage
Chen et al. · Module coverage audit against the published four-stage workflowThe results are not compatible as a time series.
Different evaluators introduce different sources of error.
None of the benchmarks observes secret or commercial laboratory practice.
Three separate benchmark perspectives; experimental validity is the clearest bottleneck