AI Researchers Got 85x More Efficient and Invented Nothing
TasteVal finds AI experimental taste doubling every 3 months since Dec 2025. The same paper labeled 2,000+ model experiments and found zero invented ideas.
By the paper's central estimate, Opus 5.5 can now reach a human expert's best result on a hard ML task in less than half the GPU time the expert needed. That is the headline number in TasteVal, a benchmark Oliver Jaffe and Dane Sherburn of P-Zero Research posted to arXiv on October 5. The more useful number sits in section 3.2. The authors labeled 540 official submissions from 18 models and found zero that contained an invented component. They checked roughly 1,600 more experiments from the top-decile runs, including failed and rejected ones, and found none there either.
So the fastest-improving measure of AI research skill in the paper tracks how quickly models reach known methods. A forecast that plugs the TasteVal trend into a model of self-improving AI inherits that limit.
What TasteVal measures
The setup splits the work. The model under test plays Researcher: it reads the task, proposes experiments in plain language, and interprets results. A fixed Coder (Opus 4.8) writes and runs the code on a single H100. The Researcher cannot execute code or see the test set. Each run gets 40 H100 hours and 120 wall-clock hours. The eight tasks span curating pretraining data, pretraining and fine-tuning language models, modeling human preferences, and aligning models against adversarial prompts. The tasks themselves are withheld to keep them uncontaminated.
Taste is scored as compute efficiency. If the expert baseline needs 40 GPU hours to reach a score and a model gets there in 20, the model has a 2x compute multiplier. The baseline is the best of at least two attempts per task from 24 experts who had recently worked at organizations including OpenAI, Google DeepMind, NVIDIA, Microsoft Research, and the University of Oxford. They were not allowed AI help for designing experiments or interpreting results.
Across 20 models released from 2023 to 2026, the multiplier rose about 85-fold, from 0.03 for GPT-4 to 2.30 for Opus 5.5 (95% CI 1.15 to 4.37). A Chow test puts a trend break at December 11, 2025, the release of GPT-5.2. Before it, the multiplier doubled every 14.0 months. Since then it has doubled every 3.0 months (95% CI 1.7 to 5.0).
The second metric is easier to miss. The performance multiplier ignores compute and asks how good the final submission is. It shows no significant break (p = 0.12) and doubles every 14.6 months. Final results kept roughly their pre-break pace while time-to-result sped up.
Zero invented
The authors sorted official submissions into four classes. Tuned follows an existing recipe, including hyperparameter sweeps. Composed combines published components used as published. Modified structurally changes at least one imported component. Invented contains at least one structural component that does not exist in the literature.
Composed submissions are the majority in every compute-multiplier bucket from 0.5x up. Modified submissions are rare before GPT-5.6 Sol, at most one in 30 per model. Invented is empty. The human experts produced no Invented submissions either: 20% of their official submissions were Tuned, 72% Composed, and 8% Modified.
The limitations section is direct about what this means. Progress on TasteVal "has primarily been driven by increasing ability to perform shallow and moderate optimization," the authors write, and "progress may slow if sustaining the current trend requires progressively deeper optimization, such as proposing novel ideas."
The transcripts show where models still trail the experts. The strongest human runs measured run-to-run spread in their first experiment and later discarded any single-run improvement smaller than it. Fable 5.1 rarely re-ran an experiment with a different seed. It also bundled several training runs into one experiment, while humans usually made a single change per experiment.
Where the forecast leans
Section 5 substitutes the post-break rate into two forecasting frameworks. In the AI Futures Model, the probability of a taste-only singularity rises from 51% of simulations to 88%, and median ASI arrival moves from 2030.5 to 2028.9. The authors label these "naive extrapolations, not forecasts" and list their assumptions. One of them, A2, assumes the parts of research taste TasteVal does not measure, such as deciding which problems are worth solving, improve at the same rate per unit of general capability as the part it does measure.
The zero-invented result is the reason to be careful with A2. A taste-only singularity needs AI to keep finding better ideas for building AI. TasteVal measures something narrower: how quickly a model reaches the expert's score by selecting and combining existing methods on single-GPU tasks with a clean validation signal. The extrapolation needs both skills to improve together, and the benchmark observes only the second.
Two more details belong next to that. In "messy" variants of four tasks where the scoring function was withheld, Opus 5.0 kept 49% of its normal multiplier and Opus 4.5 kept 40%, though with four tasks neither drop is statistically significant. The novelty labels also came from a model, Fable 5.0 with web search, so a different labeler could draw the Modified line elsewhere. An empty Invented bin across more than 2,000 labeled experiments is still hard to explain away.
What the number is good for
None of this makes the 3.0-month doubling small. Opus 5.5 beat the expert baseline at roughly one thirtieth of the cost per run: $282 on average against $9,123 for the experts, counting API spend, GPU time, and expert salary. A model that reaches a strong researcher's result in less than half the GPU time is already useful for work that consists mostly of sweeps and ablations.
Because the eight tasks stay private, the benchmark and its four-class novelty audit can be rerun on each new model without contamination.
