Anthropic says Claude now leads 26% of its own AI R&D, up from under 1% in February
Anthropic published new internal measurements showing that Claude now "leads" about a quarter of the company's own AI research and development work, up from under 1% just six months earlier.
What's new
Anthropic built what it calls the Anthropic R&D Automation Index to track how much of its AI R&D — the work of researching, building, and improving its own models — is performed by Claude rather than by human researchers. The company states plainly: "Claude 'leads' 26% of Anthropic's AI R&D work," measured as of August 2026, versus under 1% in February 2026.
The index was constructed bottom-up: Anthropic says it built "this list of tasks in a bottom-up manner from work records including Slack and various sources of internal documentation," then rated each task on an automation-level scale (AL0 through AL5) and weighted the ratings by how much person-time each task actually consumes. "Leads" sits high on that scale — Anthropic's definition is that Claude can complete most of a given task end-to-end from a high-level prompt, while a human still supervises the outcome, short of fully autonomous operation.
A broader collaboration measure is even higher: Anthropic reports "the share of work at or above 'AI collaborates' is above 90%," meaning Claude is involved in large chunks of nearly all R&D work under close human direction, even where it isn't yet leading. Anthropic also disclosed operational detail behind the measurement effort: roughly 30,000 agents run simultaneously on its main internal platform, and only about 1 in 47,000 agent decisions — 0.002% — were blocked by oversight controls.
Anthropic is explicit about the index's limits. It says "the automation ratings depend on the judge model" used to score tasks, that "there remains real room for disagreement on borderline cases," and that human reviewers agreed with the model's ratings only 59% of the time (versus 35% human-to-human agreement) — meaning the underlying classification is noisy even by Anthropic's own account. It also notes the frozen baseline "does not, on its own, tell us whether new kinds of work are appearing."
Context
This release follows a pattern of AI labs, especially Anthropic, publishing self-referential metrics about how much of their own work is now AI-driven — often cited alongside separate claims that a large majority of code merged into Anthropic's codebase is now Claude-authored. It lands alongside renewed public and press attention this week to Anthropic's broader claim that Claude is materially accelerating the company's own model-development cycle, a storyline that has drawn coverage from wire services and business press over the past few days.
Why it matters
The 26%-to-under-1% jump in six months is one of the more concrete, quantified data points any frontier lab has published about AI systems accelerating their own development — a dynamic often discussed abstractly as "recursive self-improvement" but rarely measured this specifically. Anthropic's own caveats matter as much as the headline number: a judge-model-dependent, self-reported index with only 59% human-rater agreement is a rough instrument, and the company's own framing stops short of claiming full autonomy anywhere in the pipeline. Still, if the underlying trend continues at anything close to this pace, it has direct implications for how quickly frontier labs can iterate on model capability and safety work simultaneously — including the 6-12% of R&D compute Anthropic says it now allocates specifically to safety research.
Corroborating sources
- Anthropic
https://www.anthropic.com/institute/measuring-pace-of-ai-development
“Claude “leads” 26% of Anthropic’s AI R&D work.”