Anthropic says its Claude AI system is now leading 26% of the company’s model research and development work, while humans continue to set objectives and supervise the results. The disclosure, reported Sept. 18 by NBC4 Washington, is an unusually specific account of how a frontier AI developer is using agents to help build the systems that may follow.
The figures will inevitably feed claims that AI is beginning to improve itself. Anthropic’s own account draws a sharper boundary. Claude can carry out substantial, well-defined work from high-level instructions, the company says, but it is not autonomously designing and training its successor. In Anthropic’s definition, that would be recursive self-improvement: a system independently directing the development of a more capable replacement. The company says it has not reached that threshold.

What the 26% figure does — and does not — measure
Anthropic describes two related but very different measures of adoption. Claude reportedly leads 26% of its model R&D work under human supervision. Separately, the company says roughly 90% of its research and development is conducted in collaboration with Claude. The first figure concerns work where the model is the primary executor; the second is a much broader measure that includes people using Claude as a collaborator.
Those numbers should not be read as evidence that a quarter of Anthropic’s research operation is unsupervised. The company says people remain responsible for assigning goals, evaluating output and making consequential decisions. In practical terms, the agents can take a task, write code, run an experiment, inspect results and iterate within a bounded assignment. They remain much less reliable at deciding which research question is worth pursuing, recognizing when a plausible result is misleading, or balancing technical tradeoffs that are poorly specified at the outset.

That limitation is more than a semantic caveat. A model that efficiently completes a human-framed task can speed up an organization without being able to chart its own development path. Anthropic says the gaps in goal selection and judgment are still substantial in both engineering and research. Its account of AI-assisted development argues that those functions would have to become far more capable before autonomous recursive self-improvement could be said to exist.
Rapid agent deployment, with incomplete productivity measures
The reported scale is nevertheless notable. Anthropic said about 30,000 agents were conducting research and engineering work as of August. NBC4 Washington reported that Claude-led work had risen from none in February to about one-quarter of R&D by August. That is a six-month ramp from a limited role to an embedded part of the company’s technical workflow.
Software development is where the company offers its clearest operational metrics. Anthropic says that, as of May, Claude had authored more than 80% of code merged into its codebase. It also says the typical engineer merged eight times as much code per day in the second quarter of 2026 as in 2024.
Neither number is a clean measure of eightfold engineering productivity. Anthropic explicitly cautions that lines of code can exaggerate gains: generated code may be verbose, later discarded or require substantial review, testing and maintenance. Code merged per day is closer to a throughput indicator than a measure of business value, reliability or long-term software quality. The more defensible conclusion is narrower: within Anthropic’s environment, AI assistance appears to have changed the volume and division of coding work dramatically.
The company dates the beginning of this shift to the research preview of Claude Code in February 2025. Its newer systems are being used not only for routine implementation but for broader research tasks, such as constructing experimental infrastructure and executing specified experiments. That progression helps explain why the company is framing the issue as an AI-development question rather than merely a coding-tool story.
Internal experiments show promise, not an independent benchmark
Anthropic also cites internal evaluations intended to show that agents are improving on open-ended work. It says Claude’s success rate on its most open-ended task category reached 76% in May 2026, a 50-percentage-point increase in six months. The company does not present that as a universal measure of scientific or engineering autonomy; it is an internal category and result, rather than a cross-lab benchmark.
One AI-safety research exercise produced a more concrete comparison. Anthropic says Claude-powered agents recovered 97% of a defined performance gap over 800 cumulative hours and roughly $18,000 in compute. Two human researchers recovered about 23% over roughly a week. The comparison is suggestive, but its inputs are not interchangeable: hundreds of cumulative agent-hours can be parallelized in a way two people over a week cannot. The outcome also came from a defined project, not a demonstration that the system can independently choose valuable safety research directions.
Anthropic acknowledges another constraint. It says lessons from the exercise did not transfer cleanly to production-scale models. That is a useful warning against treating a single research result as proof that agent-assisted discovery will scale predictably across model training, evaluation and deployment.
A call for comparable reporting
The disclosure is also a bid to establish a reporting norm. Anthropic wants frontier AI developers to publish comparable measures regularly, allowing observers to track how much development work is AI-assisted, how much is agent-led, what humans still control and where systems fail. At present, the figures are company-reported operational data, not independently audited measurements. They are valuable because few major labs have provided this level of detail, but they cannot by themselves establish an industry-wide rate of progress.
Anthropic’s caution is partly about governance. A system that materially accelerates the creation of more capable systems could compress the time available for safety testing, organizational oversight and public scrutiny. Yet the company’s description of its current workflow remains recognizably human-directed: people identify goals, agents execute increasingly large portions of the work, and people judge whether the output should be trusted. Whether that arrangement remains stable as agents take on less structured work is the question its proposed reporting regime is meant to make visible.
