Anthropic, OpenAI and Google have reportedly been discussing a shared industry body to develop standards, testing and audits for advanced artificial-intelligence systems before they are released. The talks, reported Sept. 14 by Yahoo News on the basis of accounts from The Information and CNN, could give the biggest frontier-model developers a more formal role in deciding what evaluations a model must pass before deployment.
There is no announced organization, agreed rulebook or public commitment from the three companies. That is more than a semantic distinction: a discussion group can explore common language around testing, while an institution would need decisions on membership, financing, technical authority and whether anyone can enforce its findings. The reported proposal arrives as arguments over the speed and safety of AI development have become unusually public, including from a former researcher who worked at both Anthropic and OpenAI.
Reported talks predate the latest safety dispute
Yahoo reported that representatives of the companies had met as a working group since at least July and had met again in the preceding week. Its account says the talks began before Anthropic chief executive Dario Amodei published his public call to pace frontier AI development, and before Jacob Coxon, a safety researcher formerly employed by Anthropic and OpenAI, left Anthropic.
The sequence is useful because it separates two related but different developments. According to the report, the companies were already exploring a common standards arrangement before Amodei’s Sept. 6 essay and Coxon’s resignation during the following week. The safety debate may have added urgency and visibility, but the available reporting does not support treating the proposed body as a direct response to either event.
ABC News’s interview with Coxon, published Sept. 13, provides the clearest on-record account in this set of reports. Coxon said he resigned over AI-safety concerns, welcomed Amodei’s appeal for developers to slow down and argued that neither former employer was acting responsibly. Those are Coxon’s assessments, not findings from an external audit, but they show why voluntary testing arrangements are under close scrutiny.
Amodei’s essay, as described by ABC, called for a slower pace of development at the frontier. It did not establish that Anthropic had adopted a binding slowdown, nor did it describe a completed joint program with rivals. Likewise, Coxon said conversations about severe AI risks had become more frequent during his time inside the companies, but that observation does not reveal what specific safeguards each company uses or whether they work.
What a pre-release testing body could do
The reported concept is more concrete than a general pledge to develop AI responsibly. Yahoo said Google DeepMind founder Demis Hassabis had proposed a public-private partnership that would evaluate advanced models before release, with industry funding and independent technical experts. Such a body could, in principle, establish common test procedures and publish judgments or recommendations on whether a system met a specified threshold.
Testing and auditing are often grouped together, but they perform different functions. Testing asks whether a model displays a particular capability or failure under defined conditions. An audit examines whether the testing process, records, controls and conclusions are credible. A body intended to influence release decisions would therefore need to settle basic operational questions: what capabilities trigger review, who designs the tests, what evidence is shared, how expert independence is protected, and what happens if a developer proceeds despite an unfavorable result.
None of those details has been agreed publicly. The reported Hassabis plan’s use of independent experts points toward one answer on technical credibility, but independence depends on more than job titles. Funding, appointment rules, access to models and the treatment of confidential results would all shape who ultimately holds power. A pre-release evaluator also differs from a regulator unless a government grants it legal authority or companies contractually bind themselves to its decisions.
Yahoo reported that OpenAI chief executive Sam Altman favored an industry testing and auditing body and believed leading labs might need to create it without government backing. The same report said the three companies declined to comment to CNN. That leaves the account of Altman’s internal position and the wider discussions attributed to the reporting, rather than confirmed through a joint announcement.
Shared safeguards could also concentrate influence
A common testing institution could make it easier for companies to compare results and avoid each developer defining safety for itself. It could also concentrate rule-making among companies with the money, compute access and staff needed to operate at the frontier. Yahoo reported that Cohere chief executive Aidan Gomez criticized the prospect of the largest firms setting the rules together, arguing it could put smaller rivals at a disadvantage.
The objection is not simply about whether AI needs safeguards; it is about who writes them and how they apply. A demanding evaluation regime may be defensible if its tests are transparent, proportionate and governed independently. But if incumbent developers set costly requirements without meaningful outside representation, the same system can become a barrier to entry. The current reports do not establish how the participants would handle representation, conflicts of interest or enforcement.
Political support is unresolved as well. Yahoo reported resistance among some companies and said House Speaker Mike Johnson had found no consensus on AI guardrails; it also described a stalled effort around a draft executive order. The companies’ reported talks were said to be continuing without administration support. For now, the tangible news is a working-level conversation among three influential developers, not a new regulator or a settled standard. Its credibility will depend on the technical tests it proposes and on whether the labs are willing to accept oversight beyond their own control.
