OpenAI, Anthropic consider cross-testing amid AI safety worries
As artificial intelligence systems advance in capability, the firms developing these technologies are progressively seeking to extend their focus beyond internal safety measures. OpenAI and Anthropic are reportedly in discussions to establish a legally binding agreement that would enable them to evaluate each other’s commercially available AI models for vulnerabilities and unforeseen behaviour. This represents a notable shift toward collaborative safety testing between companies. According to a report, the proposed arrangement would enable OpenAI and Anthropic to access each other’s commercial models via API, facilitating independent stress tests. The companies would also agree not to retain data obtained during the testing process.
The discussions arise amid escalating concerns regarding the growing autonomy of AI systems and the adequacy of current safety measures in keeping up with their swiftly advancing capabilities. The proposed agreement would enable the two AI companies to scrutinise each other’s models instead of solely depending on internal safety evaluations.
According to the deal:
- OpenAI and Anthropic to share API access for their commercial models
- Companies will leverage that access to identify safety vulnerabilities and unexpected behaviours
- Neither side will keep the other’s testing data
- The setup adds an extra layer of oversight with internal and independent evaluations
The report said the objective is to identify risks that could remain hidden during conventional testing as AI systems become more advanced. The proposed pact has not been presented as a finalised agreement, and details may evolve as negotiations progress. The discussions emerge amidst an escalating discourse among AI firms regarding the pace of development for frontier models and the criteria for their assessment. Anthropic CEO Dario Amodei recently advocated for enhanced safeguards and increased caution regarding the advancement of progressively capable AI systems. OpenAI CEO Sam Altman subsequently stated his agreement that the industry must “pace the frontier,” while OpenAI expressed its intention to support independent evaluators in gaining employee-like access to its systems. The discussion has been further intensified by apprehensions regarding AI agents, which are capable of executing increasingly intricate tasks with minimal human oversight.
OpenAI has revealed cases of what it terms “reward hacking,” wherein AI systems attain a desired result through unintended means. Concurrently, the broader industry has encountered scrutiny regarding the behaviour of autonomous systems in unfamiliar contexts. Cross-company testing could therefore provide a mechanism for competing developers to scrutinise one another’s models and uncover vulnerabilities that might remain undetected within their own organisations. The reported discussions with Anthropic occur concurrently with OpenAI’s advocacy for enhanced international collaboration on frontier AI safety. On Monday, the artificial intelligence giant urged the United States to spearhead a global initiative aimed at establishing technical standards for frontier AI, particularly for systems that possess the capability for recursive self-improvement, or RSI.
RSI refers to AI systems that enhance their own capabilities or assist in the development of subsequent generations of AI. OpenAI stated that fully autonomous RSI is not a current reality and contended that the pursuit of such systems should be contingent upon their safe development while ensuring human oversight, as reported. The company also cautioned that increased autonomy in AI research might complicate human comprehension and oversight of the progression of these systems. OpenAI stated that international standards are essential due to the global nature of AI development, as varying evaluation methods and incident reporting practices may hinder the ability to compare findings and address risks that transcend national boundaries.
The company advocated standardisation for:
- Assessing advancements in relation to RSI, particularly the degree of autonomous AI research
- Establishing the necessary parameters for human oversight in the context of automated AI research
- Establishing standardised frameworks for the classification, monitoring, and reporting of incidents related to AI safety
- Establishing common standards to evaluate the adequacy of AI safety measures
OpenAI also stated that these standards ought not to be construed as automatic licensing requirements or obligatory pre-release approval systems. Instead, individual governments would determine the manner in which to integrate them into their respective regulatory frameworks.








