Anthropic and OpenAI Propose Embedding Safety Evaluators Inside AI Labs
The two AI companies say they want independent researchers to have unprecedented internal access, but experts caution that true oversight will require transparency, independence and eventual regulation.
Anthropic and OpenAI have proposed embedding independent safety evaluators directly within their organizations, according to a report by TechCrunch. The plan would give outside researchers access to the companies' AI systems and development processes in ways not previously offered.
TechCrunch reported that researchers have welcomed the level of access being proposed, describing it as unprecedented for the AI industry. However, the same researchers cautioned that access alone does not guarantee meaningful oversight.
According to the report, experts warned that for embedded evaluators to provide genuine safety checks, several conditions would need to be met, including transparency about the evaluators' findings, structural independence from the companies they are assessing, and ultimately the introduction of external regulation.
The report raises questions about whether evaluators embedded within a company can remain truly independent, given their proximity to and reliance on the organizations they are tasked with scrutinizing. TechCrunch's reporting frames the initiative as a test of whether AI labs can credibly self-police as their systems grow more powerful.
Neither Anthropic nor OpenAI's specific commitments regarding the scope, funding or authority of these evaluators were detailed in the available reporting. TechCrunch did not specify a timeline for when such embedded evaluation programs might be implemented.
Sources
EGazette summarizes reporting from multiple sources; follow the links for the originals.
Related articles

Debate Over AI Regulation Continues Following Amodei's Proposal
Anthropic CEO's plan for slowing AI development sparked industry discussion, though disagreements over regulatory approach persist, according to The Verge.

Anthropic CEO's 'Pace the Frontier' AI Safety Plan Draws Mixed Reactions
Dario Amodei's proposal for independent safety evaluators and lab coordination gains some industry backing while facing pushback from Nvidia's Jensen Huang

Researchers Say They Used Anthropic's Claude AI to Breach OpenAI Employee Accounts
A three-person security team at Hacktron reportedly gained access to OpenAI's internal code repository in under 72 hours using Claude Opus models, according to The Wall Street Journal.

Australia's PM Asks Apple's Tim Cook to Back Online Safety, AI Rules
Anthony Albanese urges big tech support for Canberra's internet safety and artificial intelligence regulations, according to Al Jazeera.

AI industry workers express skepticism about existential risk warnings
Multiple employees at leading AI companies doubt predictions that the technology could pose catastrophic threats to humanity.

Trump Announces Plans for 'AI Force' and Artificial Intelligence Tsar
President says his administration will not "hinder or stifle" the growth of artificial intelligence, amid ongoing warnings about the technology's risks
Comments
Loading comments…