
Technology and Science · updated 1h ago · 3 min read
Anthropic CEO Dario Amodei Proposes Embedded Third-Party Safety Evaluators With Employee-Like Access
Dario Amodei proposed embedding independent safety evaluators inside frontier AI companies. Evaluators would audit models, report safety incidents, and assess alignment to external standards.
How independent embedded evaluators really are.
3 of 4 outlets skipped it: allegations that Anthropic uses AI doom narratives to fund its evaluator.
14outlets compared
Same story, two versions
tap a side to read it in full
CNBC
“Neither Anthropic's proposal nor OpenAI's existing third-party evaluation framework gives outside evaluators independent authority to halt the development or deployment of a model.”Read the original ↗
TechCrunch
“Third-party evaluators who spoke to TechCrunch broadly welcomed the proposal, but said details need to be ironed out”Read the original ↗
CNBC emphasises limited authority and enforcement gaps. TechCrunch stresses potential independence but flags details must be ironed out.
Embedded evaluators proposed
Anthropic CEO Dario Amodei proposed embedding third-party safety evaluators inside frontier AI companies on an ongoing basis, after former Anthropic researcher Jacob Coxon resigned warning that frontier labs were racing toward systems they might not control.
“employee-like access to verify safety practices and report incidents”
Amodei said the evaluators should have “employee-like access” with desks, access badges, and company laptops, and he committed Anthropic to providing third-party evaluators access comparable to internal risk teams.

Julie Andersen Hill, dean of the University of Wyoming College of Law and an expert on banking regulation, said the comparison to banking oversight is not accurate because without the power to flip the “kill switch” on the whole operation, the evaluators would not have comparable enforcement power.
Hill said bank examiners have offices inside institutions, access to internal systems and employees, and can direct a bank to stop a practice, restrict growth, force management changes, and in extreme cases close it.
Albert Ziegler, head of AI at cybersecurity company XBOW, said his team has received early access to unreleased models from Anthropic and OpenAI and that “we don't have any veto power,” with the ultimate decision remaining with the company.
Support, but internal rifts
OpenAI CEO Sam Altman reposted Amodei’s X post saying “Committing to having independent evaluators with employee-like access is a great idea, and we will do the same,” while SpaceXAI CEO Elon Musk reposted with “Dario is right.”
Despite the public alignment, El Cronista said the initiative is causing internal tensions at OpenAI and Anthropic, with some staff fearing external evaluators could compromise flagship technology security.
El Cronista reported that OpenAI noted it had already taken concrete steps including “a pause in training certain frontier technologies to moderate development,” while Anthropic has not halted its research.
Miles Brundage, CEO of the AI Verification and Evaluation Research Institute, said “binding mandate” would probably be required across the industry because “most external organizations have a lot, a lot less access than the full-time employee with the lowest level of permissions.”
In parallel, Protos reported a viral allegation that Anthropic’s safety evaluator is conflicted, citing a claim that Substack author Kevin Bass said “AI doom” media narratives were used to pump Dustin Moskovitz’s equity to fund foundations supporting METR.
Regulation stakes and oversight
The debate over embedded evaluators is tied to Washington’s frontier regulation push, with The Washington Post describing a coalition of AI researchers and 19 advocacy groups writing Senate leaders drawing red lines on frontier regulation being developed by Senate Majority Leader John Thune and Sen. Amy Klobuchar.
“a coalition of AI researchers and 19 advocacy groups has written Senate leaders”
In the same policy context, CNBC said Amodei wrote that evaluators are needed to provide “a neutral third party who can actually see the details,” but Hill said the plan does not give evaluators comparable enforcement power or legal authority to prevent a model from being trained or released.C
El Cronista said the internal question is how to balance safety and data protection while granting external experts access to “credentials, laptops, office access, and access to internal tools,” which some employees fear could expose confidential information and enable extraction of hard-to-replace knowledge.
CNBC also reported that Amodei explicitly cited embedded bank supervisors as a precedent, and Hill said continuous supervision is expensive for regulated companies, potentially strengthening large incumbents while making it harder for smaller competitors to enter.C
Meanwhile, El Cronista described a July investigation into an OpenAI agents attack on the Hugging Face developers’ site, saying METR and Redwood Research reviewed “tens of thousands of logs from over 1,200 agents,” and that OpenAI said the collaboration set an important precedent for future investigations.