Published

Anthropic CEO Dario Amodei Proposes Embedded Third-Party Safety Evaluators With Employee-Like Access
Image: The Washington Post

Technology and Science · updated 1h ago · 3 min read

Anthropic CEO Dario Amodei Proposes Embedded Third-Party Safety Evaluators With Employee-Like Access

Happened

Dario Amodei proposed embedding independent safety evaluators inside frontier AI companies. Evaluators would audit models, report safety incidents, and assess alignment to external standards.

Split on

How independent embedded evaluators really are.

Left out

3 of 4 outlets skipped it: allegations that Anthropic uses AI doom narratives to fund its evaluator.

14outlets compared

Brief IABusiness InsiderCNBCDiarioBitcoinEl CronistaExpansiónhttpsKultureGeek

Same story, two versions

tap a side to read it in full

CNBCCNBC

Neither Anthropic's proposal nor OpenAI's existing third-party evaluation framework gives outside evaluators independent authority to halt the development or deployment of a model.
Read the original

TechCrunchTechCrunch

Third-party evaluators who spoke to TechCrunch broadly welcomed the proposal, but said details need to be ironed out
Read the original
VS

CNBC emphasises limited authority and enforcement gaps. TechCrunch stresses potential independence but flags details must be ironed out.

Embedded evaluators proposed

Anthropic CEO Dario Amodei proposed embedding third-party safety evaluators inside frontier AI companies on an ongoing basis, after former Anthropic researcher Jacob Coxon resigned warning that frontier labs were racing toward systems they might not control.

employee-like access to verify safety practices and report incidents

CNBCCNBC

Amodei said the evaluators should have “employee-like access” with desks, access badges, and company laptops, and he committed Anthropic to providing third-party evaluators access comparable to internal risk teams.

Image from Brief IA
Brief IABrief IA

Julie Andersen Hill, dean of the University of Wyoming College of Law and an expert on banking regulation, said the comparison to banking oversight is not accurate because without the power to flip the “kill switch” on the whole operation, the evaluators would not have comparable enforcement power.

Hill said bank examiners have offices inside institutions, access to internal systems and employees, and can direct a bank to stop a practice, restrict growth, force management changes, and in extreme cases close it.

Albert Ziegler, head of AI at cybersecurity company XBOW, said his team has received early access to unreleased models from Anthropic and OpenAI and that “we don't have any veto power,” with the ultimate decision remaining with the company.

SourcesCNBCCNBC

Support, but internal rifts

OpenAI CEO Sam Altman reposted Amodei’s X post saying “Committing to having independent evaluators with employee-like access is a great idea, and we will do the same,” while SpaceXAI CEO Elon Musk reposted with “Dario is right.”

Despite the public alignment, El Cronista said the initiative is causing internal tensions at OpenAI and Anthropic, with some staff fearing external evaluators could compromise flagship technology security.

Image from Business Insider
Business InsiderBusiness Insider

El Cronista reported that OpenAI noted it had already taken concrete steps including “a pause in training certain frontier technologies to moderate development,” while Anthropic has not halted its research.

Miles Brundage, CEO of the AI Verification and Evaluation Research Institute, said “binding mandate” would probably be required across the industry because “most external organizations have a lot, a lot less access than the full-time employee with the lowest level of permissions.”

In parallel, Protos reported a viral allegation that Anthropic’s safety evaluator is conflicted, citing a claim that Substack author Kevin Bass said “AI doom” media narratives were used to pump Dustin Moskovitz’s equity to fund foundations supporting METR.

Regulation stakes and oversight

The debate over embedded evaluators is tied to Washington’s frontier regulation push, with The Washington Post describing a coalition of AI researchers and 19 advocacy groups writing Senate leaders drawing red lines on frontier regulation being developed by Senate Majority Leader John Thune and Sen. Amy Klobuchar.

a coalition of AI researchers and 19 advocacy groups has written Senate leaders

The Washington PostThe Washington Post

In the same policy context, CNBC said Amodei wrote that evaluators are needed to provide “a neutral third party who can actually see the details,” but Hill said the plan does not give evaluators comparable enforcement power or legal authority to prevent a model from being trained or released.C

El Cronista said the internal question is how to balance safety and data protection while granting external experts access to “credentials, laptops, office access, and access to internal tools,” which some employees fear could expose confidential information and enable extraction of hard-to-replace knowledge.

CNBC also reported that Amodei explicitly cited embedded bank supervisors as a precedent, and Hill said continuous supervision is expensive for regulated companies, potentially strengthening large incumbents while making it harder for smaller competitors to enter.C

Meanwhile, El Cronista described a July investigation into an OpenAI agents attack on the Hugging Face developers’ site, saying METR and Redwood Research reviewed “tens of thousands of logs from over 1,200 agents,” and that OpenAI said the collaboration set an important precedent for future investigations.