
An Amazon Web Services AI data center in New Carlisle, Indiana, on Oct. 2, 2025.
Noah Berger | Getty Images
The burgeoning field of artificial intelligence is poised to create a new, high-stakes role: the embedded AI evaluator. These individuals, granted access to the most sensitive systems and large language models of industry giants like Anthropic and OpenAI, will wield significant insight into the capabilities of frontier AI. However, a critical question looms: will this access translate into meaningful control?
Anthropic CEO Dario Amodei has championed a plan to integrate third-party safety evaluators directly within leading AI companies. This initiative aims to temper the rapid advancement of increasingly sophisticated AI models. The proposal emerged in the wake of concerns raised by former Anthropic researcher Jacob Coxon, who warned of a reckless race towards AI systems potentially beyond human control.
In a recent essay, Amodei outlined a vision where these embedded evaluators would receive access comparable to internal risk assessment teams. Crucially, they would be granted the right to publish their findings without corporate censorship, subject only to essential redactions. Amodei drew a parallel to the banking sector, citing the precedent of regulatory “supervisors” embedded within financial institutions to monitor operations.
However, legal experts contend that this analogy may fall short. Julie Andersen Hill, Dean of the University of Wyoming College of Law and an authority on banking regulation, argues that without the ultimate power to halt operations—akin to a “kill switch”—the comparison to robust banking oversight is inaccurate. “If you don’t give them that kind of power, I don’t know what they are doing,” Hill stated, emphasizing the stark difference in enforcement capabilities.
In the banking world, government examiners operate with a continuous presence, full access to internal systems and personnel, and the authority to direct operational changes, restrict growth, mandate management shifts, and, in severe cases, shutter an institution. In contrast, Amodei’s proposal for AI evaluators, while allowing for investigation and reporting, does not grant them the legal authority to prevent model training or deployment. This fundamental disparity in power, Hill noted, distinguishes the two regulatory models.

Amodei himself acknowledged the need for “a neutral third party who can actually see the details,” while also conceding that Anthropic retains control over the information it chooses to disclose. The current framework suggests that while evaluators will gain extraordinary access to cutting-edge AI, their formal authority remains limited. Neither Anthropic’s proposal nor OpenAI’s existing third-party evaluation protocols bestow independent power to halt model development or release.
Both Anthropic and OpenAI, which has also pledged a similar safety commitment though details are still forthcoming, declined to comment on their embedded evaluator plans.
Inside the AI Black Box: What Evaluators Are Discovering
Albert Ziegler, head of AI at cybersecurity firm XBOW, which conducts its own proactive evaluations of emerging models, reports that his team has received early access to pre-release versions from leading AI developers, including Anthropic and OpenAI. XBOW’s methodology involves independent testing within their own secure environment, with findings typically shared with the model providers. Ziegler clarified that the day-to-day reality of these evaluations is less dramatic than the existential narratives often portrayed.
While Amodei’s timeline predicts potential internet disruption by misaligned AI agents within six to 12 months, Ziegler’s team has primarily encountered issues such as models producing nonsensical outputs under unusual formatting requests or the need for more frequent intervention by external safety checkers. “But the kind of insidious, catastrophic consequences produced by subterfuge combined with unprecedented abilities that people are afraid of — that’s not something we’ve seen ourselves,” he stated.
Ziegler explained that while black-box testing can identify a model’s capacity to perform specific dangerous tasks, assessing the broader systemic risk requires deeper access. This includes understanding the model’s underlying instructions, its granted tools and permissions, its safety protocols, and logs of attempted actions. Even with such comprehensive access, he conceded, emergent risks might only surface under highly specific and unforeseen circumstances, underscoring the inherent challenges of predictive safety assessment.
“It’s true that we don’t have any veto power,” Ziegler admitted. However, he emphasized that evaluators can identify and document risks overlooked by developers, thereby “compelling an informed decision before release.” Ultimately, the decision-making authority rests with the AI company.
Even the established banking regulatory model, Hill noted, is not without its flaws. Supervisors have, at times, failed to avert major financial crises, and are often criticized for becoming too close to the institutions they are mandated to oversee. Furthermore, the continuous oversight required is a significant financial burden, potentially favoring larger, established firms and hindering smaller competitors.
Navigating Conflicts of Interest in AI Safety
The current proposals have also faced skepticism from critics who suggest that AI companies may be leveraging safety concerns to advocate for a regulatory environment that could shield them from competition and liability. Anthropic, in particular, has been scrutinized for potential conflicts of interest regarding a proposed evaluator. Critics argue that if an AI company selects its own evaluators, dictates their scope of access, and retains the right to disregard their findings, the arrangement resembles an internal compliance department more than independent oversight.
“If Anthropic wants an internal compliance department, there’s nothing currently stopping them from having one,” Hill remarked. “They don’t need the government to do that. It seems like what the incumbent AI people want is just somebody to watch them, but then communicate to the public, ‘Look, we’ve looked behind the curtain and there’s nothing bad going on there.’ That’s somewhat unusual in a regulatory sense.”
Amodei mentioned the nonprofit organization Model Evaluation and Threat Research (METR) as a potential candidate for embedded evaluation. Anthropic has previously collaborated with METR, most recently requesting a review of cybersecurity incidents involving its Claude model. The recent move of Joe Benton, a former Anthropic researcher, to METR highlights both the growing expertise within the safety evaluation field and the interconnectedness of its key players.
This dynamic underscores a fundamental tension: the entity developing the AI selects the evaluators, defines their purview, and retains ultimate control over the outcome. Deborah Raji, a researcher specializing in algorithmic auditing at UC Berkeley, argues that mere access does not equate to independence. In established auditing systems, she explained, auditors must adhere to strict standards of competence, independence, and conduct to ensure the integrity of the process.
“You are effectively not qualified to be an actual auditor if you can’t meet the standards of independence conduct,” Raji asserted. “It discredits the whole process.” She advocates for an independent authority to vet evaluators, define their scope of examination, and dictate reporting protocols. “You can’t just wake up one day and decide that you’re qualified to be a bank examiner, and the company being audited can’t randomly assign you to be a qualified bank examiner either,” she added. “Otherwise, we’d have the equivalent of the companies asking a random friend to check their homework.”
METR states that it does not accept direct financial support from AI companies or their executives. However, its own “Frontier Risk Report” acknowledges that some of its employees maintain close social ties with individuals at AI companies, and that it shares a research center with some lab employees. While these relationships are not presented as compromising its work, they do illuminate the compact and interconnected nature of the nascent AI evaluation landscape.
‘Allows for all kinds of strange things.’
Raji expressed concern over the implications of METR’s existing relationships, citing suspicions of financial and ideological entanglements, personal conflicts of interest—such as familial ties between evaluators and AI company board members—and a perceived over-reliance on Anthropic and OpenAI. She characterizes the current structure as “really unusual, and allows for all kinds of strange things.”
Despite these concerns, Raji acknowledged METR’s transparency regarding its contractual arrangements with OpenAI for investigations like the Hugging Face incident. However, she pointed out that the existing framework still allowed OpenAI to define the scope of the audit, control access, and influence the publication of findings. Ziegler, conversely, believes that his firm’s early-access engagements have not compromised XBOW’s independence, as developers are motivated to identify vulnerabilities rather than seek validation of predetermined outcomes.
Christina Ho, chief assurance officer at accounting firm Oath and former board member of the Public Company Accounting Oversight Board, notes that this tension between independence and client relationships is not unique to AI. Auditors are compensated by the entities they scrutinize, creating an inherent conflict. However, post-financial crisis legislation, such as Sarbanes-Oxley, has introduced liability for auditors in cases of failure. AI introduces a further layer of complexity: expertise. Traditional audits focus on adherence to established processes, whereas AI requires evaluating the actual system and its output, a domain with a currently limited pool of qualified professionals.
Hill argues that even if potential conflicts can be managed, the proposed AI oversight frameworks lack a detailed legal structure, which is crucial for effective regulation. “It doesn’t really work to just turn supervisors loose without any standards to hold them to,” she stated. Frontier AI currently lacks a comprehensive body of operating rules defining prohibited actions, evaluator discretion, or consequences for serious findings. While access and publication rights can offer external scrutiny, they cannot provide regulatory credibility if the company retains ultimate control. “You can’t have it both ways,” Hill cautioned. “You can’t have all of the control and then expect the credibility as if you’ve given up control.”
The defining factor, according to Raji, is whether an adverse finding leads to tangible consequences. “The goal of an audit is to get the audit target to face some kind of consequential judgment,” she explained. “If you do an audit and nothing happens, that’s audit washing.”
In Hill’s view, if the existential risks posed by AI are as profound as industry leaders suggest, the ultimate test for any oversight mechanism is clear: “If you really believe that AI has the power to destroy society, then you have to have an independent supervisor that has the ability to pull the plug on it,” she concluded.

Raji emphasized that METR’s transparency regarding its contractual arrangements with OpenAI for investigations like the Hugging Face incident is commendable. However, she pointed out that the existing framework still allows OpenAI to define the scope of the audit, control access, and influence the publication of findings. Ziegler, conversely, believes that his firm’s early-access engagements have not compromised XBOW’s independence, as developers are motivated to identify vulnerabilities rather than seek validation of predetermined outcomes.
Christina Ho, chief assurance officer at accounting firm Oath and a former board member of the Public Company Accounting Oversight Board, notes that this tension between independence and client relationships is not unique to AI. Auditors are compensated by the entities they scrutinize, creating an inherent conflict. However, post-financial crisis legislation, such as Sarbanes-Oxley, has introduced liability for auditors in cases of failure. AI introduces a further layer of complexity: expertise. Traditional audits focus on adherence to established processes, whereas AI requires evaluating the actual system and its output, a domain with a currently limited pool of qualified professionals.
Hill argues that even if potential conflicts can be managed, the proposed AI oversight frameworks lack a detailed legal structure, which is crucial for effective regulation. “It doesn’t really work to just turn supervisors loose without any standards to hold them to,” she stated. Frontier AI currently lacks a comprehensive body of operating rules defining prohibited actions, evaluator discretion, or consequences for serious findings. While access and publication rights can offer external scrutiny, they cannot provide regulatory credibility if the company retains ultimate control. “You can’t have it both ways,” Hill cautioned. “You can’t have all of the control and then expect the credibility as if you’ve given up control.”
The defining factor, according to Raji, is whether an adverse finding leads to tangible consequences. “The goal of an audit is to get the audit target to face some kind of consequential judgment,” she explained. “If you do an audit and nothing happens, that’s audit washing.”
In Hill’s view, if the existential risks posed by AI are as profound as industry leaders suggest, the ultimate test for any oversight mechanism is clear: “If you really believe that AI has the power to destroy society, then you have to have an independent supervisor that has the ability to pull the plug on it,” she concluded.

Original article, Author: Tobias. If you wish to reprint this article, please indicate the source:https://aicnbc.com/25802.html