Experts: Anthropic and OpenAI Need Independent Safety Evaluators

Over 100 AI experts are demanding better resources and protections for AI safety testers. They emphasize the need for independent oversight of powerful AI models. A consortium of academics and evaluators, including Geoffrey Hinton, published a letter urging AI providers to ensure third-party evaluators have objectivity, transparency, and independence. This initiative follows recent commitments by AI companies to embrace more rigorous third-party testing, with proposals suggesting “employee-like access” for evaluators to audit cutting-edge models and development processes. The call for standardization aims to build confidence in AI risk assessment.

Here’s the article rewritten in a CNBC style, with added depth and a professional tone, adhering to your specific requirements:

Over 100 artificial intelligence experts and evaluators are raising a unified alarm, warning that the burgeoning field of AI safety testing is facing a critical resource and protection deficit. This comes as concerns escalate from within the industry about the potential dangers posed by frontier AI models, bringing the need for independent oversight into sharp focus.

“Our primary objective is to establish a common understanding of fundamental principles and ensure that independent oversight can serve as a robust mechanism for managing AI risk on a broad scale,” stated Conrad Stosz, chair of the AI Evaluator Forum consortium, which spearheaded the coordinated letter. The consortium, comprising leading academics and independent evaluators, published a comprehensive letter outlining their concerns and demands, shared exclusively with CNBC. Signatories include luminaries such as Geoffrey Hinton, alongside members from esteemed institutions like Johns Hopkins University, Stanford University, and the non-profit evaluation body METR. The core of their message is a call to action for foundation model providers, urging them to guarantee that third-party AI evaluators are afforded the necessary scientific objectivity, transparency, independence, and stringent protections to conduct their work effectively and credibly.

Stosz emphasized that this initiative is a direct response to the recent commitments made by foundation model companies to embrace more rigorous third-party AI safety testing. The critical role of AI evaluators has been thrust into the spotlight following a recent suggestion by Anthropic CEO Dario Amodei to provide select evaluators with “employee-like access” to inspect and audit cutting-edge foundation models and their underlying development processes. While some industry leaders are advocating for government regulation to curb unchecked AI development, the debate remains contentious.

The consortium, Stosz clarified, is not prescribing a singular approach to AI safety development. Instead, their focus is on establishing “basic principles” and “greater standardization” for evaluators and other independent researchers operating outside the major AI labs. Amodei’s proposed access, according to Stosz, represents a significant expansion beyond current evaluator privileges. This could potentially involve foundation model developers granting third-party evaluators direct access to company computing infrastructure, facilitating candid discussions with employees, and allowing them to “examine sensitive internal data and unreleased systems.”

“This level of access would significantly enhance our confidence and certainty regarding the actual risks, particularly concerning systems that are utilized internally and not yet released,” Stosz explained, citing the example of an unreleased OpenAI model implicated in a recent security incident.

Prominent figures in the AI landscape, including OpenAI CEO Sam Altman and Microsoft CEO Satya Nadella, have publicly endorsed Amodei’s proposal. However, the practicalities of such an ambitious undertaking, including the selection criteria for AI evaluators and the depth of access granted to closely guarded technologies, remain largely unaddressed. The signatories are advocating for evaluator work to be conducted independently from the companies they assess, with enhanced transparency surrounding the technologies and a clear mandate to be “shielded from retaliation from the companies they embed with.”

### Consolidated Power Dynamics in AI Oversight

Vinh Nguyen, a senior fellow for AI at the Council on Foreign Relations and former chief AI officer for the National Security Agency, underscored the indispensable role of independent evaluators in uncovering critical information that can mitigate potential security failures and avert economic disruption. “When a concentrated group of powerful labs controls capabilities that could compromise cybersecurity, critical infrastructure, and the very systems underpinning our national security and economy, the government and the public cannot afford to rely solely on those labs’ own assurances of safety and security,” Nguyen stated.

Stosz reiterated that third-party evaluators are not intended to supplant internal evaluation efforts or the work of developers in mitigating identified issues. While acknowledging the possibility that foundation model companies might disregard the public letter and its call to action, he stressed that their credibility is on the line. “There are very few organizations possessing the requisite technical expertise, scale, and capability” to perform the comprehensive evaluations required, he noted.

The letter outlines “Minimum Conditions for Embedding Evaluators,” asserting that “all frontier AI companies should embed evaluators to independently assess AI risks, including evaluating the systems themselves and any significant incidents of real-world harm, as well as the companies’ training, deployment, oversight, operational, and safeguard practices.”

To ensure credibility, the letter specifies that embedded third-party evaluations must adhere to strict standards of scientific objectivity, transparency, independence, and robust protection against undue influence from the companies being evaluated. Key requirements include:

* **Meaningful Independence:** Frontier AI companies should engage evaluators who are demonstrably independent, maintain full editorial control, and proactively disclose and mitigate potential conflicts of interest. This necessitates that embedded evaluation organizations are not owned or governed by frontier AI companies, have no significant commercial dealings with them, and do not accept payment or rewards contingent on specific findings.
* **Diverse Viewpoints and Expertise:** Companies should incorporate a range of perspectives and expertise by embedding multiple evaluation organizations across critical risk areas. This includes allowing and encouraging evaluators to share their differing conclusions amongst themselves and with company employees.
* **Transparency and Disclosure:** Embedded evaluators must operate with a high degree of transparency regarding their methodologies, findings, access protocols, and the broader terms of engagement. Frontier AI companies are expected to facilitate this transparency, limiting the scope of non-disclosure agreements and permitting prompt, unfiltered communication with company boards and oversight bodies. Public release of findings and evidence should be allowed, with limited, time-bound redactions solely for protecting critical interests such as intellectual property, sensitive customer information, individual privacy, security, and public safety.
* **Protection Against Retaliation:** Embedded evaluators must be safeguarded from retaliatory actions by the companies they assess, whether for adopting reasonable evaluation methods, uncovering sensitive information, or reaching conclusions unfavorable to the company. This includes protection against retaliatory litigation and funding mechanisms that ensure continued financial support even in such circumstances.
* **Equitable Access:** Frontier AI companies must grant embedded evaluators access equivalent to that of their own highly privileged employees for evaluation purposes, with necessary exceptions for protecting third-party sensitive data. This means providing access to the same relevant systems, data, tools, and physical spaces as senior internal employees responsible for comparable risk assessments, alongside candid, direct communication with relevant staff. The letter stresses that this list is not exhaustive and advocates for increasing standardization, codification, and enforcement of such conditions to ensure credible evaluations. It also highlights the necessity for embedded evaluations to complement, rather than replace, broader external oversight efforts by frontier AI companies, including greater public transparency and expanded access for independent researchers.

Original article, Author: Tobias. If you wish to reprint this article, please indicate the source:https://aicnbc.com/25877.html

Like (0)
Previous 3 hours ago
Next 1 hour ago

Related News