Post

Conversation

THE IMPORTANCE OF EMBEDDED EXAMINERS This morning, Dario Amodei released his proposal for pacing the frontier. Chief among his regulatory reforms is the idea that frontier labs should have embedded examiners with access to the same capabilities as the company's own safety researchers and the freedom to publish whatever they might find. This, I believe, is the key innovation in frontier AI safety: continuous visibility inside the lab, rather than a test at the door on release day. It is also the idea ARI, the group I lead, put at the center of its own frontier plan that we released last month. That plan called for examiners who are on premises with access to most of the things that company employees themselves have. The analogy was to bank examiners. Dario reaches for the same analogy. Today, for many large banks, the Federal Reserve requires bank examiners who are on site full time. They can attend meetings, demand information, and serve as an early warning system for banks that might pose a risk to the financial health of the country and maybe the world. Bank regulators have worked through some of the issues that such an idea often raises, such as the fear that the examiners might be captured by the banks they oversee, or that their presence would somehow impede the banks' operations. Yesterday, , one of the most astute commentators on AI policy, had a related proposal that would have private organizations like METR, Apollo, or Transluce perform that function. That is a sensible approach, at least in the interim, it seems to me. Dario's version is like Anton's in this respect. The examiners at Anthropic will be third parties like METR, not the government. That is the right first step. Government examiners remain a possible solution if the private ecosystem proves incapable of scaling to the task. In any event, now that Dario has proposed it and is actually adopting this idea for Anthropic, it is clear that the time has come to make this a central part of frontier governance. Simply testing models that are going to be released to the consuming public is inadequate. The OpenAI–Hugging Face incident showed us that. There are models used internally by the companies that could pose serious safety risks to the broader public. Internal examiners who can flag these issues early for possible government questioning, if not intervention, are essential. We need broader, entity-based regulation if we are to ensure that AI goes well. It is my sincere hope that other companies will voluntarily adopt the approach Anthropic is taking, and that all of us work to bolster the funding and capability of the organizations doing evaluations. They are now carrying far more weight than their budgets suggest, and that is the next problem to fix.
Amitav Krishna
Post your reply