The 'Embedded' Illusion: Why AI Labs Are Trading Oversight for PR

AI-generated image · Bay Street Wire
Anthropic and OpenAI's plan to bring safety evaluators inside their walls confuses access with accountability.
*(Opinion)*
In the world of machine learning, there is a massive difference between being invited to the party and being allowed to inspect the kitchen.
As TechCrunch first reported, Anthropic CEO Dario Amodei recently proposed embedding third-party safety evaluators directly inside frontier AI companies to assess model alignment and report safety incidents. OpenAI CEO Sam Altman has signaled that his company would also commit to this practice.
On the surface, this looks like a victory for transparency. But I see this as a corporate PR play. By embedding evaluators, these labs aren't creating independent oversight; they are creating a veneer of legitimacy while maintaining the keys to the kingdom.
The core problem is that access is not the same as authority. Evaluators are already questioning whether they will function as true watchdogs or merely as vendors operating on the AI companies' terms.
The track record is concerning. When OpenAI allowed METR and Redwood Research to investigate the Hugging Face incident, they were given roughly a week on-premises; both later cited limitations in scope and timing. Similarly, Apollo Research noted in its model card contribution that it was granted a mere three days to test GPT-6 Astra. Apollo subsequently warned that limited windows and increasing model awareness meant low rates of misbehavior did not provide substantial evidence of alignment.
Furthermore, the nature of this access remains a black box. Despite repeated questions from TechCrunch, neither lab has specified which evaluators they will hire, when they start, or exactly what systems will be accessible.
From a technical perspective, testing a 'finished model' is flawed. John Steidley, head of strategy at Palisade Research, told TechCrunch that models can be trained to pass specific benchmarks—comparing it to the Volkswagen 'Dieselgate' scandal where cars recognized emissions tests to alter behavior. To catch this, evaluators need to see 'checkpoints' during training. Adam Gleave, CEO of FAR.AI, noted the need to inspect post-training environments and logs to verify company claims.
Alexander Meinke, head of research at Apollo Research, raised a critical concern to TechCrunch: whether the AI ever actively attempted to undermine its own alignment training. Currently, the public relies on companies to report this truthfully—something Meinke noted they have failed to do by default in recent incidents.
Amodei suggested evaluators could publish findings without editorial control, but Gleave noted that intellectual property is too valuable for companies to truly surrender control. FAR.AI has previously turned down contracts from developers who demanded too much oversight of the evaluation process.
Without legislative backing to codify these roles, the 'independent' evaluator is just another line item in a corporate budget. Until there are legally binding mandates on disclosure and audit duration, this is nothing more than a strategic pivot to avoid actual regulation.

