What the AI Labs Aren’t Telling You About Safety

AI labs in 2026 are not telling you the full story about their safety work. They are not telling you because the full story is not flattering, because the full story would invite regulatory scrutiny, and because the full story…

A single closed dark door at the end of a long corridor, dim warm amber light source beyond the door, deep navy shadows throughout, no people visible.

AI labs in 2026 are not telling you the full story about their safety work. They are not telling you because the full story is not flattering, because the full story would invite regulatory scrutiny, and because the full story would slow down the product release. The patterns in what the labs are not saying are predictable enough to be worth cataloguing.

OpenAI, Anthropic, Google DeepMind, Meta, xAI, and the Chinese labs (DeepSeek, Qwen, the Baidu ERNIE team) are all publishing safety reports. The reports are 30 to 100 pages long. They are filled with charts, frameworks, and case studies. They are not filled with the things the safety teams would tell you if the safety teams were the ones writing the press release. The gap between the report and the reality is where the next safety incident is coming from.

What the safety reports include

Three categories, in roughly that order of how often they appear. The framework category: where the lab describes the safety framework it has adopted (the Anthropic Responsible Scaling Policy, the OpenAI Preparedness Framework, the Google DeepMind Frontier Safety Framework). The frameworks are real. The frameworks are also aspirational, in the sense that they describe the safety work the lab plans to do, not the safety work the lab has done. The case study category: where the lab describes a safety incident or a safety intervention in detail. The case studies are real. The case studies are also the ones the lab chose to publish, which is not the same as the case studies the lab did not choose to publish. The metric category: where the lab publishes a chart of safety metrics over time. The metrics are real. The metrics are also the ones the lab chose to measure, which is not the same as the metrics the lab did not choose to measure.

What the safety reports leave out

Three categories of stuff that the safety reports systematically under cover. The near miss category: the incidents that almost happened, the interventions that almost did not work, the safety team members who almost quit. Near misses are the most valuable category of safety information because they are the ones the next lab can learn from. They are also the ones the labs are least likely to publish. The disagreement category: the safety teams disagree with the product teams about what to ship, the safety teams are overruled, the product ships anyway. The disagreements are the most valuable category of safety information because they are the ones that predict the next incident. They are also the ones the labs are least likely to publish. The off the record category: the conversations the safety teams have with regulators, with academic collaborators, with the press, that never make it into the safety report. The off the record category is the most valuable because it is the most candid. It is also the one the public never sees.

How to evaluate the safety claims

Three moves if you are using an AI model in production and you need to evaluate the safety claims. Read the safety report, but treat it as a marketing document, not as evidence. Look for the specific commitments, the specific metrics, the specific incident response procedures, and ask whether each one is verifiable. Ask the lab directly about the near misses, the disagreements, and the off the record conversations, and treat the silence as data. Build your own safety evaluation, with your own red team, with your own safety metrics, because the lab’s safety evaluation is not your safety evaluation. The enterprise that builds its own safety evaluation is the enterprise that knows what the model is actually doing in production.

Abstract safety metrics as glowing cyan columns of varying heights on a dark navy surface, dramatic chiaroscuro lighting from above.
AI safety in 2026: the gap between the safety report and the reality is where the next safety incident is coming from. The near misses, the disagreements, the off the record conversations.

The bottom line

AI lab safety reports are real documents with real information, but they are not the full story. The full story lives in the near misses, the disagreements, and the off the record conversations. Enterprises that treat the safety report as a starting point and build their own safety evaluation are the ones that know what the model is actually doing in production.

Sources & Further Reading

All claims in this article are sourced from primary documentation, vendor advisories, and reputable security researchers.

Spotted an error? Email the editor. Corrections are issued with a visible correction note.

Editorial standards. Every article on humanrequired.org is reviewed by a human editor before publication. AI may assist with drafting or research; final editorial control is human. Read the full standards.

Continue reading