The AI Safety Industrial Complex Is Asking the Wrong Questions

The most famous AI safety organisations in the world are funded by the same companies building the most powerful AI systems. That is not a coincidence, and it is not nothing.

A single brass question mark on dark wood, dim warm amber side light, deep navy shadows, no people, no logos.



The most famous AI safety organisation in the world is funded by the same companies building the most powerful AI systems. That is not a coincidence, and it is not nothing.

The phrase “AI safety” in 2026 means different things to different people. To the press, it usually means the risk of an artificial general intelligence turning on its creators, the kind of storyline that gets clicks and that policymakers like to talk about. To the researchers inside the major labs, it usually means alignment work, the technical problem of making sure a sufficiently advanced system does what its operators want. To the people actually deploying AI in production, it usually means something else entirely, and that something else is almost never what the famous safety organisations are working on.

The gap between the safety conversation and the safety problem is wide, and it is widening.

What AI safety usually means in 2026

The headline safety story is existential risk. Will a frontier model, given enough capability, decide that humanity is in the way? Will a misaligned optimisation process paperclip us out of existence? Will a single rogue agent, given access to enough compute, end civilisation? These questions dominate the press cycle and the policy discourse. They are also, at least as currently framed, mostly unanswerable. We do not have a working definition of “decide”. We do not have a working definition of “civilisation ending”. We certainly do not have an empirical basis for predicting the behaviour of a hypothetical system that does not yet exist.

What we do have, today, in 2026, is a working definition of “the model leaked a customer’s data through a prompt injection”. We have a working definition of “the agent booked the wrong flight for the CEO and we cannot get a refund”. We have a working definition of “the model refused to summarise a document because it thought the topic was unsafe”. These are not existential risks. They are not even catastrophic. They are, however, the actual safety problems of actual AI deployments, and they are the ones that nobody in the safety conversation is talking about at scale.

Why the existential risk framing serves the labs

The existential risk framing is not, on its face, absurd. A sufficiently advanced system might, in principle, pose risks we cannot model. Taking that possibility seriously is reasonable. The problem is not the framing. The problem is what the framing displaces.

An organisation whose safety mandate is “prevent the creation of an unfriendly superintelligence” has a very different budget, staff, and research agenda from an organisation whose safety mandate is “make sure the AI products we ship do not break in the next 12 months”. The first kind of organisation can justify a hundred person research team, a 10 year time horizon, and a billion dollar endowment. The second kind needs engineers who are also product managers, who ship fixes, who measure things in weeks.

A Venn diagram on a dark background with two circles, one labeled 'AI safety orgs' and one labeled 'AI deployment teams', the circles barely overlap, suggesting almost no shared concerns.
The two circles share a name and almost nothing else.

When the existential risk organisations define the conversation, they capture the funding, the press, the policy attention. The deployment teams, who are dealing with the actual problems, are left to figure it out themselves, with no budget, no researchers, and no public narrative that recognises their work as safety work.

The real safety problems, in order of frequency

Prompt injection delivered through untrusted content, the single most common production failure mode of agentic systems. Data leakage through tool calls, embedding APIs, telemetry, error reports, vector stores, agent memory, every one of which is a place the prompt can exit the system. Jailbreaks that bypass system prompts and get the model to ignore its instructions. Agent autonomy that goes further than the user expected, the kind of failure where the agent books three things instead of one and you discover it from the credit card statement. Hallucinated tool calls, where the model invents a tool result that never happened and acts on it. Confused deputy problems, where the agent uses one user’s permissions to do something another user wanted. Supply chain attacks on MCP servers, model files, training datasets. Bias that survives alignment work and shows up in production decisions at scale.

None of these are existential. All of them are happening, right now, in production systems that real people are using. None of them are the focus of the headline safety conversation. All of them would benefit from serious research, serious funding, and serious talent.

Where the safety orgs are and where the actual problems are

The headline safety organisations are concentrated in San Francisco, London, and a few university towns. The actual safety problems are everywhere an AI system is in production, which is to say, in every enterprise in the developed world. The companies with the most to lose from AI failures are the financial institutions, the healthcare systems, the law firms, the governments deploying AI for decisions that affect real people. None of these organisations have AI safety teams. Most have a compliance officer and a vendor relationship with a model provider. That becomes the state of the field in 2026.

The result is that the safety conversation is detached from the safety work. The researchers are studying problems that may or may not matter in 30 years. The practitioners are dealing with problems that matter today, with no research support, no shared knowledge base, and no agreed vocabulary.

The regulatory capture problem

The same companies that fund the safety organisations also fund the policy organisations that advise governments on AI regulation. The result, in 2026, is a regulatory environment that reflects the labs’ interests more than the public’s. Frontier model evaluations are the headline compliance regime. Existential risk stands as the headline concern. The things that affect ordinary users, like transparency about training data, like meaningful consent for personalised inference, like the right to know when a decision was made by an AI, are background noise.

This is not unique to AI. Regulatory capture is a feature of every industry that gets big enough to influence its own regulators. The fact that it is happening here does not mean the people involved are acting in bad faith. Many of them are sincerely trying to balance a hard problem. The outcome, however, sits as the same. The conversation is shaped by the people who can afford to participate in it.

What useful AI safety work would look like

Useful AI safety work in 2026 would look more like medical device safety than like AI alignment research. It would be empirical, not theoretical. It would study real systems in production, with real users, under real conditions. It would measure failure rates, not hypothetical future risks. It would publish incident reports the way the aviation industry does, with technical detail, anonymised data, and a shared taxonomy of failure modes. It would fund red teams that try to break deployed systems, not imagined ones. It would study the boring middle, where most of the actual harm happens, instead of the speculative extremes.

This is unglamorous work. It does not get a TED talk. It does not get a Senate hearing. It is, however, the work that would actually make AI systems safer for the people using them today.

A practical split: who is asking the right questions

The organisations asking the most useful safety questions in 2026 are not the headline ones. They are the small academic groups studying jailbreak robustness. The red teams at major model providers. The compliance teams at financial institutions deploying AI for credit decisions. The civic tech organisations documenting AI failures in public services. The independent security researchers publishing proof of concept exploits. The product teams at companies that have shipped AI to real users and learned the hard way what breaks.

These are not the voices in the safety conversation. They should be.

The bottom line

AI safety is a serious field with a serious problem. The problem is not that the work is not being done. The problem is that the work being funded and platformed is not the work that would make the systems people are using today safer. The gap between the safety conversation and the safety work sits as the single largest unaddressed risk in the AI industry, and it is not going to be closed by the organisations that benefit from keeping it open.

If you are deploying AI, your safety work is your own. The headline organisations are not going to do it for you. The funding is not going to flow to the problems that affect you. The press is not going to cover the failures in your stack. The regulators are not going to require the controls you actually need.

What you can do is treat AI safety like any other production safety discipline. Measure it. Red team it. Document the failures. Share the lessons. Build the muscle that the field as a whole has not yet built. The companies that do this will deploy AI that does not break. The companies that do not will read about themselves in a future post incident report, and the headline safety organisations will not be there to help, because the problem you had was not the problem they were studying.


Sources & Further Reading

All claims in this article are sourced from primary documentation, vendor advisories, and reputable security researchers.

Spotted an error? Email the editor. Corrections are issued with a visible correction note.

Editorial standards. Every article on humanrequired.org is reviewed by a human editor before publication. AI may assist with drafting or research; final editorial control is human. Read the full standards.

Continue reading