The AI Bubble Is Real. It Is Also Not the Problem.

There are two AI bubbles. The one in the press is about GPUs and data centres. The one nobody is talking about is the consulting and integration layer that has built up around AI deployment. The first one might pop.…

A single brass bubble on dark wood, dim warm amber side light, deep navy shadows, no people, no logos.



There are two AI bubbles. The one in the press is about GPUs and data centres. The one nobody is talking about runs as the consulting and integration layer that has built up around AI deployment. The first one might pop. The second one will quietly deflate, and the deflation will hurt more people.

The GPU bubble amounts to the easy one to describe. The big labs are spending $80 billion to $120 billion a year on AI infrastructure. The data centre build out counts as the largest in the history of the industry. The power agreements are bigger than the grid capacity of several US states. The narrative is that the capex will pay off in model revenue, that the labs will sell enough access to recoup the spend, that the bet is sound. The narrative is also the bet. Nobody knows if it pays off. The market is pricing it as if it does. That runs as the bubble.

The other bubble stands as the one nobody is writing about. It sits as the services layer. It counts as the consultancies, the system integrators, the boutique shops, the freelance prompt engineers, the “AI transformation advisors” who are charging $400 to $1,200 a day to help companies figure out which model to use, how to write the prompts, how to wire the API, how to add RAG, how to build the agent. The services bubble is bigger in headcount than the GPU bubble. It is also less durable. It is going to deflate, and when it does, the people working in it will need somewhere to go.

Where the headline bubble money is going

The headline bubble is straightforward. Capex. NVIDIA, AMD, Broadcom selling chips to Microsoft, Amazon, Google, Meta, Oracle, plus a long tail of neoclouds and sovereigns. The chips go into servers. The servers go into data centres. The data centres need power, water, cooling, network, land, permits, all of which cost money. The total addressable spend is hundreds of billions per year, growing fast, with the capital providers betting that the resulting compute will be sold at a margin.

The bet is not insane. The bet is also not guaranteed. The history of large capex cycles is that some pay off and some do not, and the difference is usually the elasticity of demand. The 1990s fibre bubble built infrastructure that is still in use, but only after several bankruptcies and a decade of write downs. The 2000s telecom bubble built similar infrastructure with a similar outcome. The cloud data centre build of the 2010s paid off because the demand for cloud services turned out to be enormous. The AI data centre build will pay off if the demand for AI inference turns out to be enormous, and will not pay off if the demand plateaus.

You do not need to pick a side on the GPU bet to see the services bubble. The services bubble exists regardless of whether the GPU bet pays off. The services bubble is its own phenomenon.

Where the actual bubble money is going

A vintage brass pressure gauge lying on its side on a dark wood surface, the cracked glass face with a single hairline fracture, the needle pushed into a red zone past the maximum mark, a warm tungsten desk lamp glowing softly in the background against a deep navy wall
The big cloud gets the press. The small cloud has the headcount.

Walk into any large enterprise in 2026 and ask who is doing the AI work. The answer is, more often than not, a consultancy. Accenture has a 30,000 person AI practice. The Big Four have built out their own. McKinsey, BCG, Bain, all have AI transformation groups. Deloitte’s AI practice grew faster than the rest of the company. The boutique shops, the “AI native” consultancies, have raised hundreds of millions to do the same work at smaller scale. The freelance market, Upwork, Toptal, specialised AI staffing firms, has rates that are 2x to 4x what the same person would earn as a full time employee.

The work, in most cases, is straightforward integration. Pick a model. Wire up an API. Write the prompt. Add the agent loop. Connect the data source. Set up the evaluation. Make sure the thing does not hallucinate too much. Make sure the credentials are scoped. Make sure the logs are there. None of this is research. None of it is novel. All of it is being done by people who learned it in the last 24 months and are charging a premium because the demand exceeds the supply.

The $400/day prompt engineer

The premium is real. A senior prompt engineer with a year of experience can command $400 to $700 a day on a contract basis, $1,200 a day through a consultancy. A “GenAI architect” with three to five years of related experience is $800 to $1,500 a day. These rates are 2x to 4x what the same person would earn as a full time engineer at the same company. The premium exists because the skills are rare, the demand is high, and the budget is uncapped.

The premium will not last. The skills are not actually rare. Prompt engineering is, in its core form, the ability to write clear instructions and iterate against a model’s output. The skill is teachable in a week. The iteration speed is bounded by the model’s response time. The premium is a function of FOMO and budget, not of underlying capability. The premium will compress as more people learn the skill, as the models get better at understanding rough instructions, and as the budget reality catches up with the hype.

Why so much of this work is one time

The consulting work on AI integration is, in the majority of cases, one time. A company hires a consultancy to figure out which model to use, wire up the integration, write the prompts, build the agent. The work takes 3 to 6 months. The work is done. The company has the integration. The company does not need the consultancy for the next integration, because the next integration will be done by the engineers who learned from the first one.

The pattern counts as the same one that played out in cloud migration, in ERP rollouts, in CRM implementations, in every other enterprise technology wave. The first wave of work is done by external consultants, who are expensive because the in house talent does not exist. The second wave is done by the in house team, who learned from the consultants and can now do the work at internal cost. The third wave is done by the next generation of engineers, who learned the work in school and never needed the consultants at all. The consulting market peaks and then compresses.

What happens when the one time work is done

The 30,000 person Accenture AI practice is sized for a market that does not exist at that scale for more than 2 to 3 years. The boutique shops that raised at $200 million valuations will see their revenue compress as the work moves in house. The freelance rates will fall to the underlying engineering rate, plus a smaller premium for genuine expertise. The people who built careers in AI consulting in 2024 to 2026 will, by 2028, need to find somewhere to go.

Some will go in house. Companies that have built AI practices will absorb the consultants, usually at lower rates than they were charging externally. Some will go to the model providers, who will always need people who can implement their products. Some will go to the next wave, whatever it is, and bring the AI integration skills with them. Some will leave the industry. The pattern becomes the same as every other consulting wave.

The honest forecast for the services market

The AI services market in 2026 is, by most estimates, somewhere between $80 billion and $150 billion in annual revenue. The forecast for 2030 is for the market to be smaller, not larger, in real terms. The reason is that the one time work will be done, the in house teams will be staffed, and the boutique shops will either have grown into product companies or will have been acquired or shut down. The forecast is uncomfortable for the people working in the market. The forecast is correct.

The forecast does not mean AI deployment is slowing. AI deployment is accelerating. The work is moving from external consulting to internal teams, from large engagements to small ones, from one time projects to continuous product work. The work is going somewhere. The somewhere is not Accenture.

The real opportunity, which runs as the unsexy middle

The companies and people who will do well in the next 5 years are the ones working on the unsexy middle. The model evaluation work that nobody wants to do. The integration plumbing that holds the system together. The data labelling and curation that makes the model useful in a specific domain. The security and compliance work that makes the deployment defensible. The reliability engineering that keeps the agent from breaking in production. The user research that finds what the model is actually good at versus what the demo suggested.

None of that becomes the press release work. All of it serves as the work that lasts. The press release work is what gets the consulting premium today. The unsexy middle is what pays the salary in 2030.

The bottom line

The AI bubble is real. It is also not the bubble that matters. The bubble that matters sits as the services layer that has built up around AI deployment, and the bubble that matters will deflate, and the deflation will hurt people. The people who planned for the deflation, by building durable skills, by working on the unsexy middle, by moving to internal teams before the consulting rates collapsed, will be fine. The people who treated the consulting premium as a permanent feature of their career will be reading about themselves in a future labour market report.

The GPU bet may or may not pay off. The services bet will not pay off. The services work will be done, the rates will compress, the headcount will redistribute. The work amounts to the same work. The compensation runs as the temporary part.


Sources & Further Reading

All claims in this article are sourced from primary documentation, vendor advisories, and reputable security researchers.

Spotted an error? Email the editor. Corrections are issued with a visible correction note.

Editorial standards. Every article on humanrequired.org is reviewed by a human editor before publication. AI may assist with drafting or research; final editorial control is human. Read the full standards.

Continue reading