We all know that the AI hyperscalers are planning to spend hundreds of billions of dollars on massive data centers for generative AI.
But Gemini told me that a new paradigm is developing for business-oriented AI.
AI is already used in dedicated applications. This is orders of magnitude less expensive than generative AI which tries to absorb and interconnect all of human knowledge.
Dedicated AI applications are vastly less expensive to run and maintain than generalized frontier Generative AI models.
The economics of AI are split into two phases: Training (teaching the model) and Inference (running the model for a user).
| Factor | Dedicated AI (e.g., Pharma Discovery) | Frontier Generative AI (e.g., GPT-4/Omni variants) |
|---|---|---|
| Model Scope | Narrow and highly optimized. Built to understand specific matrices of data (like molecular bindings). | Massive and generalized. Built to understand everything from conversational humor to Python code. |
| Parameter Count | Typically millions to a few billion parameters. | Hundreds of billions to trillions of parameters. |
| Training Costs | Millions of dollars. High, but a bounded corporate R&D expense. | $100 million to $500+ million per training run. |
| Inference Costs (Running it) | Low. Simulating molecular structures is computationally heavy, but it is run on demand by a small, closed group of scientists. | Astronomical. Millions of users hitting a trillion-parameter model simultaneously creates a massive, ongoing cash-burn on server infrastructure. |
But the new paradigm is different from both.
It’s called “distillation.” The customer rents or buys only the applications their business uses over and over again. It’s like having a nurse practitioner monitor a patient with diabetes and only send the patient to a specialist if a complex new problem appears.
The advantages of distillation:
- It solves 90% of everyday, typical problems.
- It’s simple enough to install on the customer’s own servers for security and intellectual property protection.
- It’s much, much cheaper than generative AI. Gemini says that using generative AI for (say) summarizing e-mails is like using a particle accelerator to pop popcorn or a rocket to turn on a light switch.
Gemini told me, “The industry is aggressively pivoting toward “smaller, highly-distilled open-source models” (like Meta’s Llama series) that corporations can host on-premises or in smaller cloud environments to slash their inference bills.”
When I asked for an explanation:
That sentence is packed with dense tech-industry jargon, but it represents a massive shift in how businesses are trying to avoid going broke using AI.
Let’s unpack it by looking at the four key concepts that explain exactly what is happening:
1. “Smaller, highly-distilled…” (The Tech Downsizing)
When the generative AI boom started, the rule was “bigger is better.” Companies built massive frontier models with over a trillion parameters (the internal connections the AI uses to make decisions).
-
The Problem: Running a trillion-parameter model requires an absurd amount of expensive computer chips (GPUs) running 24/7.
-
The Solution: “Distillation.” Think of this like boiling down a massive, 50-gallon vat of soup into a highly concentrated, rich bouillon cube. Engineers take a giant, expensive model and train a much smaller model (say, 8 billion parameters instead of 1 trillion) to copy its behavior. The smaller model retains 95% of the smarts for specific tasks but requires a fraction of the computational muscle to run.
2. “…open-source models (like Meta’s Llama series)…” (The Free Blueprints)
Instead of renting access to an AI hidden behind a walled garden (like OpenAI’s GPT-4), Meta (Facebook) did something highly disruptive: they gave away the code and “weights” of their powerful Llama models for free. Anyone can download them.
- Because it is open-source, any corporation can take Meta’s base model and tweak it to do a specific job without paying a subscription or royalty to a gatekeeper.
3. “…that corporations can host on-premises or in smaller cloud environments…” (Cutting Out the Middleman)
Because these distilled models are so small and free to download, companies don’t need to route all their data through Microsoft or OpenAI’s massive, expensive supercomputers.
-
On-premises means a company can run the AI directly on their own server in their own building.
-
Smaller cloud environments means they can rent a couple of standard servers on Amazon Web Services (AWS) or Google Cloud to run it quietly in a private corner of the internet. It protects their data security and keeps them independent.
4. “…to slash their inference bills.” (The Bottom Line)
As mentioned before, inference is the technical term for running the AI after it has already been trained—it’s the cost incurred every time an employee types a prompt and the AI generates an answer.
-
If a global bank has 50,000 employees asking a giant frontier model questions all day, their monthly bill to a hyperscaler would be astronomical.
-
By switching to a smaller, distilled, open-source model hosted on their own servers, they can do the exact same daily work (like summarizing emails or checking code) while cutting their computing energy and hosting costs by 80% to 90%.
The Macro Picture: It’s a classic economic correction. The hyperscalers built the massive, multi-billion-dollar infrastructure, but enterprise customers are realizing they don’t need a nuclear-powered rocket ship just to drive down the street. They are choosing smaller, cheaper, customized tools to protect their profit margins. [end Gemini quote]
The obvious flip side of the customers choosing smaller, cheaper, customized tools to protect their profit margins is lower earnings for the hyperscalers who are running the AI data centers.
The hyperscalers are reacting to this. Instead of selling a high-priced, luxury software subscription, the hyperscalers pivot to being digital utilities. They sell the computation of the “distilled” AI using their generative AI, their secure cloud service, and the data pipelines.
The real margins for hyperscalers over the next few years aren’t coming from simple chat windows. They are coming from Agentic Workflows—autonomous AI agents that talk to other AI agents to execute complex corporate tasks (e.g., automatically reconciling an inventory mismatch across a global supply chain).
They expect that this will become so inexpensive that it will pervade businesses everywhere. What they lose on “token” prices they will make up on volume.
Will this widespread adoption be fast enough to cover the immense expenditures in time to satisfy the markets? Or will it resemble the internet fiber optic cable build-out of the late 1990s which crashed in 2000 and took years to recoup the losses?
According to 2026 enterprise tech surveys, over 60% of large corporations are actively deploying or projecting to deploy autonomous agentic workflows within the next 24 months. Enterprise surveys show that while 97% of executives say their companies deployed AI agents, only 29% are seeing significant, scaled organizational ROI so far.
Will the revenue from this thriftier paradigm be enough to justify the expenditures without crushing hyperscaler stock valuations ?
To put the scale of the challenge in perspective, look at the staggering asymmetry between the money flowing out of hyperscalers versus the money flowing in from enterprise AI:
-
The Expenditures: The aggregate 2026 capital expenditure (CapEx) for just the five largest infrastructure giants (Amazon, Alphabet, Microsoft, Meta, and Oracle) is tracking toward $660 billion to $800 billion.
-
The Revenue: The total combined 2026 revenue for the entire cohort of major standalone generative AI model vendors (OpenAI, Anthropic, Cohere, Mistral, and Perplexity) is projected to be less than $35 billion.
Even though 80% of enterprise applications shipped are now embedding agentic components, the actual revenue generated by these distilled, efficient workflows is currently a drop in the bucket compared to the massive physical data center build-out.
What will the stock market do when the numbers show up in the financial reports of the hyperscalers? Only time will tell.
Wendy
