A Thriftier Paradigm for AI

We all know that the AI hyperscalers are planning to spend hundreds of billions of dollars on massive data centers for generative AI.

But Gemini told me that a new paradigm is developing for business-oriented AI.

AI is already used in dedicated applications. This is orders of magnitude less expensive than generative AI which tries to absorb and interconnect all of human knowledge.

Dedicated AI applications are vastly less expensive to run and maintain than generalized frontier Generative AI models.

The economics of AI are split into two phases: Training (teaching the model) and Inference (running the model for a user).

Factor Dedicated AI (e.g., Pharma Discovery) Frontier Generative AI (e.g., GPT-4/Omni variants)
Model Scope Narrow and highly optimized. Built to understand specific matrices of data (like molecular bindings). Massive and generalized. Built to understand everything from conversational humor to Python code.
Parameter Count Typically millions to a few billion parameters. Hundreds of billions to trillions of parameters.
Training Costs Millions of dollars. High, but a bounded corporate R&D expense. $100 million to $500+ million per training run.
Inference Costs (Running it) Low. Simulating molecular structures is computationally heavy, but it is run on demand by a small, closed group of scientists. Astronomical. Millions of users hitting a trillion-parameter model simultaneously creates a massive, ongoing cash-burn on server infrastructure.

But the new paradigm is different from both.

It’s called “distillation.” The customer rents or buys only the applications their business uses over and over again. It’s like having a nurse practitioner monitor a patient with diabetes and only send the patient to a specialist if a complex new problem appears.

The advantages of distillation:

  1. It solves 90% of everyday, typical problems.
  2. It’s simple enough to install on the customer’s own servers for security and intellectual property protection.
  3. It’s much, much cheaper than generative AI. Gemini says that using generative AI for (say) summarizing e-mails is like using a particle accelerator to pop popcorn or a rocket to turn on a light switch.

Gemini told me, “The industry is aggressively pivoting toward “smaller, highly-distilled open-source models” (like Meta’s Llama series) that corporations can host on-premises or in smaller cloud environments to slash their inference bills.”

When I asked for an explanation:

That sentence is packed with dense tech-industry jargon, but it represents a massive shift in how businesses are trying to avoid going broke using AI.

Let’s unpack it by looking at the four key concepts that explain exactly what is happening:

1. “Smaller, highly-distilled…” (The Tech Downsizing)

When the generative AI boom started, the rule was “bigger is better.” Companies built massive frontier models with over a trillion parameters (the internal connections the AI uses to make decisions).

  • The Problem: Running a trillion-parameter model requires an absurd amount of expensive computer chips (GPUs) running 24/7.

  • The Solution: “Distillation.” Think of this like boiling down a massive, 50-gallon vat of soup into a highly concentrated, rich bouillon cube. Engineers take a giant, expensive model and train a much smaller model (say, 8 billion parameters instead of 1 trillion) to copy its behavior. The smaller model retains 95% of the smarts for specific tasks but requires a fraction of the computational muscle to run.

2. “…open-source models (like Meta’s Llama series)…” (The Free Blueprints)

Instead of renting access to an AI hidden behind a walled garden (like OpenAI’s GPT-4), Meta (Facebook) did something highly disruptive: they gave away the code and “weights” of their powerful Llama models for free. Anyone can download them.

  • Because it is open-source, any corporation can take Meta’s base model and tweak it to do a specific job without paying a subscription or royalty to a gatekeeper.

3. “…that corporations can host on-premises or in smaller cloud environments…” (Cutting Out the Middleman)

Because these distilled models are so small and free to download, companies don’t need to route all their data through Microsoft or OpenAI’s massive, expensive supercomputers.

  • On-premises means a company can run the AI directly on their own server in their own building.

  • Smaller cloud environments means they can rent a couple of standard servers on Amazon Web Services (AWS) or Google Cloud to run it quietly in a private corner of the internet. It protects their data security and keeps them independent.

4. “…to slash their inference bills.” (The Bottom Line)

As mentioned before, inference is the technical term for running the AI after it has already been trained—it’s the cost incurred every time an employee types a prompt and the AI generates an answer.

  • If a global bank has 50,000 employees asking a giant frontier model questions all day, their monthly bill to a hyperscaler would be astronomical.

  • By switching to a smaller, distilled, open-source model hosted on their own servers, they can do the exact same daily work (like summarizing emails or checking code) while cutting their computing energy and hosting costs by 80% to 90%.

The Macro Picture: It’s a classic economic correction. The hyperscalers built the massive, multi-billion-dollar infrastructure, but enterprise customers are realizing they don’t need a nuclear-powered rocket ship just to drive down the street. They are choosing smaller, cheaper, customized tools to protect their profit margins. [end Gemini quote]

The obvious flip side of the customers choosing smaller, cheaper, customized tools to protect their profit margins is lower earnings for the hyperscalers who are running the AI data centers.

The hyperscalers are reacting to this. Instead of selling a high-priced, luxury software subscription, the hyperscalers pivot to being digital utilities. They sell the computation of the “distilled” AI using their generative AI, their secure cloud service, and the data pipelines.

The real margins for hyperscalers over the next few years aren’t coming from simple chat windows. They are coming from Agentic Workflows—autonomous AI agents that talk to other AI agents to execute complex corporate tasks (e.g., automatically reconciling an inventory mismatch across a global supply chain).

They expect that this will become so inexpensive that it will pervade businesses everywhere. What they lose on “token” prices they will make up on volume.

Will this widespread adoption be fast enough to cover the immense expenditures in time to satisfy the markets? Or will it resemble the internet fiber optic cable build-out of the late 1990s which crashed in 2000 and took years to recoup the losses?

According to 2026 enterprise tech surveys, over 60% of large corporations are actively deploying or projecting to deploy autonomous agentic workflows within the next 24 months. Enterprise surveys show that while 97% of executives say their companies deployed AI agents, only 29% are seeing significant, scaled organizational ROI so far.

Will the revenue from this thriftier paradigm be enough to justify the expenditures without crushing hyperscaler stock valuations ?

To put the scale of the challenge in perspective, look at the staggering asymmetry between the money flowing out of hyperscalers versus the money flowing in from enterprise AI:

  • The Expenditures: The aggregate 2026 capital expenditure (CapEx) for just the five largest infrastructure giants (Amazon, Alphabet, Microsoft, Meta, and Oracle) is tracking toward $660 billion to $800 billion.

  • The Revenue: The total combined 2026 revenue for the entire cohort of major standalone generative AI model vendors (OpenAI, Anthropic, Cohere, Mistral, and Perplexity) is projected to be less than $35 billion.

Even though 80% of enterprise applications shipped are now embedding agentic components, the actual revenue generated by these distilled, efficient workflows is currently a drop in the bucket compared to the massive physical data center build-out.

What will the stock market do when the numbers show up in the financial reports of the hyperscalers? Only time will tell.

Wendy

5 Likes

The tech leaders n billionaires said 2-3 years ago, that AI LLMs would be commoditized.

These are the L&Bs of which I speak:
Jensen Huang, Elon, n the All In Pod crew: Chamath, David Sacks, Jason Calacanis, David Freiburg.

NVDA recently open sourced a lot of software n developer tools in conjunction with a deal in the EU.

ChatGPT accessed NVDA news releases:
("NVIDIA has been releasing a large collection of open models, datasets, and AI software tools while simultaneously partnering with European governments to build a “sovereign AI” ecosystem. These efforts came together during GTC Paris and other European announcements over the past year. �*)

Why would NVDA do this?
IMO, it’s cause the software is being commoditized.
By releasing the software n development tools, NVDA is building/maintaining the moat around its hardware.

:thinking:
ralph is biased cause he owns a couple NVDA shares.

4 Likes

Something like this is what bankrupted Global Crossing during the dot-com bubble bust. Prices dropped so fast that they could amortize the loans used to build the infrastructure. This could be the fate of the likes of OpenAI.

GoogleAI:

The massive overbuild of AI data centers and GPUs risks a “dark fiber” scenario similar to the dot-com era. If hardware depreciation and loan amortization outpace AI revenue growth—and compute costs plummet—companies like OpenAI face severe debt and valuation pressures. [1, 2, 3, 4]

The parallel between telecom in the 1990s and AI today highlights a very real structural risk: [1]

  • The Dot-Com Telecom Trap: During the late 1990s, companies like Global Crossing raised billions to lay undersea fiber-optic cables, banking on infinite demand. Technology improved so rapidly that the cost of transmitting data plummeted before they could pay off their loans, resulting in catastrophic overcapacity and the largest bankruptcy in telecom history. [1, 2, 3, 4, 5]
  • The AI Infrastructure Overbuild: Hyperscalers and AI startups are locked in an “arms race” that requires tens of billions of dollars in CapEx for GPU clusters and power infrastructure. Some estimates suggest that overbuilding infrastructure too quickly before North American power grids or sustainable business models can support it could lead to massive defaults on construction loans. [1, 2, 3]
  • The Revenue Reality Check: While major players like OpenAI are projecting billions in annualized revenue, the skyrocketing costs of training next-gen models, renting compute from cloud providers, and paying down debt are putting immense pressure on corporate balance sheets. [1, 2, 3]

Experts at institutions like the Bank for International Settlements (BIS) have explicitly warned that an aggressive, debt-fueled AI investment boom could turn into a severe market bust if returns fail to materialize.[1, 2]

The Captain

7 Likes

Since the early days of tech there’s been a push/pull between local/remote computing and generic/custom software. I accessed my first computer, a PDP-11, by a teletype.

PC’s brought hardware and software to the desktop. Companies now push software to the cloud and subscriptions. Desktops and phones are modern teletypes. Sun Microsystems proclaimed the network is the computer in the 1980s and still failed. When I was in college we had a home-grown email system. My kids’ schools use Gmail.

It’s interesting to see this play out with AI. Will each company hire their own AI engineers for custom intelligence or will businesses arise to build AI targeted at specific industries? Will the compute be local with companies running their own data centers or will they rent cloud computing that they “control”.

I suspect we will start to see companies begin to sell customized AI solutions and perhaps they would be good investments. Do any already exist?

2 Likes

I remember both Albaby1 and I (and probably others) predicting that the huge, one-size-fits-all model would not last. There is no reason for a medical office dealing only in their domain to also have AI that can do virtual home makeovers or choose the best new car.

We all know what will happen, we just don’t know when. Reality will - at some point - step in. It has to. I don’t know if it will be when two of them report some sort of disastrous quarters, or a report comes out from some research firm showing the folly of it all, or if a mild recession is sparked elsewhere (say, by oil supply constriction) but eventually it has to happen.

Even if the companies are profitable, and even with magic bookkeeping tricks pushing the expenses into the future, the costs are real, the revenue has to be real, or the inexorable hand of profitability will make itself heard.

6 Likes

What a fun discussion!… because most here have had doubts about the AI big scale explosion for some time.

  1. Now we will likely see the insanity of most of the already too deeply invested in giant AI quintupling down further so as to emerge as the sole winner of that race, saving their marbles maybe, and
  2. Apple (as the best publicized) going gosh almighty fast to dominate (with some carefully chosen allies?) the anti-godzilla smallish enterprise scale AI.

Buckle up, brace your heads, because (as many here have been gently saying for months) we may see a whole lot of whiplash in the financial markets soon…..?

3 Likes

The Chinese are masters at “distillation”.

intercst

1 Like

That would be agentic AI, and platforms are already being sold for fields such as legal or HR.

DB2

3 Likes