2
1 Comment

When AI Gets Cheaper, Who Actually Wins?

GLM 5.2 just made the token price war official. But the enterprise buyer celebrating cheaper models may be solving for the wrong number.

---

A client called us last week with a screenshot.

GLM 5.2 pricing. Token costs that would have looked like a typo two years ago.

"We're renewing in six months. I'm sending this to you now so we can talk about repricing."

It's a reasonable move. If the underlying model costs drop by 70%, shouldn't your AI vendor fee follow?

Here's the problem with that logic: it assumes the thing you're paying for is the model.

---

What the token price war actually affects

The GLM 5.2 announcement is real. So is the broader trend — Claude, GPT, Gemini, and now domestic Chinese models are in a race where the floor keeps dropping. Model pricing is commoditizing faster than anyone predicted three years ago. The Jevons paradox is in full effect: as models get cheaper, enterprises use more of them, and total AI spend goes up even as per-token cost collapses.

But here's what we've tracked across our deployments: compute costs — the thing that's actually getting cheaper — averaged 11% of total project spend. The range was 7% to 16%.

The other 89% doesn't appear on any model provider invoice.

It distributes like this: data engineering — cleaning, normalization, deduplication, pipeline construction — runs 35 to 45% of total cost. Permission and integration work — security reviews, API connectors, auth flows, the vendor assessments that no one scoped — runs 25 to 35%. Ongoing maintenance and tuning once the system is live: 15 to 20%.

GLM 5.2 getting 70% cheaper is a real number. Seventy percent of eleven percent is about 7.7 percentage points off your total project cost. That's not nothing. It's also not the number your CFO thinks it is when they see that screenshot.

---

Model commoditization ≠ deployment commoditization

The buyer's intuition is: AI models are getting cheaper, therefore AI is getting cheaper.

That's only true if the model is the bottleneck. In almost every non-AI-native enterprise we've worked with — auto retail, medical devices, manufacturing, real estate — the model was never the bottleneck.

The bottleneck was organizational. Data that lived in seven systems with three different naming conventions for the same customer. Security reviews that had a six-week queue. The senior analyst who was the only person who knew what "production variance" actually meant in context — and whether it was measured from order confirmation, shipping, or physical run completion.

We wrote a piece earlier in this series about a manufacturing client who wanted to automate variance reporting. We spent two weeks in their systems before touching a model. We found three different definitions of the core metric across three different systems. The twelve-hour human process we'd been asked to automate existed, in large part, to reconcile those definitions every week.

There was no "production variance" to automate. There was a twelve-hour human arbitration process wearing the costume of a reporting task.

GLM 5.2 doesn't change that. Neither does any model price drop. The data still needs cleaning. The definitions still need reconciling. The organizational process still needs mapping before any model can do anything useful with it.

---

What actually happens when models get cheaper

Here's the counterintuitive outcome of the token price war: cheaper models don't simplify enterprise AI. They expand it.

When compute costs drop, the rational enterprise response is to use more models, in more places, for more decisions. A company that was running one AI process in 2024 might be running five by 2026 — because the marginal cost of adding a new one has fallen to nearly zero.

But each new model deployment carries the same organizational overhead that the first one did. You still need data pipelines. You still need integration work. You still need the process audit that tells you which decisions in the workflow are actually deterministic and which are judgment calls wearing the costume of rules. You still need the accountability structure — the document that answers the medical director's question: "When the AI got that routing decision wrong, who authorized it to make that call?"

More models, same per-model setup cost, larger total footprint. The integration complexity scales with the number of systems, not with the per-token price.

This is the part of the conversation that doesn't fit on a screenshot.

---

The real negotiation

When our client sends us a GLM price comparison as a repricing signal, they're opening a negotiation. But they're negotiating on the wrong line item.

What they're actually buying from us isn't model access. They could buy model access directly — and increasingly they do, for the raw inference layer. What they're buying is the work that makes the model usable inside their actual organization: the data readiness work, the integration development, the process audit before anyone touches a model, the done-state specification that prevents a reporting agent from generating 847 versions of the same report over a weekend because nobody defined what "complete" looked like.

We've written about all of these separately. The common thread: none of them get cheaper when the model gets cheaper.

A 40% reduction in API costs on 11% of total spend is a 4.4% improvement in total project economics. That's the actual math behind the screenshot. It's real. It's also a different conversation than repricing the engagement.

---

Who actually benefits from the margin collapse

The AI margin collapse is genuinely bad news for companies whose revenue is tied to model access — the API wrappers, the thin-layer aggregators, the companies who were charging a premium for access to something that's now a commodity.

It's not automatically good news for enterprise buyers, because their real costs were never primarily in the model layer.

And it's quite good news for exactly one category of vendor: the ones who never charged for the model to begin with.

The shops whose value is in the deployment layer — the data engineering, the integration work, the process understanding, the organizational navigation that makes AI actually run inside a real company — those shops are largely immune to model commoditization. The client asking us to reprice based on GLM 5.2 is applying pressure to the line item that represents 11% of our total value. The other 89% isn't moving.

This was always true. The margin collapse just makes it visible.

---

The sentence that matters more than the screenshot

We have a diagnostic we run before every engagement: "If we close this deal, what one sentence does the buyer say to their board to justify the spend?"

The clients who send us token price comparisons are usually trying to justify a renewal to a CFO who's been reading about the AI cost collapse. The sentence they need to say to their board isn't "we renegotiated our API costs." It's something about what the system is actually doing — the cost line it's eliminating, the process it's replaced, the decision accuracy that's improved.

The model price is the easiest number in the room. It's on a vendor's website. It's in a press release. It benchmarks cleanly.

The harder number — the one that actually determines whether the project paid for itself — is the one nobody modeled before the project started. The fully-loaded cost of the analyst who spent 40% of her time being a data pipeline. The calendar time absorbed by IT security review. The ongoing tuning work that everyone assumed would be zero and always turns out to be 15% of the first year.

Cheaper models are a real tailwind. They're just not the wind the CFO screenshot is pointing at.

---

One thing we might be wrong about

The cost ratios we've described — 11% compute, 89% everything else — come from deployments in traditional industries with legacy data infrastructure. It's possible that as tooling matures — better data connectors, faster vendor assessments, standardized integration layers — the non-compute costs will compress. We'd be genuinely happy to be wrong about this. The structure we're describing isn't inherent to AI. It's the current shape of the problem in the industries we work in.

It's also possible that the model commoditization effect reshapes the market in ways that benefit enterprise buyers differently than we're predicting. If cheaper models accelerate the adoption of AI-native data infrastructure — if companies start building with data readiness as a first-class concern because they know they'll be running multiple models — the 89% might shrink.

But that's a three-to-five year story. The CFO with a GLM screenshot is asking about this renewal cycle. And for this renewal cycle, the number that moved is the one that represents about a dime on the dollar of what they're actually spending.

---

Working notes from B2B AI deployment in North America. Part of an ongoing series on what we keep noticing across wildly different industries — and what the industry isn't ready to say out loud.

posted toAvatar for product Carbuki
Carbuki
  1. 1

    The cost breakdown is a useful perspective, but I'd keep validating what enterprise buyers are actually hiring AI vendors to do. If they believe they're buying model access, cheaper models become a pricing discussion. If they're buying organizational capability—making AI work reliably inside a complex business—then the model is just one input, not the product.