1
0 Comments

The Benchmark Said 94%. The Client's Data Said Something Else.

Why the number that closes the procurement meeting is almost never the number that matters after deployment — and what to do about it.


The evaluation committee had done their homework.

Three AI vendors. Side-by-side comparison. Accuracy rates, latency benchmarks, integration scores. A spreadsheet someone had clearly spent real time building. When they called us in for the final pitch, the VP of Operations led with the table.

"Your intent recognition score is 89%. Vendor B is at 96%. Walk me through why we shouldn't just go with the higher number."

I asked one question: "Which dataset were those scores run on?"

Silence. Then: "The vendors provided their own test results."

We were being compared against benchmarks that were run on different data, in different conditions, measuring different things.


This is the moment I want to talk about. Not because it's rare — it happens in almost every competitive evaluation we've been through. But because the people in that room were not naive. They were doing exactly what procurement process tells you to do: collect comparable data, make an evidence-based decision.

The problem is that AI benchmark scores are not comparable data. They are marketing collateral that looks like comparable data.


The auto dealer evaluation: 89% vs. 96%

The dealership group above — seven stores in the Bay Area, high inbound call volume — ran a proper evaluation. They requested benchmark documentation from each vendor. Vendor B submitted a 94-page technical report. Their intent recognition score of 96% was tested against a standard voice AI benchmark corpus: 50,000 calls, mix of automotive, retail, and general customer service queries, clean audio, controlled conditions.

Our 89% was tested against 3,200 actual calls pulled from their existing phone system over the previous 90 days. Background noise from the service bay. Customers calling about specific car models only sold in that region. Dialect variation from the local market. A specific pattern of how their callers described transmission problems that nothing in the standard corpus had ever seen.

We asked them to run a blind test: take 200 calls from their own system and run both implementations on them.

Vendor B scored 71%. We scored 84%.

The benchmark had measured something real. It just wasn't their business.


The medical device evaluation: the number that procurement needed

A respiratory device company we worked with — FDA-regulated, hospital system clients — had a formal procurement committee. Three-stage evaluation. They needed a benchmark report as a required deliverable before any vendor could advance to the commercial discussion.

We could have submitted a standard benchmark. Instead we spent two weeks with their IT and clinical teams extracting three months of actual patient intake calls. De-identified, HIPAA-compliant, run through a proper annotation process. The resulting benchmark covered 1,100 calls, specific to their patient population, their device vocabulary, their intake workflow.

Our score on that benchmark was 81%. Our score on the industry-standard benchmark was 93%.

We submitted the 81%.

The procurement team pushed back. "Every other vendor submitted higher numbers. Why should we trust a lower score?"

Our answer: "Because it's the only score in this evaluation that will still be accurate six months after go-live."

They advanced us to commercial discussion. The committee chair told us later it was the first time a vendor had come in and argued against their own benchmark.


THE PATTERN WE KEEP SEEING

Benchmark scores in AI procurement function like credit scores in a loan application — they're proxies that help an institution make a defensible decision, not indicators of what will actually happen.

The problem is that credit scores are standardized. There is a shared definition of creditworthiness. AI benchmarks are not. Each vendor runs their own, on their own data, in their own conditions, using their own evaluation methodology. When a procurement team compares them side by side, they are comparing apples to conceptual descriptions of fruit.

This is not a vendor ethics problem. Standard benchmarks serve a legitimate purpose: they let you assess a system's general capability floor before you invest in a full evaluation. The failure is treating that floor as a ceiling — as if the number that measures general capability will hold in your specific deployment context.

It almost never does.


WHAT THIS ACTUALLY MEANS FOR PROCUREMENT

The number you bring into a procurement meeting is not the same as the number that will appear in your quarterly ops review.

First, ask every vendor which dataset their benchmark was run on. If they can't tell you, or if the answer is "standard industry corpus," what you have is a capability estimate, not a deployment prediction. Useful for shortlisting. Not useful for final decision.

Second, require a proof-of-concept on your own data before any commercial discussion. This is more work upfront. It is significantly less work than a failed deployment. In every evaluation we've run where we got a PoC on real client data before commercial terms, the gap between the standard benchmark and the actual performance was material — usually 8-15 percentage points in either direction.

Third, define what "accuracy" means in your context before you collect any numbers. Intent recognition accuracy means different things in a call center handling billing disputes versus a call center handling clinical intake. If the benchmark score doesn't carry a definition of what was being measured and what success looks like in that context, the number is not information. It is a well-formatted guess.


ONE THING WE MIGHT BE WRONG ABOUT

Our position is that client-specific benchmarks are more valuable than standard benchmarks. But this only works if the client has enough clean, labeled historical data to run a meaningful evaluation — and a lot of the companies we talk to don't.

If your historical call data is a mess (wrong labels, inconsistent tagging, low volume), then a client-specific benchmark might actually be less reliable than a well-run standard one. We've seen PoCs fail not because the AI was wrong, but because the "ground truth" data we were evaluating against was itself inaccurate.

In those cases, the right answer is probably: use the standard benchmark to shortlist, then invest in six to eight weeks of data cleanup before running any PoC. We've recommended this twice. Both times the client said the timeline was too long. One of them signed with the highest-benchmark vendor instead. We're still not sure how that deployment is going.


Working notes from B2B AI deployment in North America. Part of an ongoing series on what we keep noticing across wildly different industries — and what the industry isn't ready to say out loud. More at [zenaicorp.com](https://zenaicorp.com/en).

posted toAvatar for product Carbuki
Carbuki