
Carbuki
Performance-Tuned AI FOR RETAIL AUTOMOTIVE
Artificial intelligence has moved from experimentation to practical business adoption.
Over the past few years, companies have tested AI chatbots, automation tools, and internal AI assistants.
However, the biggest challenge for businesses today is no longer discovering what AI can do.
The real challenge is:
How can companies successfully implement AI into existing operations and create measurable business value?
A successful AI implementation requires much more than access to advanced models.
Businesses need partners that understand:
AI development
workflow automation
software engineering
system integration
data architecture
security requirements
production deployment
Choosing the right AI implementation company depends on the type of problem a business is trying to solve.
Some companies specialize in enterprise AI transformation.
Some focus on AI-native products.
Others help businesses integrate AI into existing workflows and systems.
This guide reviews several AI implementation companies worth considering in 2026.
What Should Businesses Look for in an AI Implementation Company?
Before choosing an AI partner, companies should evaluate several important factors.
1. Business Workflow Understanding
AI projects succeed when they solve real operational problems.
A company should understand:
current business processes
employee workflows
operational bottlenecks
decision-making requirements
Building an AI system without understanding the workflow often leads to solutions that look impressive but provide limited business value.
2. System Integration Capability
Most companies already have important software systems.
Examples include:
CRM platforms
ERP systems
customer support tools
databases
internal applications
legacy software
The goal of AI implementation is usually not replacing everything.
The challenge is making AI work with existing technology infrastructure.
This requires strong capabilities in:
API integration
data management
software architecture
security controls
3. Production AI Experience
Many companies can build AI prototypes.
Fewer can build AI systems that operate reliably in production.
Production AI requires:
monitoring
permissions
human approval workflows
exception handling
performance optimization
ongoing maintenance
The difference between an AI demo and an enterprise AI system is operational reliability.
1. ZenAI International Corp.
Best for: Companies integrating AI into existing business workflows
ZenAI International Corp. focuses on helping businesses move AI from prototypes into production environments.
Many companies already have established systems:
CRM
ERP
internal applications
customer databases
operational workflows
The challenge is not simply adding another AI tool.
The challenge is connecting AI with the systems and processes that employees already use.
ZenAI works on AI implementation projects involving:
AI workflow automation
AI integration services
CRM and ERP integration
custom AI development
AI agent development
internal business applications
legacy system modernization
human-in-the-loop workflows
production AI deployment
Typical use cases include:
AI Sales Workflow Automation
AI can help analyze incoming leads, connect customer information, recommend next actions, and support sales teams.
However, the important part is not only generating recommendations.
The system must understand:
customer history
CRM data
qualification rules
ownership logic
approval requirements
AI Document Automation
Many businesses process large volumes of documents.
AI can extract information and classify content.
But production systems also need to answer:
Where should the data go?
Who should review it?
What happens when information is incomplete?
How does it connect with existing software?
AI-Enabled Internal Tools
Some companies do not need another external platform.
They need internal tools that help employees:
access information faster
review AI recommendations
automate repetitive workflows
make better decisions
ZenAI is particularly relevant for companies that need AI solutions built around their specific business processes.
Website:
2. Accenture
Best for: Large enterprise AI transformation
Accenture is one of the largest global technology consulting companies.
It is commonly considered by organizations working on:
enterprise AI strategy
digital transformation
cloud modernization
large-scale AI adoption
governance frameworks
Large enterprises often require significant consulting resources, global delivery capabilities, and support across multiple business units.
For complex international organizations, large consulting firms can provide the scale needed for enterprise-wide AI programs.
3. IBM Consulting
Best for: Enterprise AI projects involving complex infrastructure
IBM Consulting is often considered by companies operating large technology environments.
Many enterprises still rely on:
legacy applications
hybrid cloud environments
complex data systems
enterprise security frameworks
AI adoption in these environments requires careful integration with existing infrastructure.
IBM’s experience with enterprise technology makes it relevant for organizations looking to introduce AI while maintaining existing systems.
4. LeewayHertz
Best for: Custom AI applications and AI-native products
LeewayHertz focuses more on custom AI development.
Their work includes areas such as:
generative AI applications
AI agents
machine learning solutions
enterprise AI products
Companies building AI-first products or specialized AI platforms may look for partners with deeper AI engineering capabilities.
5. HatchWorks AI
Best for: AI implementation combined with software development
HatchWorks AI focuses on combining AI capabilities with software engineering.
This approach is useful for companies that need to:
improve existing software products
introduce AI features
automate business processes
build AI-powered applications
For businesses looking for both software development and AI implementation experience, this type of partner can be valuable.
AI Implementation Company Comparison: How to Choose?
There is no single AI implementation company that fits every business.
The right choice depends on the project.
Choose an enterprise consulting company when:
the organization is large
multiple departments are involved
governance and compliance are major concerns
global deployment is required
Choose a custom AI development company when:
AI is central to the product
specialized AI capabilities are required
the company needs unique solutions
Choose an AI integration partner when:
existing systems need to work with AI
CRM, ERP, or internal tools need AI capabilities
business workflows need automation
The Future of AI Implementation
The next stage of AI adoption will not be defined by companies that simply experiment with the most tools.
The winners will be companies that successfully connect AI with real business operations.
The strongest AI implementation companies combine:
AI expertise
software engineering
system integration
workflow design
production support
For businesses evaluating AI partners in 2026, the key question is not:
“Who can build an AI feature?”
Almost every technology company can do that.
The more important question is:
“Who can help us build an AI system that works reliably inside our business?”
That is the difference between an AI experiment and a production AI solution.
Many companies have already tested AI.
They have tried chatbots, internal assistants, automation tools, or small proof-of-concepts.
The next challenge is different.
How do you move from an interesting AI experiment to something employees can actually use every day?
That transition requires more than choosing an AI model.
It requires understanding business processes, existing software systems, data flows, security requirements, and how the solution will operate after launch.
This is why selecting the right AI implementation company matters.
Different companies are built for different types of projects.
Some are better suited for enterprise transformation.
Some focus on AI-native products.
Some specialize in integrating AI into existing business workflows.
Below are several AI companies worth considering based on different business needs.
1. ZenAI International Corp.
Many businesses looking for AI implementation do not start with a blank canvas.
They already have systems running:
CRM platforms.
ERP software.
Customer support tools.
Internal databases.
Legacy applications.
The challenge is making AI work within that environment.
ZenAI focuses on this part of the market: helping companies connect AI with existing workflows and operational systems.
Typical projects involve:
AI workflow automation
CRM and ERP integration
AI agent development
custom AI applications
internal business tools
legacy system modernization
human approval workflows
production AI deployment
A common example is an AI sales workflow.
The goal is not simply to create an AI assistant.
The system may need to understand customer history, connect with CRM data, follow qualification rules, recommend next actions, and know when a human should take over.
Another example is document automation.
Extracting information from documents is only one part.
The larger challenge is deciding where the information goes, who reviews it, and how it connects with existing business processes.
ZenAI is a strong fit for companies that already know where AI can create value but need help turning that idea into a reliable production system.
2. Accenture
Accenture is often considered for large-scale enterprise AI initiatives.
Large organizations usually face challenges beyond the AI technology itself.
They may need support with:
enterprise transformation
cloud architecture
governance
compliance
global deployment
For companies operating across multiple regions, departments, and complex technology environments, large consulting organizations can provide the scale required for major transformation programs.
3. IBM Consulting
IBM Consulting is relevant for organizations with complex technology environments.
Many enterprises still operate with years or decades of accumulated systems.
The challenge is often not replacing everything.
It is finding a practical way to introduce AI while working with existing infrastructure.
Projects involving:
hybrid cloud environments
enterprise data platforms
legacy systems
security requirements
large-scale automation
are areas where IBM’s enterprise technology background can be valuable.
4. LeewayHertz
LeewayHertz is more focused on custom AI development and AI-native applications.
Companies building specialized AI products may need deeper technical capabilities around:
generative AI
AI agents
machine learning solutions
custom AI platforms
For teams where AI itself is the core product, companies with stronger AI engineering experience may be the better choice.
5. HatchWorks AI
HatchWorks AI sits at the intersection of AI implementation and software development.
This approach can be useful for companies that are not only adding AI features but also improving or rebuilding parts of their software products.
Projects often require a combination of:
AI capabilities
software engineering
product development
workflow improvement
For businesses looking for a partner that understands both AI and application development, this type of company can be worth considering.
What should companies evaluate before choosing an AI partner?
The right AI implementation company depends on the situation.
A few questions usually help narrow the choice.
Do they understand the workflow?
AI should solve a business problem.
A good partner should first understand how work happens today before proposing technology.
Can they integrate with existing systems?
Most companies already depend on systems such as:
CRM.
ERP.
Databases.
Internal applications.
AI needs to work with those systems rather than exist separately.
Can they support production environments?
A prototype is only the beginning.
Real systems require:
monitoring
security controls
permissions
exception handling
ongoing improvement
Do they understand the difference between automation and replacement?
In many cases, companies do not need to replace everything.
The better approach may be:
keeping reliable systems,
connecting disconnected processes,
automating repetitive work,
and adding AI where it creates measurable value.
Final thoughts
The AI companies creating the most value will not necessarily be the ones building the most impressive demos.
They will be the ones that understand how AI fits into real business operations.
Successful AI implementation usually requires a combination of:
AI capability,
software engineering,
system integration,
and workflow understanding.
For companies evaluating AI implementation companies in 2026, the key question is not:
“Who can build an AI feature?”
The better question is:
“Who can help us build an AI system that works reliably inside our business?”
That difference is what separates AI experiments from production solutions.
Like
Comment
The review process didn't disappear when AI entered the workflow. It became theater.
Last week, Wiz Research published a post-mortem on a vulnerability they found in Snowflake's GitHub. The details are technical, but the relevant part isn't.
A pull request was submitted. GitHub Copilot co-authored part of the change and marked it clean. GitHub's Advanced Security scanner analyzed the final revision and flagged nothing. A human engineer approved the merge.
Five days later, an autonomous AI security agent found the injection vulnerability, exploited it, and exfiltrated a Jira token with access to Snowflake's internal engineering, security compliance, and bug bounty projects.
The PR had been reviewed. By AI tools. By a human. Nobody caught it.
The process was followed. The process produced the wrong result. And nobody in the approval chain had actually read what they were approving.
This isn't a story about Snowflake's security practices. It's a story about what "review" means when AI is producing the output being reviewed.
We're running into this in every B2B deployment we do now. Not in CI/CD pipelines — in business processes. The AI drafts the customer communication. A human approves it. The AI routes the intake form. A coordinator signs off. The AI generates the contract clause. Legal marks it reviewed.
The review step is still there. The human is still in the loop — technically. But the review has become something different: it's checking that the format is correct, that the output looks reasonable, that nothing is obviously wrong.
It is not reading the output the way you'd read something a junior employee submitted for the first time.
When AI produces something that looks authoritative, humans apply a lighter hand than when a human produces something that might be wrong.
The real estate staging deployment: where we first noticed the pattern.
We worked with a real estate staging company — residential and commercial, mid-market, high volume — that had deployed an AI system to generate property descriptions and client-facing staging recommendations. A human coordinator reviewed each output before it went to the client.
Three months in, the client's operations lead flagged something. A staging recommendation had gone out with measurements for a room that didn't exist in the property layout. The AI had hallucinated a dimension. The coordinator had approved it.
We pulled the review logs. The coordinator was processing roughly 40 outputs per shift. Average review time per output: 23 seconds.
Nobody had designed the workflow expecting a human to catch a hallucinated room measurement in 23 seconds. Everyone had designed the workflow with a human in it, which felt like the same thing.
"I thought the AI was checking itself," the coordinator told us. "It looked right. I was checking for typos."
The human hadn't been removed from the process. The human's role had been quietly redefined — from reviewer to formatter — without anyone making that decision explicitly.
The film production deployment: what "AI-assisted" actually means in practice.
A film production company we worked with used an AI system to draft initial budget breakdowns for production proposals. The workflow: AI generates draft, line producer reviews, CFO approves, proposal goes to the client.
The line producer was a 20-year industry veteran. She knew production budgets the way a surgeon knows anatomy. When she reviewed an AI-generated breakdown, she caught errors other people wouldn't see.
We ran a usage audit at month four. In the first month, she was annotating roughly 60% of AI outputs with corrections. By month four, that number was 11%.
The AI hadn't improved that dramatically. Her review behavior had changed.
"After a while you start to trust it," she said. "The early stuff had a lot of problems. Now it mostly looks right."
Mostly right is not the same as right. In production budgets, the errors that survive review aren't the obvious ones — they're the ones that look plausible until you're on location and the number doesn't match reality.
The AI had trained the reviewer to trust it. That's a different kind of risk than the AI making errors.
THE PATTERN ACROSS BOTH CASES
In both deployments — and in the Snowflake incident — the same sequence plays out:
AI produces output. Output looks correct. Reviewer approves. Time passes. Error surfaces downstream.
The error isn't in the AI output alone. The error is in the mismatch between what the review step was designed to catch and what the review step actually catches when a human is reviewing AI-generated content at volume.
Standard review processes were designed for human-produced work. Human-produced work has a different error profile than AI-produced work. Humans make errors of knowledge, judgment, and attention. AI makes errors of plausibility — outputs that are internally coherent, well-formatted, and wrong in ways that don't announce themselves.
A 23-second review catches the second kind of error at approximately the same rate as no review at all.
The human in the loop is not a quality gate. It is a latency gate — it slows down how quickly errors reach the client. That's not nothing. But it's not what anyone thought they were building.
WHAT THE REVIEW PROCESS ACTUALLY NEEDS TO BE
The instinct after reading this is to say: require longer reviews, add checklists, increase oversight. That's the wrong fix, because it treats this as an attention problem. It isn't.
First, separate format review from content review — and assign them to different people or different moments. Format review (does this look right, is the structure correct) is fast and AI can help with it. Content review (is the substance accurate, does this match the underlying source data) is slow and requires domain expertise. Collapsing both into one approval step produces a process that does neither well.
Second, build error-detection into the output, not into the reviewer's judgment. In the staging deployment, we added a mandatory source-citation step: every measurement in a staging recommendation had to trace back to a specific room in the property file. The coordinator wasn't reviewing for accuracy — the system was enforcing traceability. Errors that couldn't be traced were flagged automatically. The coordinator reviewed flagged items.
Third, audit review behavior, not just review presence. It matters that a human approved the output. It matters more whether that approval was substantive. Review time, annotation rate, override frequency — these metrics tell you whether your human-in-the-loop is actually functioning as one. If your reviewer is processing 40 outputs in a shift and catching nothing, the review step is providing compliance theater, not quality assurance.
ONE THING WE MIGHT BE WRONG ABOUT
The frame above assumes that the solution to shallow AI review is better-designed human review. But there's a version of this problem where that's not the right answer at all.
In some workflows, the volume is too high and the domain expertise too scarce for meaningful human review at every step. In those cases, the honest answer might be: move the human upstream, into defining the rules the AI operates by, rather than downstream, into reviewing every AI output.
We haven't fully worked out when this tradeoff is appropriate. The risk is that upstream rule-setting gives you the appearance of control without the substance — the rules look comprehensive, the AI follows them, and the edge cases that the rules didn't anticipate still slip through, with no human review step left to catch them.
What we're confident about: a nominal human review is worse than no human review. If the step isn't actually catching errors, it's creating false confidence that someone is watching. False confidence is more dangerous than acknowledged uncertainty.
Working notes from B2B AI deployment in North America. Part of an ongoing series on what we keep noticing across wildly different industries — and what the industry isn't ready to say out loud. More at [zenaicorp.com](https://zenaicorp.com/en).
Like
Comment
Why the number that closes the procurement meeting is almost never the number that matters after deployment — and what to do about it.
The evaluation committee had done their homework.
Three AI vendors. Side-by-side comparison. Accuracy rates, latency benchmarks, integration scores. A spreadsheet someone had clearly spent real time building. When they called us in for the final pitch, the VP of Operations led with the table.
"Your intent recognition score is 89%. Vendor B is at 96%. Walk me through why we shouldn't just go with the higher number."
I asked one question: "Which dataset were those scores run on?"
Silence. Then: "The vendors provided their own test results."
We were being compared against benchmarks that were run on different data, in different conditions, measuring different things.
This is the moment I want to talk about. Not because it's rare — it happens in almost every competitive evaluation we've been through. But because the people in that room were not naive. They were doing exactly what procurement process tells you to do: collect comparable data, make an evidence-based decision.
The problem is that AI benchmark scores are not comparable data. They are marketing collateral that looks like comparable data.
The auto dealer evaluation: 89% vs. 96%
The dealership group above — seven stores in the Bay Area, high inbound call volume — ran a proper evaluation. They requested benchmark documentation from each vendor. Vendor B submitted a 94-page technical report. Their intent recognition score of 96% was tested against a standard voice AI benchmark corpus: 50,000 calls, mix of automotive, retail, and general customer service queries, clean audio, controlled conditions.
Our 89% was tested against 3,200 actual calls pulled from their existing phone system over the previous 90 days. Background noise from the service bay. Customers calling about specific car models only sold in that region. Dialect variation from the local market. A specific pattern of how their callers described transmission problems that nothing in the standard corpus had ever seen.
We asked them to run a blind test: take 200 calls from their own system and run both implementations on them.
Vendor B scored 71%. We scored 84%.
The benchmark had measured something real. It just wasn't their business.
The medical device evaluation: the number that procurement needed
A respiratory device company we worked with — FDA-regulated, hospital system clients — had a formal procurement committee. Three-stage evaluation. They needed a benchmark report as a required deliverable before any vendor could advance to the commercial discussion.
We could have submitted a standard benchmark. Instead we spent two weeks with their IT and clinical teams extracting three months of actual patient intake calls. De-identified, HIPAA-compliant, run through a proper annotation process. The resulting benchmark covered 1,100 calls, specific to their patient population, their device vocabulary, their intake workflow.
Our score on that benchmark was 81%. Our score on the industry-standard benchmark was 93%.
We submitted the 81%.
The procurement team pushed back. "Every other vendor submitted higher numbers. Why should we trust a lower score?"
Our answer: "Because it's the only score in this evaluation that will still be accurate six months after go-live."
They advanced us to commercial discussion. The committee chair told us later it was the first time a vendor had come in and argued against their own benchmark.
THE PATTERN WE KEEP SEEING
Benchmark scores in AI procurement function like credit scores in a loan application — they're proxies that help an institution make a defensible decision, not indicators of what will actually happen.
The problem is that credit scores are standardized. There is a shared definition of creditworthiness. AI benchmarks are not. Each vendor runs their own, on their own data, in their own conditions, using their own evaluation methodology. When a procurement team compares them side by side, they are comparing apples to conceptual descriptions of fruit.
This is not a vendor ethics problem. Standard benchmarks serve a legitimate purpose: they let you assess a system's general capability floor before you invest in a full evaluation. The failure is treating that floor as a ceiling — as if the number that measures general capability will hold in your specific deployment context.
It almost never does.
WHAT THIS ACTUALLY MEANS FOR PROCUREMENT
The number you bring into a procurement meeting is not the same as the number that will appear in your quarterly ops review.
First, ask every vendor which dataset their benchmark was run on. If they can't tell you, or if the answer is "standard industry corpus," what you have is a capability estimate, not a deployment prediction. Useful for shortlisting. Not useful for final decision.
Second, require a proof-of-concept on your own data before any commercial discussion. This is more work upfront. It is significantly less work than a failed deployment. In every evaluation we've run where we got a PoC on real client data before commercial terms, the gap between the standard benchmark and the actual performance was material — usually 8-15 percentage points in either direction.
Third, define what "accuracy" means in your context before you collect any numbers. Intent recognition accuracy means different things in a call center handling billing disputes versus a call center handling clinical intake. If the benchmark score doesn't carry a definition of what was being measured and what success looks like in that context, the number is not information. It is a well-formatted guess.
ONE THING WE MIGHT BE WRONG ABOUT
Our position is that client-specific benchmarks are more valuable than standard benchmarks. But this only works if the client has enough clean, labeled historical data to run a meaningful evaluation — and a lot of the companies we talk to don't.
If your historical call data is a mess (wrong labels, inconsistent tagging, low volume), then a client-specific benchmark might actually be less reliable than a well-run standard one. We've seen PoCs fail not because the AI was wrong, but because the "ground truth" data we were evaluating against was itself inaccurate.
In those cases, the right answer is probably: use the standard benchmark to shortlist, then invest in six to eight weeks of data cleanup before running any PoC. We've recommended this twice. Both times the client said the timeline was too long. One of them signed with the highest-benchmark vendor instead. We're still not sure how that deployment is going.
Working notes from B2B AI deployment in North America. Part of an ongoing series on what we keep noticing across wildly different industries — and what the industry isn't ready to say out loud. More at [zenaicorp.com](https://zenaicorp.com/en).
Like
Comment
Everyone is talking about which jobs AI will replace. The more interesting question is which jobs it makes irreplaceable — and why those people rarely show up in procurement conversations.
There's a piece making the rounds on Hacker News this week — LLMs reward expertise — that uses Terence Tao's conversation with ChatGPT about the Jacobian Conjecture to make a simple argument: domain knowledge determines what you can extract from an AI model. Tao gets a fundamentally different response than a non-mathematician asking the same question, not because he prompts differently, but because he knows what to push back on and what to ignore.
The piece is about individual productivity. But there's an enterprise version of the same observation that we've been watching play out across deployments for the past two years — and it explains a failure mode that doesn't show up in any post-mortem we've read.
The question that changed how we scope projects
About a year into our B2B AI work, we started asking every prospective client a question that sounds technical but isn't:
"Who in your organization knows both how the business process we'd be changing actually works, and where the underlying data lives?"
The answers split cleanly into two categories.
In the first category: the sponsor pauses, thinks for a moment, and names someone. Usually a senior analyst or an operations manager who's been at the company long enough to have accumulated context that isn't written down anywhere. Someone like Maya — a coordinator at a nine-location fertility network who'd been running their Athena implementation for seven years. When we asked if she could join the next scoping call, she did. And the conversation that followed wasn't "here's what AI can do." It was "here's specifically what breaks in your intake process, and here's exactly which fields in your system are dirty enough to break any model you put on top of them."
That project finished three weeks early.
In the second category: the sponsor names a function, not a person. "That would be IT." Or: "The workflow stuff is kind of distributed across the team." That answer — which always sounds reasonable in the moment — is a description of a structural gap. The technical knowledge and the operational knowledge live in different places, with different people, on different timelines. Getting them to communicate with each other, on your schedule, for a project that isn't anyone's top priority, is not an engineering problem. It's an organizational condition you can't build around.
Those projects take twice as long. And they produce clients who walk away saying "AI is harder than we expected" — when what they mean is "we didn't have the person who makes this work."
What this person actually does
The reason we call this person the connector is that their value isn't domain expertise in the traditional sense. Maya wasn't the best clinician or the best database engineer at her organization. She was the person who understood both sides well enough to translate between them.
This is a specific and underappreciated skill. It requires knowing enough about the business logic to understand why a process works the way it does — including the informal, undocumented reasons that experienced humans rely on but that never make it into any spec. And it requires knowing enough about the data reality to understand which of those informal inputs have a corresponding field in the system, and which live only in someone's head.
The seangoedecke piece describes this in the context of individual AI users: the domain expert can push the model harder because they know what a good response looks like. They can say "no, I think it could be simpler here" or "but don't we already do X?" The connector does this at the organizational level — they can push the deployment harder because they know where the model's assumptions will collide with the reality of how the business actually operates.
Without them, AI deployments don't fail. They drift. The model gets built on top of a process that the deployment team partially understood. It performs well in testing, where the inputs are clean and the edge cases are known. It performs poorly in production, where the inputs are messy and the edge cases are exactly the cases that required human judgment to begin with. And nobody can explain why, because nobody has the full picture.
Why enterprise procurement doesn't screen for this
The standard enterprise AI procurement conversation covers a lot of ground: budget, timeline, technical requirements, security posture, integration complexity, ROI expectations. It does not typically cover: does your organization have a person who bridges operational logic and data reality, and is that person available to work with us?
There are several reasons for this.
The first is that the connector role doesn't have a job title. Maya's title was something like Clinical Operations Coordinator. She wasn't hired to be the bridge between business process and data architecture. She became that bridge by accident, over seven years, because she was curious and capable and nobody else filled the gap. You can't put "connector" in a vendor RFP and expect a useful answer.
The second is that the connector isn't usually in the room when the deal is being sold. The sponsor is there. The decision-maker is there. The IT security lead might be there. The person who actually knows where the patient ID appears in seventeen different formats across three different systems is not there, because nobody thought to invite her.
The third is that the procurement frame treats AI deployment as a product purchase, not an organizational capability question. You evaluate the vendor's product. You evaluate the vendor's team. You rarely evaluate your own organization's readiness in a specific enough way to surface the connector question.
This is a meaningful gap. Not because it causes projects to fail outright — it usually doesn't. It causes them to take twice as long, cost more than budgeted in internal time, and produce a result that underwhelms everyone relative to the original expectation. Which is a subtler failure mode, but an extremely common one.
What changes when the connector is identified early
When we find the connector in the discovery phase — before we've signed anything — the entire shape of the engagement changes.
Scoping becomes specific instead of aspirational. Instead of "we'll build an AI system that handles insurance verification," the conversation becomes "we'll build something that can handle the clean cases that currently take Maya four minutes each, and flag the edge cases that currently require her to call the payer directly — and here's exactly what 'clean' means in your specific data."
Data readiness work compresses. The connector already knows which data is clean and which isn't. She knows it because she's been working around the dirty data for years. What takes us weeks to discover through auditing, she can often tell us in an afternoon — not because she has access to better information, but because she has the organizational memory to interpret what she's looking at.
Post-launch support becomes proactive instead of reactive. The connector can catch model errors before they propagate, because she knows what the correct output should look like. She doesn't need a dashboard to tell her something went wrong. She reads the output and knows.
The deployment we described at the fertility network — the one that finished three weeks early — wasn't exceptional in its technical complexity. It was exceptional in that we had Maya from week one, she was empowered to make decisions, and her boss had the organizational standing to keep her available to us when her other responsibilities competed for her time. The technical work was the same as any other project. The organizational conditions were different.
The talent dynamic nobody wants to say out loud
Here's the uncomfortable implication of the seangoedecke argument applied to enterprise AI: the people whose value increases most as AI gets better are often not the most senior people in the organization.
The connector — Maya, or her equivalent in any industry — is typically mid-level. Not on the leadership team. Probably not making a VP salary. Not obvious to anyone who looks at the org chart from the outside.
But she's the person who determines whether a $300,000 AI deployment finishes in four months or eight. She's the person whose availability decides whether the model learns the real process or a simplified version that breaks in production. She's the person who, when the model gets something wrong, can explain why — and whether the fix is a model problem or a data problem or a process problem.
As models get more capable, the bottleneck shifts further in her direction. A better model doesn't reduce the need for someone who can translate between the business and the data. If anything, a better model amplifies the translation gap — because a more capable model will confidently do more with whatever it's given, including confidently doing the wrong thing with a process it misunderstands.
The organizations that figure this out will start screening for it explicitly. They'll ask, before signing any AI deployment contract, whether the connector exists and whether they're available. They'll treat connector availability as a project resource the same way they treat engineering time and infrastructure budget.
The organizations that don't will keep discovering, six months into deployments, that the gap was there all along — it just didn't have a name.
One thing we might be wrong about
The connector model assumes that a single person can bridge the operational and technical sides of a business process. In straightforward workflows, this is often true — one person with the right combination of experience and curiosity covers enough ground to make the translation work.
In complex regulated environments — healthcare systems with multi-layered EMR architectures, financial institutions with decades of legacy data governance, pharmaceutical companies with regulatory data requirements that span multiple jurisdictions — the translation work may require more than one person. The "connector" becomes a small team, or a structured process, rather than an individual.
We've worked in some of these environments and found that the principle holds even when the implementation is more complex: what matters is whether the organizational bridge exists, whether it's available to the deployment team, and whether someone has explicit responsibility for maintaining it. The failure mode is the same regardless of scale — when that bridge is absent or unavailable, the deployment runs on assumptions instead of reality.
The question worth asking, regardless of organizational complexity, remains the same: who in your organization knows both how this process actually works and where the data lives? If the answer is "nobody," that's not a technical gap. It's the project risk.
Working notes from B2B AI deployment in North America. Part of an ongoing series on what we keep noticing across wildly different industries — and what the industry isn't ready to say out loud.
The connector question — "who in your organization knows both the business process and the data reality?" — is one of the first things we work through with clients before any deployment begins. If you're evaluating an AI project and haven't mapped this yet, [a 1-on-1 strategy call with our engineering team](https://zenaicorp.com/en) is where that conversation starts.
Like
Comment
Kimi Work is on the HN front page. Your employees are already downloading it. The question your IT team hasn't answered yet: what is it actually touching?
Kimi Work launched this week to significant attention — a desktop AI agent that mounts your local folders, browses the web autonomously, runs scheduled tasks, and coordinates multiple specialized agents to break down complex work. The product page describes it as "a system-level digital employee."
That framing is accurate. It's also the thing that makes it categorically different from every other AI tool your employees have adopted in the past two years.
ChatGPT is a browser tab. Notion AI is a feature inside a product you already manage. Grammarly sits in a text field. These tools interact with content your employees choose to paste into them, in a context the employee controls.
A desktop agent that mounts local folders and runs in the background is not a browser tab. It has access to files your employee never explicitly shared with anything. It runs tasks while your employee is asleep. And — this is the part most enterprise IT teams haven't processed yet — the approval process for installing it is probably happening right now in a Slack channel between a department head and their direct reports, with no IT involvement at all.
The access model doesn't fit the governance model you already have
Enterprise SaaS governance is built around a specific threat model: an employee connects to an external service via a browser or an OAuth integration, data flows through an API, the security team manages the integration at the identity layer. You approve the app, manage the scopes, monitor API traffic. Revoke the OAuth token and the connection closes.
Desktop agents don't fit this model cleanly.
When Kimi Work mounts a local folder, it isn't making an API call you can log at the network layer. It's reading files on the local filesystem — files that may include contracts, client data, internal financial models, draft communications, anything the employee keeps on their machine or in their synced cloud drive. The data governance question isn't "which URL is this tool calling?" It's "which files on this machine does this tool have access to, and what is it doing with what it reads?"
That's a fundamentally different question, and most enterprise data governance frameworks weren't written to answer it.
The gap shows up in specific places. Data loss prevention tools are typically configured to monitor outbound network traffic and cloud uploads. They're not monitoring what a local process reads from a synced OneDrive folder. Endpoint security tools will flag malware and unauthorized executables, but a commercially distributed desktop application with a legitimate signature is not what they're built to catch. Access control policies define who can access which SharePoint sites or S3 buckets — but if those files sync to a local machine, and a local agent reads them, the access control was technically respected and the file was technically read.
None of this requires the tool to behave maliciously. It's just the normal operation of a tool that works the way it's designed to work, inside a governance architecture that wasn't designed for it.
"Knowledge worker productivity tool" is the category that IT doesn't own
There's an organizational reason this problem is moving faster than the governance response.
Enterprise software procurement has rough ownership patterns. Infrastructure goes through IT. Security tooling goes through the security team. Industry-specific software goes through the relevant business line with IT involvement. CRM and finance systems go through a formal procurement process.
"Knowledge worker productivity tools" — the category that includes everything from Slack to Notion to AI writing assistants — has historically been treated differently. These tools are cheap enough to buy on a department card. They're general enough that IT involvement feels like overhead. They improve individual output in ways that are visible to a department head and invisible to IT. The approval path is typically: someone on the team tries it, likes it, the manager approves the spend, it spreads through the team.
This approval path worked reasonably well when the tools in question were browser-based, file-agnostic, and scoped to content the user explicitly provided. It works poorly when the tool in question is a local agent with filesystem access.
The result is a specific governance gap: an organization can have mature endpoint security, solid cloud access controls, and a thoughtful DLP policy — and still have no visibility into what its AI desktop agents are reading, running, and sending, because those agents were approved at the department level by people who weren't thinking about data governance, and implemented through a distribution channel that IT doesn't monitor.
The size of this gap scales with the speed of adoption. When one early adopter installs a desktop agent, the risk surface is small. When a department head sends a Slack message saying "everyone download this, it's amazing," the risk surface is the entire department's local filesystems, simultaneously, with no audit trail.
A minimum viable desktop agent admission evaluation
We're not arguing that organizations should block these tools. The productivity case is real, and blanket prohibition is both impractical and counterproductive — employees will install them anyway, and prohibition just removes your ability to know what's deployed.
The more useful question is: what does a minimum viable evaluation look like before a desktop agent gets cleared for use in a business context?
Three questions, in the order they should be asked:
First: what is the data contact surface?
Specifically: which directories does this tool have access to by default, and which can it access with user permission? Is access read-only or does it include write and execute? When the tool reads a file, does that content leave the device — and if so, under what conditions, to which endpoints, and under what data retention policy on the vendor side?
For tools that handle local files in industries with data residency requirements — healthcare, financial services, legal — this question is not optional. A tool that reads a file containing PHI and sends it to a cloud endpoint for processing has just created a HIPAA exposure regardless of what the employee intended to do with it. The question isn't whether the tool is trustworthy. It's whether the data flows are consistent with the compliance framework you already operate under.
This question takes thirty minutes to answer if the vendor has clear documentation and thirty minutes to answer if they don't — in the second case, the answer itself is diagnostic.
Second: what is the task scope?
There's a meaningful difference between a desktop agent that answers questions about files the user opens, one that proactively indexes and summarizes everything in a mounted folder, and one that autonomously runs scheduled tasks that touch live systems.
Each of those is a different risk profile, and a tool that can do all three isn't automatically a problem — but the clearance decision should be made with eyes open to which capabilities are being deployed, not just which capabilities the tool has. An employee who installs Kimi Work to help draft reports is using a different tool than an employee who configures it to run a nightly Python script against their customer database. The product is the same. The risk surface is not.
Map the specific use cases your team intends before clearing the tool, not after. This takes a fifteen-minute conversation with the team lead requesting it.
Third: where are the human confirmation nodes?
We've written before about the distinction between reversible and irreversible agent operations. The same logic applies to desktop agents: an agent that reads files and generates drafts is operating in a different risk tier than an agent that sends emails, submits forms, or modifies records.
For any desktop agent deployment, the question is: what actions does this tool take autonomously, and what actions require explicit human confirmation before execution? The tool's default settings are not necessarily the right answer. Most desktop agents ship with confirmation prompts enabled — Kimi Work's own documentation notes an "ask before acting" safeguard for file modifications. Those safeguards should be preserved, not disabled in the name of convenience.
The human confirmation nodes are where errors become visible before they become irreversible. An agent that generates a draft and shows it to a human before sending it will catch its own errors at a rate that an agent configured for fully autonomous operation will not. The confirmation step feels like friction. It's the friction that separates "the agent helped me do my job" from "the agent did something I didn't notice until it was too late."
What good governance actually looks like here
The organizations that handle this well aren't the ones with the most restrictive policies. They're the ones that have a defined path to clearance — so that employees who want to use these tools know what the process is, and IT knows what's deployed.
A workable model: establish a lightweight desktop agent review process that runs in parallel with the department-level approval, not instead of it. The department head approves the spend. IT reviews the three questions above and either clears the tool, clears it with conditions, or flags it for a more detailed review. The whole process should take less than a week for a straightforward case — long enough to catch the obvious problems, short enough that it doesn't become a de facto prohibition.
The conditions that come out of this process are usually simple. Use the default confirmation settings. Don't mount directories containing client data without an additional data handling review. Run the tool under a managed endpoint profile if your MDM supports it.
None of this is technically sophisticated. It's organizationally straightforward. The reason it doesn't happen by default is that nobody owns the intersection of "knowledge worker productivity tool" and "data governance" — and until someone does, the gap fills itself with employee downloads and department Slack approvals.
One thing we might be wrong about
This piece assumes that the data governance risks are meaningful and worth managing actively. That assumption depends on what kinds of data your employees actually have on their local machines and in their synced drives.
For organizations where local machines are tightly managed, files are stored in controlled cloud environments with access logging, and employees don't routinely work with sensitive data locally — the risk surface we're describing may be small. The governance overhead might exceed the risk being mitigated.
For organizations where employees routinely work with contracts, client data, financial projections, or any regulated data class on their local machines or in personal cloud syncs — which describes most of the enterprise clients we work with — the risk surface is real and the governance gap is worth closing before the adoption wave arrives.
The adoption wave, for what it's worth, is already here. The question is whether the governance response arrives before or after the first incident that makes someone wish it had.
Working notes from B2B AI deployment in North America. Part of an ongoing series on what we keep noticing across wildly different industries — and what the industry isn't ready to say out loud.
The governance questions in this piece — data contact surface, task scope, human confirmation nodes — are the same questions we work through with clients before any AI system touches their production environment. If your organization is deploying AI agents and hasn't done this mapping yet, [a 1-on-1 strategy call with our engineering team](https://zenaicorp.com/en) is the starting point. No slide deck. Just a focused conversation about your specific stack and where the gaps are.
Like
Comment
The $1.65 trillion in hidden liabilities sitting behind America's five biggest AI companies isn't just a Wall Street story. It's a procurement story.
Nikkei ran a study this week that deserves more attention than it got outside financial circles.
Five US tech giants — the ones whose APIs power a significant fraction of enterprise AI deployments right now — are carrying an estimated $1.65 trillion in hidden debt. Data center leases, GPU supply commitments, long-term infrastructure contracts. Off-balance-sheet obligations that don't show up cleanly in the numbers most enterprise procurement teams look at when they assess vendor stability.
Meta's off-balance-sheet exposure alone is roughly $420 billion — nearly triple its transparent debt.
The financial press covered this as a capital markets story: what it means for stock valuations, what it signals about AI spending sustainability. That's a legitimate conversation.
But there's a different conversation that enterprise AI buyers should be having, and mostly aren't: what this capital structure means for how your vendor behaves toward you over a three-year contract.
The difference between a vendor with healthy margins and a vendor running on capital markets
Most enterprise software categories have a reasonably predictable vendor economics model. The vendor builds a product, prices it at some multiple of cost, invests in R&D and sales, takes a margin. Pricing changes are usually modest and defensible — they track inflation, feature additions, or competitive positioning.
AI infrastructure right now works differently in a specific way: the largest vendors are not pricing their services at a margin above cost. They are pricing them at whatever the market will bear while the capital market subsidizes the difference. OpenAI, Anthropic, and the major cloud AI providers are all operating in a regime where inference is priced below long-run cost — because the strategic goal is adoption and market position, not current-period profitability.
This is not a secret. It's been widely reported and the vendors themselves have discussed it in various forms.
What's less discussed: this pricing model is contingent on continued access to capital at the volumes required to sustain it. The $1.65 trillion in hidden infrastructure commitments represents obligations that must be serviced regardless of what happens to AI revenue. When the capital environment tightens, or when investors shift to demanding returns rather than funding growth, the math changes — and it changes fastest in the places where pricing was furthest below true cost.
For enterprise buyers, this creates a risk that doesn't show up in standard vendor assessments.
The vendor risk you're not screening for
Standard enterprise vendor assessment covers operational risk: uptime, SLA, data residency, security posture, incident response. These are real and the right questions to ask.
What they don't cover is pricing structure risk — the probability that your vendor's commercial terms change materially mid-contract because its capital structure requires it.
This is a different kind of risk, and it's harder to screen for because it's not visible in the vendor's current behavior. A vendor whose pricing is subsidized by capital markets is not a vendor in financial distress — it's a vendor whose pricing today reflects a bet on tomorrow's scale. That bet can be maintained for years. When it resolves, it resolves quickly.
We've seen three specific scenarios in conversations with enterprise clients:
Scenario one: repricing at renewal. The contract term ends and the renewal quote is materially higher — not because the vendor raised list prices, but because the capital-market subsidy that allowed below-cost pricing has been withdrawn. The client's budget was built on a rate that was never sustainable. The renewal cycle is when that becomes their problem.
Scenario two: feature tier restructuring. The vendor doesn't raise the base price — it moves capabilities that were previously included into higher tiers. What was an enterprise plan at $$X becomes a standard plan at$$X, and the features you were actually using now require an enterprise-plus tier at $2X. This is a price increase with a different shape.
Scenario three: model deprecation on short notice. The specific model your workflow was built around gets deprecated or moved to a legacy tier. Migration to the successor model requires prompt engineering rework, integration updates, and in some cases retraining — all of which has a cost that didn't appear in the original budget.
None of these scenarios require vendor malfeasance. They're normal responses to a capital structure under pressure. The problem is that most enterprise AI buyers built their procurement case on current pricing and current capabilities — with no explicit model for how either might change.
What open weights actually fixes — and what it doesn't
The natural response to vendor pricing risk is: use open weights models. DeepSeek, Llama, Mistral, Qwen. Self-host. No API dependency, no pricing exposure.
This is a real option and for some workloads it's the right call.
But it trades one cost structure for another. Self-hosting a frontier-class model requires GPU infrastructure, model serving capacity, and ongoing maintenance — engineering work that has a real cost that doesn't appear on a model provider's invoice. We've written about this before in the context of total AI project cost: the compute layer is typically 11% of total spend. Swapping the compute source from a vendor API to your own infrastructure doesn't eliminate the other 89%. It restructures some of it while adding new categories — infrastructure management, security patching, model version control.
Open weights also doesn't solve the capability gap risk in reverse. When a vendor deprecates a model, you have a migration problem. When an open source model community moves on, you have a maintenance problem — you're now responsible for a model version that nobody else is running in production, which means you're also responsible for finding and fixing the security issues nobody else will find for you.
The point is not that open weights is a bad choice. It's that vendor pricing risk and open weights hosting risk are different problems that require different mitigation strategies — and conflating them means you're not actually solving either.
Three questions enterprise buyers should add to their vendor assessment
Standard vendor assessments are well-designed for the risks they were built to catch. We'd add three questions that most of them currently miss:
First: what is the vendor's current cost of compute relative to its published API pricing, and how does that ratio trend as scale increases?
This is a hard question to get a direct answer to, and you probably won't. But how the vendor responds to the question is diagnostic. A vendor that can explain, in general terms, how their infrastructure economics work and what their path to sustainable pricing looks like is a different risk profile than a vendor who deflects or doesn't understand the question.
Second: what has this vendor's pricing history looked like over the past 24 months, and specifically, how have they handled model deprecation and capability tier changes?
This is answerable. Check the vendor's changelog, their pricing page history (Wayback Machine is useful for this), and the community discussion in their developer forums and Reddit. A pattern of frequent tier restructuring or short-notice model deprecation is a leading indicator of how they'll handle future capital pressure.
Third: if this vendor's pricing increased by 3× tomorrow, what would our migration path look like and what would it cost?
This is the question most procurement teams don't want to answer because the answer is usually "we don't know, and it would be expensive." That's exactly why it should be asked before signing a multi-year commitment rather than after a renewal surprise. The answer should include: which workflows depend on vendor-specific features that have no equivalent elsewhere, what the re-engineering cost would be, and what the timeline would be to migrate critical processes.
If the answer is "migration would take twelve months and cost more than two years of vendor fees" — that's a negotiating position, not a vendor selection outcome. It should affect the contract terms you accept.
What contract terms actually protect you
The vendor assessment conversation is about knowing what you're getting into. The contract terms conversation is about what you can do when the situation changes.
Three clauses worth negotiating explicitly in AI vendor contracts, particularly for multi-year commitments:
Pricing stability provisions. The contract should specify what conditions allow the vendor to reprice during the term, and what notice period is required. "Pricing may change with 30 days notice" is not a stability provision. "Pricing is fixed for the contract term for the capabilities specified in Exhibit A" is.
Model continuity provisions. If your workflow depends on a specific model or capability, that dependency should be named and the vendor should be required to provide either continuity or a defined migration support package if the model is deprecated. What constitutes "support" should be specific — not "we'll help you migrate" but "we'll provide [X] hours of migration engineering at no additional cost within [Y] timeline."
Exit ramp provisions. The contract should specify your right to terminate and what data portability and migration assistance looks like if you do. In a well-negotiated contract, this is a mutual risk management provision — the vendor has more certainty about your commitment, and you have more certainty about your options if the relationship changes.
These are negotiable in most enterprise AI contracts. They are not offered by default. You have to ask for them, and you have more leverage to ask before you sign than after.
The procurement frame that's missing
There's a pattern we've seen across the enterprise AI deals we've been part of or adjacent to: procurement teams are applying a software vendor framework to an infrastructure vendor dynamic, and the two have different risk profiles.
When you buy enterprise SaaS, the vendor's economics are relatively stable — the software is built, the margin is predictable, the pricing reflects a reasonably clear cost structure. The main risk is product direction and support quality.
When you buy API access to a model that's being priced below cost on a capital-markets timeline, you're in a different relationship. The vendor's incentive is adoption now and monetization later. That transition from "adoption" to "monetization" phase is the moment your renewal looks different from your initial contract.
The companies that are best positioned when that transition happens are the ones who saw it coming and structured their contracts accordingly — or who made architecture choices that gave them genuine optionality when the pricing environment changed.
That requires thinking about vendor risk earlier in the procurement process than most teams currently do.
One thing we might be wrong about
The scenario we're describing — AI vendors moving toward cost-reflective pricing as capital market conditions change — is directionally likely but not inevitable on any particular timeline. AI infrastructure costs are falling. It's possible that the gap between current pricing and sustainable pricing closes through cost reduction rather than price increase. In that case, the renewal risk we're describing doesn't materialize.
We're also working from aggregate data — the Nikkei figures are estimates, and the actual capital structure of any specific vendor is more complex than a headline number. Some vendors in this category are better capitalized than others. Some have revenue trajectories that make the math more sustainable.
The argument isn't "all AI vendors will have a pricing crisis." The argument is that vendor pricing stability — which has been assumed rather than examined in most enterprise AI procurement — is worth examining explicitly, particularly for multi-year commitments in business-critical workflows.
The downside of treating this risk seriously is that you spend some legal time on contract provisions that turn out to be unnecessary. The downside of not treating it seriously is that you find out three years into a deployment that the pricing your ROI case was built on has changed in ways you have no contractual recourse for.
Working notes from B2B AI deployment in North America. Part of an ongoing series on what we keep noticing across wildly different industries — and what the industry isn't ready to say out loud.
If you're currently evaluating enterprise AI vendors or structuring a multi-year AI commitment, this is one of the conversations we have with clients before they sign — alongside the data readiness, integration scope, and total cost modeling that determines whether the economics hold. [A 1-on-1 strategy call is the starting point.](https://zenaicorp.com/en)
Like
5 Comments
5 Comments
-
1
I'm curious what convinced you vendor pricing stability is the strategic risk buyers are underestimating rather than technical lock-in itself.
From the conversations you've had with enterprise teams, do procurement leaders usually recognize this capital-structure risk once it's explained, or do they still view AI vendor selection primarily as a technical evaluation?
-
1
Technical lock-in is the risk most teams can already see — it shows up in migration estimates and architecture reviews. Capital-structure risk is harder because it looks fine right now. In our experience, procurement leaders usually get it quickly once you frame it as "what does your renewal look like if this vendor's investors start demanding returns?" — that question lands. The ones who don't engage with it tend to be earlier in the process, still in the "will this work technically" phase. The two concerns aren't sequential, but that's often how the conversation is structured.
-
1
Thanks for taking the time to explain your thinking. I'd enjoy continuing the conversation outside the thread if you're open to it. What's the best email to reach you on?
-
1
You can contact this email address: zenai.intl@gmail.com
-
1
Thanks! I’ve just sent it over.
Looking forward to hearing your thoughts whenever you have a chance.
-
-
-
-
Enterprise AI agents can now read and write your files directly. The question nobody is asking: what happens after they do?
---
A client walked into our quarterly review with visible excitement.
They'd just rolled out an AI agent with direct access to their document management system — contracts, templates, compliance checklists. The agent had spent the weekend scanning 500+ master service agreement templates and updating the liability clause language to reflect a new regulatory requirement.
"It finished in four hours. Would have taken our legal ops team six weeks."
Our first question was not "how accurate was it?"
Our first question was: "After it finished, how do you audit what it changed?"
The room got quiet in a familiar way.
---
The question they thought we were asking
They thought we were asking about accuracy. Hallucinations. Whether the agent had correctly identified the relevant clause and applied the right language.
That's a reasonable concern. It's also the wrong concern — or at least, the second concern.
The first concern is simpler and more structural: when an AI agent makes changes at scale, the organization's ability to review, catch errors, and reverse course does not scale with it.
A legal ops team updating 500 contracts over six weeks produces 500 discrete moments where a human reads what was changed. Slow, expensive, human. But each moment is a checkpoint. An error on contract #47 gets caught before contract #48 is touched.
An agent updating 500 contracts over four hours produces one moment: the moment it finishes. At that point, whatever it did is done, across all 500 files simultaneously, and your review window is already upstream of the output.
The accuracy question matters. The reversibility question comes first.
---
What "done" looks like when the agent writes files
We've written before about done-state specification — the problem of agents that loop because nobody defined what completion looks like. This is a related but different failure mode.
In the loop problem, the agent doesn't stop. In the reversibility problem, the agent stops exactly when it should — it completes the task correctly, by every technical measure — and the problem is that you can't easily undo what it did.
These are not the same thing. And most enterprise teams who are currently excited about AI agents with file access are thinking about the first problem while the second one sits unexamined.
The specific capability that's newly available — agents reading and writing Word documents, Excel files, contract databases directly — is genuinely useful. It's also the capability that collapses the review window the fastest. Because unlike an agent sending a recommendation to a human for approval, an agent writing a file has already acted. The recommendation is the action.
---
The reversibility matrix
Not all agent operations carry the same reversal cost. In deployments where we've scoped agent access, we've started using a simple framework to classify operations before they go live.
Tier 1 — reversible by default. Reading data, generating draft documents, producing summaries, flagging items for human review. These operations produce outputs that don't change the underlying record. The agent can run at full autonomy. If it gets something wrong, the human who reads the draft catches it before it becomes real.
Tier 2 — reversible with effort. Overwriting files with version history enabled, updating records in a system with a full audit log, appending to a document rather than replacing it. The agent's action is real, but there's a documented path back. These operations need a confirmation window — not a full human review of every change, but a defined process for catching systematic errors before they propagate. The 500-contract update would live here if version control was active.
Tier 3 — effectively irreversible. Sending emails. Submitting forms to external systems. Overwriting files without version history. Triggering downstream workflows in other systems. Deleting records. These operations have consequences outside the organization's control perimeter. Once the email is sent, the email is sent. The agent cannot unsend it. You cannot unsend it.
The principle is straightforward: autonomous agent authority should be proportional to how reversible the operation is.
Tier 1 operations: full autonomy is fine. Tier 2: autonomous with mandatory audit trail and a human checkpoint before scale. Tier 3: human-required confirmation, no exceptions.
The failure mode we keep seeing: organizations grant Tier 3 authority because the agent has demonstrated accuracy on Tier 1 tasks. Accuracy and reversibility are independent variables. An agent can be highly accurate and still perform irreversible operations that turn out to be wrong — because the error wasn't in the agent's logic, it was in the instructions the agent received.
---
The contract case: what the audit found
Back to the 500 contracts.
We spent a week with that client after the quarterly review. The news was not bad — the agent had been largely accurate. The new liability clause was correctly identified and correctly updated in 91% of files.
The problem was the 9%.
Of the roughly 45 contracts where the update was wrong, most were wrong in a specific way: the agent had encountered clause variants that looked like the target language but were in fact negotiated exceptions — custom language a particular counterparty had insisted on in the original deal. The agent correctly identified them as "liability clause" and correctly applied the standard language.
That was the error. Not a hallucination. Not a failure of understanding. The agent did exactly what it was told. The instructions didn't account for the cases where the "standard" clause had been deliberately made non-standard.
Now: can you find those 45 contracts in a pool of 500? Yes, if you have version history. With effort, over time, as individual counterparties notice the change and flag it.
Or in the next contract renewal cycle, when the wrong clause language creates a dispute.
This is the specific cost of irreversible operations at scale: errors don't surface immediately. They surface at the worst possible moment — when the contract is in dispute, when the audit happens, when the counterparty's lawyer is already on the phone.
---
"Undo" is an organizational problem, not a technical one
The technical undo is usually available. Version history in SharePoint, Git commits, database snapshots, audit logs. The infrastructure for reversal often exists.
What often doesn't exist is the organizational process for using it.
When something goes wrong at scale — when the agent updated 500 files and 45 of them need reverting — who makes that call? Who identifies which 45? Who has the authority to initiate the rollback? Who notifies the counterparties? Who updates the record of what happened and why?
In a human process, errors are caught at the individual level. The legal ops analyst who updates a contract catches her own mistake, fixes it, moves on. The error never becomes an incident.
In an agent process running at scale, errors are systemic. They affect a class of records simultaneously. Fixing them requires a process that most organizations have never had to build — because they've never had a single actor make changes across 500 files in four hours before.
We now include what we call a reversal protocol in every agent deployment that touches production data. Before the agent goes live, we document three things: who has authority to initiate a rollback, what the rollback procedure is for each category of operation, and what constitutes a threshold for triggering review. Not "if the agent makes an error" — every agent will make errors. "What error rate, in what category of operation, requires a human to stop and assess before the agent continues?"
That document exists before the first run. Not because we expect catastrophic failure. Because the cost of writing it before is thirty minutes, and the cost of not having it after a systematic error is measured in weeks.
---
What this means if you're deploying agents with file access
The capability is real and the excitement is warranted. An agent that can read and write enterprise documents directly compresses weeks of process time into hours. We are not arguing against it.
We are arguing for sequencing.
First, classify your operations before you grant access. Map every action the agent will take against the reversibility matrix. Not every operation the agent could theoretically take — every operation this specific agent, in this specific workflow, will actually perform. Tier 1 and Tier 2 operations can run at scale from day one. Tier 3 operations require a defined confirmation process before they go live.
Second, run at scope before you run at scale. The first deployment of any agent with write access should cover a bounded set of records — not 500 contracts, but 20. Enough to validate the logic, surface the edge cases, and find the negotiated exceptions the agent wasn't told to handle. Scale is not a reward for accuracy on the first run. It's a second deployment decision that requires its own review.
Third, build the reversal protocol before you build the agent. The question "how do we undo this if it's wrong" needs to be answered before the agent runs for the first time, not after the first systematic error. The answer should include who makes the call, what the procedure is, and what the notification path looks like. If you can't answer those questions before deployment, you're not ready to deploy at Tier 2 or Tier 3.
---
The sentence nobody wants to say in the kickoff meeting
Every agent deployment kickoff we've been in follows the same pattern: the team is excited, the use case is clear, the timeline is aggressive.
The question that changes the energy is the one about failure.
Not "what if the agent gets something wrong" — teams have usually thought about that. The specific question: "If the agent runs tonight and makes a systematic error across your full document set, what does your organization do tomorrow morning?"
The teams that have a clean answer to that question are ready to deploy. The teams that pause and look at each other — they're not. Not because their agent won't work, but because their organization isn't yet designed for what happens when it does.
The agent reading and writing your files is a real capability shift. It's also the first time in most organizations' history that a single non-human actor has had that kind of access, at that kind of speed, without a human in the approval loop.
The technical question — can the agent do this accurately — is answerable with a pilot.
The organizational question — what do we do when it doesn't — needs to be answered before the pilot runs.
---
One thing we might be wrong about
The reversibility matrix we've described assumes a relatively clear line between read operations and write operations. In practice, that line is getting blurrier.
An agent that reads a file and generates a "recommendation" that auto-populates into a tracked-changes version of the same document is doing something between Tier 1 and Tier 2. An agent that drafts an email and places it in a "pending send" queue is sitting between Tier 2 and Tier 3 — the human is nominally in the loop, but if the review queue is 200 items long, the nominal review is not a real one.
The principle we're confident in: autonomous authority should scale with reversibility, not with accuracy. An agent that is 95% accurate on Tier 3 operations is not ready for full autonomy on Tier 3 operations. Accuracy and reversibility are independent. The agent needs both.
Whether organizations can consistently build and maintain that distinction as agents become more capable — that's an open question. The answer will probably vary by industry, by regulatory context, and by how much the organization has invested in the reversal infrastructure before the first systematic error arrives.
We'd rather they invest before.
---
Working notes from B2B AI deployment in North America. Part of an ongoing series on what we keep noticing across wildly different industries — and what the industry isn't ready to say out loud.
Like
Comment
GLM 5.2 just made the token price war official. But the enterprise buyer celebrating cheaper models may be solving for the wrong number.
---
A client called us last week with a screenshot.
GLM 5.2 pricing. Token costs that would have looked like a typo two years ago.
"We're renewing in six months. I'm sending this to you now so we can talk about repricing."
It's a reasonable move. If the underlying model costs drop by 70%, shouldn't your AI vendor fee follow?
Here's the problem with that logic: it assumes the thing you're paying for is the model.
---
What the token price war actually affects
The GLM 5.2 announcement is real. So is the broader trend — Claude, GPT, Gemini, and now domestic Chinese models are in a race where the floor keeps dropping. Model pricing is commoditizing faster than anyone predicted three years ago. The Jevons paradox is in full effect: as models get cheaper, enterprises use more of them, and total AI spend goes up even as per-token cost collapses.
But here's what we've tracked across our deployments: compute costs — the thing that's actually getting cheaper — averaged 11% of total project spend. The range was 7% to 16%.
The other 89% doesn't appear on any model provider invoice.
It distributes like this: data engineering — cleaning, normalization, deduplication, pipeline construction — runs 35 to 45% of total cost. Permission and integration work — security reviews, API connectors, auth flows, the vendor assessments that no one scoped — runs 25 to 35%. Ongoing maintenance and tuning once the system is live: 15 to 20%.
GLM 5.2 getting 70% cheaper is a real number. Seventy percent of eleven percent is about 7.7 percentage points off your total project cost. That's not nothing. It's also not the number your CFO thinks it is when they see that screenshot.
---
Model commoditization ≠ deployment commoditization
The buyer's intuition is: AI models are getting cheaper, therefore AI is getting cheaper.
That's only true if the model is the bottleneck. In almost every non-AI-native enterprise we've worked with — auto retail, medical devices, manufacturing, real estate — the model was never the bottleneck.
The bottleneck was organizational. Data that lived in seven systems with three different naming conventions for the same customer. Security reviews that had a six-week queue. The senior analyst who was the only person who knew what "production variance" actually meant in context — and whether it was measured from order confirmation, shipping, or physical run completion.
We wrote a piece earlier in this series about a manufacturing client who wanted to automate variance reporting. We spent two weeks in their systems before touching a model. We found three different definitions of the core metric across three different systems. The twelve-hour human process we'd been asked to automate existed, in large part, to reconcile those definitions every week.
There was no "production variance" to automate. There was a twelve-hour human arbitration process wearing the costume of a reporting task.
GLM 5.2 doesn't change that. Neither does any model price drop. The data still needs cleaning. The definitions still need reconciling. The organizational process still needs mapping before any model can do anything useful with it.
---
What actually happens when models get cheaper
Here's the counterintuitive outcome of the token price war: cheaper models don't simplify enterprise AI. They expand it.
When compute costs drop, the rational enterprise response is to use more models, in more places, for more decisions. A company that was running one AI process in 2024 might be running five by 2026 — because the marginal cost of adding a new one has fallen to nearly zero.
But each new model deployment carries the same organizational overhead that the first one did. You still need data pipelines. You still need integration work. You still need the process audit that tells you which decisions in the workflow are actually deterministic and which are judgment calls wearing the costume of rules. You still need the accountability structure — the document that answers the medical director's question: "When the AI got that routing decision wrong, who authorized it to make that call?"
More models, same per-model setup cost, larger total footprint. The integration complexity scales with the number of systems, not with the per-token price.
This is the part of the conversation that doesn't fit on a screenshot.
---
The real negotiation
When our client sends us a GLM price comparison as a repricing signal, they're opening a negotiation. But they're negotiating on the wrong line item.
What they're actually buying from us isn't model access. They could buy model access directly — and increasingly they do, for the raw inference layer. What they're buying is the work that makes the model usable inside their actual organization: the data readiness work, the integration development, the process audit before anyone touches a model, the done-state specification that prevents a reporting agent from generating 847 versions of the same report over a weekend because nobody defined what "complete" looked like.
We've written about all of these separately. The common thread: none of them get cheaper when the model gets cheaper.
A 40% reduction in API costs on 11% of total spend is a 4.4% improvement in total project economics. That's the actual math behind the screenshot. It's real. It's also a different conversation than repricing the engagement.
---
Who actually benefits from the margin collapse
The AI margin collapse is genuinely bad news for companies whose revenue is tied to model access — the API wrappers, the thin-layer aggregators, the companies who were charging a premium for access to something that's now a commodity.
It's not automatically good news for enterprise buyers, because their real costs were never primarily in the model layer.
And it's quite good news for exactly one category of vendor: the ones who never charged for the model to begin with.
The shops whose value is in the deployment layer — the data engineering, the integration work, the process understanding, the organizational navigation that makes AI actually run inside a real company — those shops are largely immune to model commoditization. The client asking us to reprice based on GLM 5.2 is applying pressure to the line item that represents 11% of our total value. The other 89% isn't moving.
This was always true. The margin collapse just makes it visible.
---
The sentence that matters more than the screenshot
We have a diagnostic we run before every engagement: "If we close this deal, what one sentence does the buyer say to their board to justify the spend?"
The clients who send us token price comparisons are usually trying to justify a renewal to a CFO who's been reading about the AI cost collapse. The sentence they need to say to their board isn't "we renegotiated our API costs." It's something about what the system is actually doing — the cost line it's eliminating, the process it's replaced, the decision accuracy that's improved.
The model price is the easiest number in the room. It's on a vendor's website. It's in a press release. It benchmarks cleanly.
The harder number — the one that actually determines whether the project paid for itself — is the one nobody modeled before the project started. The fully-loaded cost of the analyst who spent 40% of her time being a data pipeline. The calendar time absorbed by IT security review. The ongoing tuning work that everyone assumed would be zero and always turns out to be 15% of the first year.
Cheaper models are a real tailwind. They're just not the wind the CFO screenshot is pointing at.
---
One thing we might be wrong about
The cost ratios we've described — 11% compute, 89% everything else — come from deployments in traditional industries with legacy data infrastructure. It's possible that as tooling matures — better data connectors, faster vendor assessments, standardized integration layers — the non-compute costs will compress. We'd be genuinely happy to be wrong about this. The structure we're describing isn't inherent to AI. It's the current shape of the problem in the industries we work in.
It's also possible that the model commoditization effect reshapes the market in ways that benefit enterprise buyers differently than we're predicting. If cheaper models accelerate the adoption of AI-native data infrastructure — if companies start building with data readiness as a first-class concern because they know they'll be running multiple models — the 89% might shrink.
But that's a three-to-five year story. The CFO with a GLM screenshot is asking about this renewal cycle. And for this renewal cycle, the number that moved is the one that represents about a dime on the dollar of what they're actually spending.
---
Working notes from B2B AI deployment in North America. Part of an ongoing series on what we keep noticing across wildly different industries — and what the industry isn't ready to say out loud.
Like
1 Comment
1 Comment
-
1
The cost breakdown is a useful perspective, but I'd keep validating what enterprise buyers are actually hiring AI vendors to do. If they believe they're buying model access, cheaper models become a pricing discussion. If they're buying organizational capability—making AI work reliably inside a complex business—then the model is just one input, not the product.
Why CRM and ERP AI projects break after the API connection succeeds.
The demo looked good.
The CRM was connected.
The ERP API was authenticated.
The AI could summarize an account, pull inventory data, and draft a reply.
Then a real customer request arrived.
The sales record showed an open opportunity and a negotiated price. The ERP showed the account was on credit hold. A support ticket revealed an unresolved delivery issue. An email thread contained a promise that never made it into CRM.
The AI had access to information.
What it did not have was a rule for deciding what mattered most.
That is where many CRM and ERP AI projects fail.
Not at the connection.
After the connection.
Most companies already have the systems they need. Sales has a CRM. Finance and operations rely on ERP. Support has a ticketing platform. Important context still lives in inboxes, PDFs, call notes, shared drives, and spreadsheets that nobody has had time to retire.
The problem appears when one customer request crosses all of them.
A customer asks for an updated quote.
A sales rep needs inventory.
Operations needs to know whether the order can ship.
Finance needs to confirm payment status.
Support may need to flag an unresolved issue.
That is not a CRM question or an ERP question.
It is one business decision spread across several systems.
Integration Creates a Path. A Workflow Creates a Decision.
Companies often evaluate AI integration by asking whether the systems can connect.
Can AI read Salesforce?
Can it pull data from NetSuite?
Can it update HubSpot?
Can it search a custom database?
Those questions matter. But they are only the entry point.
A connector creates a path for data to move.
A workflow decides what should happen when that data arrives.
That distinction becomes critical once AI can do more than summarize notes.
When AI can create tasks, draft quotes, route work, update customer records, or trigger downstream actions, the company needs operating rules.
Which system is trusted when records conflict?
What can the AI read?
What can it recommend?
What can it write back?
Which actions require approval?
What happens when a record cannot be matched?
Who handles the exception after launch?
Those are not technical details to revisit later. They are the business logic of the system.
The “Source of Truth” Is Usually More Complicated Than One System
Most teams want a simple answer to a difficult question:
Which system is the source of truth?
In practice, the answer is often different for different decisions.
Your CRM may be the source of truth for account ownership and sales activity.
Your ERP may be the source of truth for inventory, payment status, orders, and financial controls.
A contract repository may be the source of truth for pricing exceptions.
A ticketing system may be the source of truth for unresolved service issues.
Trying to force everything into one “master” system can create its own problems. A more useful approach is to define the trusted source for each important decision.
For example:
Business decision
Likely trusted source
Who owns the account?
CRM
Can this order ship?
ERP or operations system
Is there a special contract price?
Approved contract repository
Is the account on hold?
ERP or finance system
Is there an unresolved service issue?
Support platform
Can AI send the response?
Human approval rule
The AI does not need unrestricted access to every system.
It needs the right access for the task, plus clear rules for what happens when two records tell different stories.
That is the difference between “AI knows a lot” and “AI can be trusted inside a workflow.”
The Useful Role for AI Is Often Smaller Than People Expect
The most valuable CRM and ERP workflows do not usually begin with autonomous agents making major decisions.
They begin with the work people do before a decision is made.
Reading incoming emails.
Pulling context from several systems.
Extracting fields from documents.
Finding missing information.
Flagging contradictions.
Preparing a draft.
Routing the case to the right person.
Take a sales inquiry.
A prospect submits a form after business hours. The message includes a product need, a deadline, and a rough budget.
A well-designed workflow can:
identify the company and contact;
check whether the account already exists;
detect duplicate records;
retrieve relevant account history;
confirm whether the product is available;
identify missing information;
assign the request to the right owner;
create a follow-up task;
prepare a response draft;
flag the inquiry if pricing, credit status, or an existing customer issue requires review.
That is already a meaningful improvement.
The sales team responds faster. The CRM is cleaner. High-value leads do not disappear in an inbox. Employees spend less time switching between systems.
None of that requires AI to silently approve a discount, alter credit status, or promise delivery dates.
The boundary is the point.
The Real Design Question: What Is AI Allowed to Do?
A useful CRM and ERP AI workflow needs four levels of permission.
1. What can AI see?
The answer should be limited by the task.
A sales-assist workflow may need customer history, inventory, open opportunities, and approved product information.
It does not need unrestricted access to payroll data, internal legal documents, or every finance record in the company.
Access should follow the workflow, not convenience.
2. What can AI interpret?
AI can often help classify requests, extract data from forms, summarize account history, and identify missing or inconsistent information.
This is where document-heavy operations can gain value quickly.
Invoices, purchase orders, packing lists, claims, customer forms, inspection reports, and technical files all create repetitive work before someone can make a decision.
AI can reduce the time spent on standard cases.
The important question is what happens to the unusual ones.
3. What can AI recommend?
Recommendations are often safer than automatic actions.
AI can suggest the next owner, prepare a response draft, flag a risk, recommend an exception path, or identify records that need validation.
A person remains responsible for the judgment.
This model works especially well in workflows involving pricing, finance, customer commitments, compliance, or sensitive customer issues.
4. What can AI change?
This is where companies need to be deliberate.
Creating a follow-up task may be low risk.
Updating a non-sensitive structured field may be acceptable under clear rules.
Changing a price, modifying an account status, approving a refund, or updating a critical ERP record is different.
Those actions often need approval, logging, and a clear escalation path.
The NIST AI Risk Management Framework is useful because it treats trustworthiness and risk management as part of system design, development, use, and evaluation—not as a checklist added after deployment.
Exceptions Are Not Edge Cases. They Are the Workflow.
This is one of the most expensive misunderstandings in enterprise automation.
Teams often design for the normal path:
A request comes in.
Data is retrieved.
AI prepares a response.
A system is updated.
But real business work is shaped by exceptions.
The customer record does not match.
The shipping address differs between CRM and ERP.
A document is incomplete.
An API is unavailable.
The AI has low confidence.
The customer has a special arrangement that exists only in a contract attachment.
The service issue has not been resolved.
The customer request falls outside normal policy.
These are not failures of the workflow.
They are the conditions the workflow was supposed to handle.
A production-ready system needs an answer to each one.
Should the case pause?
Should it route to finance, operations, sales, support, or a manager?
Should the AI produce a draft and wait?
Should it be logged for review?
Should the system prevent a write-back action until a person approves it?
The quality of an AI workflow is often visible in what it does when it cannot confidently continue.
Why Demos Hide the Hard Part
A demo usually has clean inputs.
The sample customer record is complete.
The ERP API returns the expected result.
The document is readable.
The workflow has one obvious path.
Real operations are less cooperative.
Data changes. Teams reorganize. Permissions shift. Fields are added. APIs break. A new exception appears that nobody anticipated.
That does not mean AI projects are doomed.
It means the project cannot end when the demo works.
After launch, someone needs to own:
monitoring failed actions;
reviewing exceptions;
updating business rules;
handling changing APIs;
managing permission changes;
checking whether the workflow is improving the intended metric;
deciding which actions are safe to automate next.
A partner that only builds the first version has solved part of the problem.
A partner that helps define the operating model has solved the part that determines whether the system survives everyday use.
Keep the Core System. Improve the Work Around It.
Many companies assume AI adoption requires replacing the CRM, ERP, DMS, TMS, or internal platform that people already depend on.
Often, that is unnecessary.
A more practical approach is to add an intelligent layer around the existing environment.
That layer can:
read approved data;
interpret incoming emails, calls, and documents;
gather context from multiple systems;
identify missing or conflicting information;
create tasks;
prepare drafts;
route exceptions;
write back through controlled steps;
maintain an audit trail for important actions.
The core system stays in place.
The workflow around it becomes easier to run.
This approach is often more realistic for businesses with customized software, old systems, limited APIs, or years of operational knowledge built into existing processes.
The question is not whether the technology stack looks modern.
The question is whether people can do their work with less friction and more control.
A Better Starting Point Than “Integrate AI With Everything”
The best first project is rarely “connect AI to the entire business.”
It is usually one workflow that already has a measurable cost.
A missed lead.
A slow quote process.
A document queue.
A customer-service handoff.
A reconciliation workflow built around spreadsheets.
An internal system people avoid using because it is too slow or too fragmented.
Start with the process map.
What triggers the work?
Which systems are involved?
Who touches the process?
Where do people lose time?
What data is trusted?
Which decisions need approval?
What happens when something goes wrong?
Then choose one outcome to measure.
Maybe it is first-response time.
Maybe it is document-processing time.
Maybe it is the number of unresolved exceptions.
Maybe it is lead-to-meeting conversion.
Maybe it is the amount of time a service team spends searching across systems.
The goal is not to make AI visible everywhere.
The goal is to make one important piece of work run better.
Before You Hire an AI Integration Partner
Ask these questions before you sign anything:
How will you determine the source of truth when CRM and ERP data conflict?
What can AI read, recommend, create, and update?
Which actions require human approval?
How will the system handle incomplete records, unusual requests, and API failures?
How will you test the workflow against real business cases?
What data should remain outside the first phase?
Who owns monitoring and exception handling after launch?
Which business metric will tell us whether the workflow was worth expanding?
A provider that only talks about model names, agent frameworks, or connectors is not necessarily wrong.
They may simply be answering an easier question.
The harder question is whether the workflow will still work when the customer is waiting, the data is messy, and the normal path breaks.
The Best CRM and ERP AI Project Usually Starts Small
Most companies do not need a complete AI transformation plan before they begin.
They need one workflow that is ready to improve.
If your team is debating several ideas, bring three things into the discussion:
a rough process map;
the systems involved;
one current number that shows the cost of the problem.
That is enough to start a serious conversation.
ZenAI helps companies pressure-test CRM and ERP AI workflows before they commit to a large build. The goal is to identify what belongs in the first phase, what should stay outside it, and what it will take to make the workflow reliable in production.
Like
Comment
About
Performance-tuned AI for retail automotive. Inbound service and sales receptionists, outbound BDC agents, and a dashboard that ties it all together.


Comment