What Building AI for Real Businesses Taught Me About the Gap Between AI Hype and Business Value

75 Views

A few months ago, my team pulled the usage logs from the AI translation system we’ve spent the last two years building, expecting them to confirm what the product pitch already said. Instead, we found something the pitch never predicted. The users who mattered most weren’t the ones typing a quick sentence into a box and glancing at the result. They were the ones uploading a full contract or a supplier agreement, then sitting there for thirty or forty minutes, checking every line like it owed them something.

That gap, between the two-minute demo everyone remembers and the forty-minute session nobody talks about, is where I’ve spent most of my career building AI systems for companies that cannot afford a wrong answer. It has taught me more about the real distance between AI hype and AI business value than any roadmap, benchmark, or analyst report ever has.

If you work anywhere near IT, procurement, or operations in a supply chain organization, you’ve probably lived a version of this gap yourself. AI pilots go beautifully in a conference room and come apart in week six of production, once the inputs get messy, the stakes get real, and someone has to sign their name to the output. That’s not usually a failure of the technology. It’s a failure of testing the technology on the wrong question.

The scale of that failure is bigger than most teams admit out loud. According to S&P Global Market Intelligence’s 2025 survey of enterprise AI deployments, 42 percent of companies abandoned most of their AI initiatives last year, up sharply from 17 percent the year before. The tools didn’t get worse in twelve months. What changed is that more companies pushed past the pilot stage, into the part of the process where an AI system has to hold up against a real supplier contract, a customs form, or a production schedule, not a curated test case.

Enterprise AI abandonment nearly tripled between 2024 and 2025 as more pilots hit real production conditions. Source: S&P Global Market Intelligence, 2025.

I saw this pattern up close building translation infrastructure. Most AI translation demos are built around a clean paragraph of marketing copy, and a single large language model will handle that beautifully nine times out of ten. Hand that same model a bill of lading, a supplier safety data sheet, or a technical tolerance spec with mixed units and abbreviations, and the failure mode changes. It doesn’t stumble. It produces something fluent and confident-sounding that is quietly wrong, which is a worse outcome than an obvious error, because nobody catches it until it has already shipped.

Data synthesized from Intento and the WMT24 evaluation benchmarks puts the hallucination rate for individual top-tier large language models on translation tasks at 10 to 18 percent. For a supply chain team, that isn’t an abstract accuracy number. That’s a mistranslated tolerance on a technical drawing, a customs declaration flagged at the border, or a safety instruction that reads correctly in English and means something else once it reaches a warehouse floor in another language.

A single model’s fluent output and a verified answer are not the same thing. Source: Intento and WMT24 evaluation benchmarks; internal consensus benchmarking.

This is where the hype and the real value genuinely diverge. Hype sells the demo: one model, one confident answer, delivered instantly. Value shows up later, in whether that answer holds up once a real document with real consequences runs through it. The clearest evidence for that difference doesn’t come from a benchmark. It comes from looking at how real AI users behave once something is actually at stake, which is the opposite of what a survey will tell you. In our own usage data, the sessions that ran the longest and involved a full document upload, rather than a pasted line of text, consistently belonged to the users with the most to lose from a wrong answer. Nobody had to tell us that. The behavior said it.

That distinction changed how we built the system. Instead of trusting one model’s fluent-sounding guess, the consensus mechanism my team built runs the same input through 22 independent AI models at once and only returns the version most of them agree on. It doesn’t remove judgment from the process. It makes disagreement visible before it becomes someone else’s problem downstream, which is the opposite of what a single confident model does.

As Ofer Tirosh, CEO of Tomedes, a professional translation company, put it when describing this shift: “We’ve evolved beyond pure comparison into active composition, and the system surfaces the most robust translation, not merely the highest-ranked candidate.”

That distinction, between picking a winner and building an answer models actually agree on, is a smaller version of the same choice every business is making right now with AI in general: trust the fastest confident output, or build in a way to check it before it costs you something.

The pattern holds at scale, too. Internal benchmarking on multi-document translation workflows shows consensus-verified output holding terminology and register consistency above 96 percent across large volumes, against an industry baseline closer to 78 percent for single-model output at equivalent scale. For a supply chain business translating the same part names, compliance terms, and safety language across dozens of documents and markets, that consistency gap compounds fast. A term that drifts on page one of a spec sheet and shows up differently on page forty isn’t a stylistic quirk. It’s the kind of inconsistency an auditor or a customs inspector will find.

The lesson I’d give any IT or operations leader evaluating an AI vendor right now, translation or otherwise, is to stop asking vendors to demo their best case and start asking how the system behaves on your worst one. Feed it your messiest document, not your cleanest one. Ask what happens when the model is uncertain, not just what happens when it’s right. And watch what your own team actually does with the tool once the impressive-demo phase is over, because that behavior, not the pitch, is where the real answer about business value lives.

AI hype rewards the two-minute demo. AI value gets built in the unglamorous stretch after it, in the forty-minute session where a real document, a real deadline, and a real cost of being wrong are all sitting on the same screen. That’s the gap I’ve spent my career trying to close, and it’s the same gap every supply chain leader evaluating AI right now would do well to go looking for, before they sign the contract, not after.