Skip to content

The Hidden Cost of AI Hallucinations

Hallucinations are not just wrong text. The hidden cost is confident action on invented state: writes, emails, and tickets without contracts, authority, or audit.

AI Reality Checks #21 · Part of Binary and Beyond. LinkedIn newsletter edition follows.

The ticket says the refund was processed.

Support did not process a refund. Finance has no matching ledger entry. The customer received a polished email that names a policy clause which does not exist.

Someone on the call says the model hallucinated.

That word is doing too much work.

A wrong sentence in a draft is an editing problem. A confident write against invented state is an operations problem. The hidden cost of hallucinations is not embarrassment in a demo. It is action: tickets created, fields updated, emails sent, commitments made, without a contract for what was true, who was allowed to act, or an audit trail that survives Monday morning.

The incident that looked like a language bug

Picture a mid-market SaaS company rolling out an "AI first response" path for billing tickets.

Week one is clean. Drafts sound senior. Average handle time drops. Leadership screenshots the dashboard.

Week five is louder. A VIP is told their renewal credit "already landed." It did not. The model invented a credit note ID that looked plausible, wrote it into the CRM note, and triggered a customer email from the same turn. Support spends two days walking the claim back. Finance opens a reconciliation thread. Legal asks for the policy version that authorised that language and gets a chat log with no tool payloads.

The eval suite still looks fine. Most answers cite something. Fluency never dipped.

What failed was not "the model said something wrong."

What failed was a system that treated invented state as safe enough to mutate production and notify a human outside the company.

That is the cost that lands on support queues, finance recon, and compliance review, not on the model card.

Hallucination is not one failure mode

Teams collapse three different failures into one word:

Invented facts in prose. The model fabricates a clause, a date, or a citation. Painful in a draft. Contained if a human owns the send button.

Invented world state from tools. The model fills gaps when a tool times out, returns empty, or returns stale data, then narrates the gap as settled fact (AI Reality Checks #18, Production Notes #09).

Invented authority. The model acts as if it may commit, refund, or promise, because the tool path was left open and the prompt said "be helpful."

Only the first one is mostly a language problem.

The second and third are integration and governance problems that happen to speak fluently.

When we still call all three "hallucination," we send the fix to prompt engineering and leave the write path alone. The next incident looks identical with a different paragraph.

Where the money actually goes

I keep a short ledger of where hallucination cost shows up after the model is "good enough":

Support absorbs the walk-back. Agents spend cycles correcting commitments the system already emailed. CSAT dips on the cases that looked automated and wrong.

Finance reconciles ghost events. Invented invoice IDs, credits, and "already applied" language create ledger hunts. Nobody budgeted AI for closing packages.

Compliance asks for a trail that does not exist. "What did the model see, which policy version was live, who approved the send?" If the answer is a completion log, you have a demo record, not an audit (AI Reality Checks #20).

Product slows the programme. After one public miss, auto-send becomes draft-only for eighteen months. The "efficiency" slide never updates. The hidden cost is the capability you freeze because you cannot defend it.

Wrong text is cheap when it stays in a draft pane.

Wrong text that becomes state is expensive in the same way any bad write is expensive: recovery, reputation, and the temporary controls that never leave.

Confident action without a contract

The dangerous pattern is not "the model does not know."

It is "the model does not know, and the path still lets it write."

That shows up as:

None of that is fixed by a better temperature setting.

It is fixed by treating model turns that can mutate state like any other integration: contracts for inputs, pending states when truth is incomplete, named authority, idempotent writes, and review matched to blast radius.

A useful way to read the failure

User request
      ↓
Model turn (may invent or compress)
      ↓
Tool / retrieval results (may be empty, stale, or partial)
      ↓
Policy gate (or missing)
      ↓
Mutation or outbound message (or blocked)
      ↓
Audit record (or only fluent text)

If the gate and audit layers are missing, every invented clause becomes eligible for action.

Immature programmes measure hallucination rate on a quiz set.

Mature programmes measure unsafe actions prevented and reconstructable decisions after a miss.

What I ask before "auto" means auto

When a stakeholder says hallucinations are "under control," I ask four questions:

  1. What may the path invent without blocking? Empty tool results should become pending, not prose that sounds finished.
  2. Which mutations require a human name on the approval? Drafts and refunds do not share one lane.
  3. Can you reconstruct inputs, tool payloads, policy version, and actor for last week's miss? Chat text alone is not enough.
  4. Who owns walk-back when the customer already has the email? Support needs a runbook, not a Slack thread titled "AI weirdness."

If those answers are vague, the hallucination rate on the slide is theatre. The production risk is still open.

How this shows up in delivery

When we build AI app delivery for product teams, the first credible milestone is rarely lower inventiveness on a benchmark.

It is a boring one: reads separated from writes, invented IDs blocked before outbound, and an audit shape legal can walk.

On the Arkreach PR analytics build, the hard problem was not generating fluent summaries of coverage. It was making article-level measurement defensible for agencies that would quote those numbers in client reviews. Invented certainty would have been worse than a slow human pass. The architecture had to survive scrutiny, not just a demo day (case study).

That is the same discipline hallucination programmes need. Treat invented state as a data and authority failure before you treat it as a wording failure.

Design responses that do not depend on the model apologising

I do not design for "the model will notice and correct itself."

I design for:

  • Explicit pending states when tools disagree or return empty
  • Read tools isolated from mutating tools in review policy
  • Schema and ID validation before any write or customer-facing claim
  • Idempotency keys on every create path the agent can touch
  • Eval cases for empty results, stale reads, partial writes, and invented identifiers
  • Outbound templates that cannot emit free-form policy language without a cited source record

If your tests only score paragraph quality, you tested the draft pane.

You did not test the cost centre that pays for confident action.

A contrast teams blur

What teams optimiseWhat the hidden cost cares about
Lower quiz hallucination rateFewer unsafe mutations from invented state
Smoother customer toneCommitments grounded in source records
Higher auto-resolveWalk-back runbooks and finance recon load
Prompt polishContracts, gates, and audit after tool failure
"The model said sorry"Reconstructable actor, policy, and payloads

Prompt work is real. It does not replace the right-hand column.

A checklist after "we fixed hallucinations"

Use this when a stakeholder says the language problem is solved:

  1. List every place invented text can become state. CRM notes, tickets, emails, field updates.
  2. Block free-form IDs and policy claims unless a tool returned them.
  3. Separate draft and mutate paths. Auto-draft is not auto-commit.
  4. Name authority for every write (user-delegated, service, or human gate).
  5. Log tool payloads and policy versions, not only model text.
  6. Add evals for empty and partial tool results, not only wrong facts in isolation.
  7. Staff walk-back with a named owner in support and finance.
  8. Measure unsafe actions prevented, not only thumbs on drafts.

If items 2 through 8 are TBD, you improved the prose.

You did not yet bound the cost.

What good looks like six months in

Calm teams treat hallucination as a class of production failure:

  • Invented identifiers never reach customers or ledgers
  • Disagreement between systems becomes pending, not a confident paragraph
  • Mutations share change control with any other write API
  • Dashboards show blocked unsafe actions and audit completeness
  • Incidents produce path changes, not another round of "be more accurate" prompting alone

They do not argue about whether the model is creative enough.

They argue about which invented claims are allowed to become action, and whether the programme budgeted for contracts, gates, and recovery after the demo looked fluent.

Related reading

The honest close

Hallucinations will not disappear because a vendor ships a smarter model.

What can disappear is the habit of letting invented state drive tickets, emails, and ledger-shaped claims without contracts, authority, or audit.

Wrong text in a draft is an editing problem.

Confident action on invented state is a production incident with a nice voice.

The hidden cost is not the sentence. It is the write that believed it.

Agency partner

Need delivery stability without adding headcount?

Quick Brown Fox helps agencies ship complex web platforms, tighten QA, and scale engineering capacity—without becoming a liability to your client relationships.