Skip to content

The Next Bottleneck After LLMs

After models are good enough, the constraint moves to integration, evaluation, data quality, review capacity, and who owns the path.

AI Reality Checks #25 · Part of Binary and Beyond. LinkedIn newsletter edition follows.

For two years the bottleneck was obvious.

Get a better model. Get a longer context. Get a cheaper token.

That race still matters at the frontier.

For most enterprise products, the model crossed "good enough" months ago. The queue moved and many teams did not notice. They are still shopping for intelligence while the organisation chokes on wiring, judgment, and ownership.

The next bottleneck after LLMs is not another model card.

It is everything that makes a fluent answer safe to run in a business.

The pilot that outran the factory

Picture a company that swapped three model vendors in six months.

Each swap improved a bake-off score. None of them fixed duplicate CRM writes, unowned RAG corpora, or a review queue staffed by one exhausted lead. Auto-resolve rate looked fine until VIP exceptions piled up and legal asked for an audit trail that did not exist (AI Reality Checks #20).

They optimised the component that was already past the constraint.

Manufacturing learned this lesson as theory of constraints. Software is rediscovering it with chat UIs.

Where the constraint actually sits

After base quality is adequate, programmes stall on:

Integration. Tools, auth, contracts, retries, partial failure (MCP Is Just the Beginning, Every Integration Is a Distributed System).

Data. Freshness, meaning, ACL, contradiction handling (Context Is the New Database, Garbage In, Garbage Out Is Now a Million-Dollar Problem).

Evaluation. Not vibes. Cases for brownouts, schema drift, empty retrieval, and duplicate side effects.

Human review capacity. Autonomy that depends on unstaffed approval is fiction.

Organisational ownership. Who is paged when the path breaks: AI team, CRM team, vendor, or "the model"?

Change cost. Prompt and policy versioning, corpus versioning, rollback when a "helpful" change ships inventiveness into production (AI Reality Checks #21).

These look less glamorous than model upgrades.

They determine whether the upgrade can be felt by a customer without creating a new incident class.

A better scoreboard

Stop leading with:

  • Arena rankings
  • Tokens per dollar alone
  • Demo wow

Lead with:

  • Mutations with reconstructable audit
  • Actions blocked by missing authority or empty retrieval
  • Review queue SLA and escape rate
  • Tool error budgets
  • Time-to-diagnose when context was wrong

If those numbers are unknown, the bottleneck is visibility, not intelligence.

How this shows up in delivery

In AI app development, the scarce resource on serious programmes is rarely model access. It is senior integration judgment and product owners who will name invariants.

Arkreach needed architecture that made analytics quotable under client scrutiny, not a newer completion endpoint (case study).

A checklist when someone proposes "just upgrade the model"

  1. What failure is currently limited by model quality vs path quality?
  2. Which evals fail today that a larger model would not fix?
  3. What ownership gaps would the upgrade amplify?
  4. Can we reconstruct decisions after the fact?
  5. Is review capacity sized for the new auto-rate?
  6. What is the rollback for corpus and prompt together?

If the honest answer to 2–5 is weak, buy wiring before you buy weights.

Why organisations misread the bottleneck

Model upgrades are vendor-shaped. Someone can buy them.

Path quality is org-shaped. It requires CRM owners, security, legal, and support to agree on invariants. That coordination cost looks like "slow AI adoption." It is often the real work.

Teams also confuse capability with throughput. A better model that triples auto-draft volume without tripling review capacity simply moves the queue to humans who cannot say no fast enough.

If your roadmap is only model swaps, you are optimising the part with a SKU.

Related reading

AI Reality Checks #25 · Part of Binary and Beyond. LinkedIn newsletter edition follows. Unblocking the path after good-enough models? Start a conversation.

Agency partner

Need delivery stability without adding headcount?

Quick Brown Fox helps agencies ship complex web platforms, tighten QA, and scale engineering capacity—without becoming a liability to your client relationships.