Skip to content

Why AI Agents Need Architecture

Agents without architecture are chat UIs over unscoped tool chains. They need state machines, auth, idempotency, budgets, and degrade paths, the same as any serious integration.

AI Reality Checks #23 · Part of Binary and Beyond. LinkedIn newsletter edition follows.

The deck says agentic.

The demo shows a chat box. The model plans, calls three tools, and posts a summary. Someone says the hard part is the reasoning. Architecture can wait.

Architecture cannot wait.

An agent without architecture is a chat UI over an unscoped tool chain. It can look autonomous in a happy path and still lack the pieces every production integration needs: explicit states, auth boundaries, idempotent writes, budgets, and a plan for when half the path fails.

I have watched teams rename a chatbot "agent," wire MCP servers, and declare the platform era open (AI Reality Checks #19). Week four, duplicate tickets and unbounded spend arrive. The model is blamed. The missing architecture is the actual root cause.

The pilot that stopped at "it can use tools"

Picture an internal ops agent: look up account, search docs, create a ticket, draft a reply.

Session one is magic. An exec asks for a status pack. The agent loops, cites two threads, opens a ticket, and emails a draft. Leadership green-lights broader access.

Week three is quieter and more expensive.

A slow create retries into two tickets. A tool timeout mid-plan leaves CRM updated and email unsent; the next turn invents completion (AI Reality Checks #21). Token spend spikes because the agent re-plans forever on a flaky search server. Compliance asks which identity performed the writes and gets "the agent." Support owns the walk-back. Finance owns the surprise bill. Nobody owns the state machine.

The model did not suddenly get worse.

The organisation discovered it shipped autonomy without the rails autonomy requires.

Autonomy is not a substitute for a control plane

Autonomy means the model may choose among allowed actions toward a goal.

It does not mean the model invents:

  • Which states exist between start and done
  • Which identity it acts as
  • Which writes may retry
  • How much money or how many calls it may burn
  • What happens when tools disagree or die mid-flight

Those belong to architecture. The same architecture you would demand for a payment sync or a fulfilment worker (Production Notes #02, Mental Models #14).

Agents add a stochastic planner in the middle. They do not delete distributed systems.

Immature programmes measure tools connected and tasks completed in the demo.

Mature programmes measure bounded runs: authenticated, idempotent, budgeted, recoverable, auditable.

The stack an agent actually needs

A useful way to read an agent feature:

Goal / user intent
      ↓
Policy + auth (role, scopes, environment)
      ↓
Orchestration (state machine / workflow)
      ↓
Model planner (proposes next step)
      ↓
Tool adapters (contracts, idempotency, timeouts)
      ↓
Business systems
      ↓
Budget + degrade + human gate (matched to risk)
      ↓
Audit (plan, tools, policy version, actor, cost)

Most demos implement the chat UI and the planner.

Most incidents live in every other box.

Trustworthy agent programmes treat the planner as one component inside a control plane, not as the control plane itself (AI Reality Checks #20).

Five architectural surfaces that keep agents boring

1. Explicit state machines

"In progress" is not a state model.

Name the states the business already knows: intake, gathering context, pending approval, mutating, compensating, done, failed. Persist them outside the chat transcript. When the session dies, the run should resume from a named state, not from vibes in the last message (Mental Models #14, Production Notes #12).

If you cannot draw the states on a whiteboard, you do not have an agent product. You have a loop.

2. Auth and authority maps

Who is the agent acting as?

User-delegated identity, narrowly scoped service account, and impersonation are different products. "The bot" is not an identity. Tool servers still need OAuth, row-level rules, and environment separation. Connectors without authority maps turn stochastic routing into privilege escalation with manners (AI Reality Checks #19).

3. Idempotency on every mutation

Agents retry. Networks flake. Planners re-issue the same create after a timeout.

Every ticket create, email send, and field update needs a key and a defined behaviour on duplicate (Production Notes #06, Production Notes #05). Without that, "helpful" becomes "duplicative" at machine speed.

4. Budgets for tokens, time, and side effects

Unbounded agents are cost centres with a chat UI.

Cap model turns, tool calls, wall clock, and mutating actions per run. When the budget trips, degrade: stop, ask a human, or return partial with a pending state. Do not hope the planner will notice spend.

Finance will notice.

5. Degrade paths and human gates

Partial tool success is normal (Production Notes #09, Production Notes #04).

Design the response: pending when systems disagree, read-only mode when write tools are unhealthy, human gate on irreversible acts. Review is not a moral stance. It is a rate limiter on blast radius.

A contrast teams blur

What the deck saysWhat architecture requires
Agentic reasoningNamed workflow states outside the transcript
Tool useScoped auth and mutating vs read tools
Autonomous loopsBudgets and hard stops
"It will figure it out"Idempotent writes and compensation
Chat history as memoryReconstructable audit and durable run state

Prompt polish can improve planning.

It cannot invent the right-hand column.

How this shows up in delivery

When we build AI app delivery for clients who want agents, the first credible milestone is rarely a longer autonomous loop.

It is a short, boring run: authenticated, state-persisted, mutation-keyed, budget-capped, with a degrade path legal can walk.

On the Arkreach PR analytics build, the work that mattered was not "let the model roam." It was an architecture where article-level measurement stayed defensible for agencies quoting numbers in client reviews: clear inputs, owned computation, audit-friendly outputs. The same principle applies when the planner can call tools. Autonomy without rails is not delivery; it is a demo that mails invoices to support (case study).

Complexity does not leave when you add an agent. It moves into orchestration, contracts, and recovery (AI Reality Checks #16).

Context and tools are not architecture either

A large window and a RAG corpus feel like progress (AI Reality Checks #22).

They are inputs. Architecture decides whether those inputs can drive a write, under which state, with which key, against which budget.

I have seen teams invest months in retrieval quality and still ship an agent that double-charges a create path because nobody owned idempotency. The context was fine. The control plane was missing.

A checklist after "we shipped an agent"

Use this when a stakeholder says the remaining work is prompt tuning:

  1. Draw the state machine. Persist it. Resume from it.
  2. Map authority for every tool (user, service, impersonation).
  3. List mutating tools and attach idempotency keys plus owners.
  4. Set budgets for turns, tools, time, and mutations per run.
  5. Define degrade behaviour for timeout, conflict, empty context, and budget trip.
  6. Separate environments. Experiment tools do not register beside prod writes.
  7. Log plan steps, tool payloads, policy version, actor, and cost.
  8. Name on-call for the orchestration path, not "the model vendor."

If items 2 through 8 are TBD, you shipped a chat loop.

You did not ship an agent architecture.

What good looks like six months in

Calm teams treat agents like any other integration surface with a planner attached:

  • Runs are durable workflows with named states
  • Mutations share change control with REST and jobs
  • Budgets and gates are written down and enforced in code
  • Dashboards show safe completions, blocked mutations, spend, and audit completeness
  • Incidents produce architecture changes, not only "be more careful" system prompts

They do not argue about whether the model is agentic enough.

They argue about which production boundaries each run is allowed to cross, and whether the programme budgeted for state, auth, keys, budgets, and degrade paths after the chat demo looked finished.

Related reading

The honest close

Calling software an agent does not exempt it from architecture.

It raises the stakes. A planner that can choose tools will find every missing key, every fuzzy identity, and every unbounded loop faster than a human click path ever did.

Build the control plane first: states, auth, idempotency, budgets, degrade paths.

Then let the model plan inside those rails.

An agent without architecture is not autonomy. It is an unscoped integration with a conversation on top.

Agency partner

Need delivery stability without adding headcount?

Quick Brown Fox helps agencies ship complex web platforms, tighten QA, and scale engineering capacity—without becoming a liability to your client relationships.