Skip to content

Prompt Engineering Won't Save Bad Systems

Better wording cannot fix missing ownership, bad data contracts, unbounded writes, or no recovery path. Systems problems stay systems problems.

AI Reality Checks #24 · Part of Binary and Beyond. LinkedIn newsletter edition follows.

The meeting always finds the same comfort.

If we just improve the prompt, the model will stop doing the dangerous thing.

Sometimes that is partly true. Tone, format, and refusal phrasing respond to wording.

Ownership does not. Idempotency does not. A missing source of truth does not. A tool that can refund without a key does not.

Prompt engineering is real craft.

It will not save a bad system.

The rewrite that felt like progress

Picture a team whose copilot keeps posting duplicate tickets.

They lengthen the system prompt: "Never create a ticket if one already exists. Be careful. Think step by step."

Duplicate rate drops for a week. Then a timeout returns ambiguity, the model tries again, and two tickets appear with slightly different titles. The prompt did its best. The write path had no idempotency key and no "find existing" contract (Production Notes #06, Production Notes #05).

They did not have a prompting problem.

They had an integration problem with a chat box on top.

What prompts can and cannot buy

Prompts can: shape style, require citations, prefer tools in order, refuse certain topics, ask clarifying questions, format JSON more reliably.

Prompts cannot: create an authoritative record that does not exist, enforce row-level auth the tool ignores, prevent a partial write, decide which system wins during disagreement, staff a review queue, or make stale RAG fresh (Context Is the New Database, Why AI Needs Better Data Than Humans Do).

If the failure mode is "the model did something the platform should have made impossible," stop editing adjectives.

Fix the platform.

The false economy

Prompt iteration is cheap and visible.

Architecture work is slower and less demo-friendly.

So organisations burn cycles on prompt theatre while the blast radius stays open: mutating tools overshared, corpora unowned, audits incomplete, evals that only score fluency (AI Reality Checks #20, AI Reality Checks #21).

You can win a Friday bake-off that way.

You cannot win a Monday incident that way.

How this shows up in delivery

In AI app development, we treat prompts as versioned policy adjacent to the path, not as the path.

On Arkreach, better copy in a prompt would not have made PR analytics defensible. Evidence trails and measurement architecture did (case study).

A litmus test

When someone proposes a prompt change as the fix, ask:

  1. Which invariant should the platform enforce even if the model is wrong?
  2. Which write becomes impossible without a key or approval?
  3. Which source is authoritative when tools disagree?
  4. What do we log that a prompt cannot invent later?
  5. What eval fails if this regression returns?

If the answers are all "we will tell the model," you are negotiating with a stochastic component.

You are not designing a system.

A working split of labour

Keep a prompt council for tone, refusal phrasing, and output schemas.

Keep a platform council for auth, keys, contracts, review policy, and eval harnesses.

When an incident lands, ask which council owns the fix. If both point at each other, you have been using prompts as a shared database for accountability.

Ship prompt changes through the same change control as config: version, review, rollback, and an eval that would have caught last week's failure.

That is craft.

It is still not architecture by itself.

Related reading

AI Reality Checks #24 · Part of Binary and Beyond. LinkedIn newsletter edition follows. Fixing the path under the prompt? Start a conversation.

Agency partner

Need delivery stability without adding headcount?

Quick Brown Fox helps agencies ship complex web platforms, tighten QA, and scale engineering capacity—without becoming a liability to your client relationships.