AI Reality Checks #22 · Part of Binary and Beyond. LinkedIn newsletter edition follows.
The agent answers from "everything we know about the account."
That phrase used to mean a query against a system of record. Now it often means a bag of chunks: CRM notes, Confluence pages, Slack exports, last week's PDF, and whatever still fits in the window.
Someone in the room calls it context.
I call it a database without admitting it is one.
Context windows and retrieval corpora are becoming the operational store for decisions agents make and actions they take. When you skip ownership, freshness, access control, and meaning, you do not get a clever assistant. You get an ungoverned replica that speaks with confidence while the real systems disagree.
The pilot that treated the corpus as a folder
Picture a product org wiring a support copilot to "the knowledge base."
In practice that means a nightly dump of help articles, a share of CRM note text, and a folder of policy PDFs someone uploaded during the hackathon. Embeddings look green. Retrieval demos cite three passages. Leadership green-lights rollout.
Month two, the quiet failures arrive.
An agent quotes a pricing exception that sales revoked in January. The chunk is still in the index. A junior rep's speculative note about a churn risk becomes "confirmed" in a summary to an exec. A contractor-visible agent retrieves an internal incident write-up that never should have left the security group. Compliance asks who owns the corpus and gets three different answers: AI platform, support ops, "whoever uploaded it."
None of that is a model failure first.
That is a store without a data programme: no owner, no TTL, no ACL, no schema for what a chunk means, and no failure mode when retrieval is stale or empty (AI Reality Checks #17, AI Reality Checks #18).
The cost lands where it always lands. Support walks back wrong guidance. Finance argues about which number was "in context." Legal asks why a restricted document appeared in a customer-facing draft.
Context is a store, not a prompt trick
A useful mental model:
Sources of truth (CRM, warehouse, ticket DB, docs CMS)
↓
Extraction / sync / ACLs
↓
Context store (embeddings, caches, window contents, tool results)
↓
Model turn
↓
Action or answer
Teams obsess over the model turn.
Incidents originate in the middle: what was admitted into the store, under whose authority, with what freshness, and with which meaning attached.
MCP and tool protocols make it cheaper to pull more into that middle layer (AI Reality Checks #19). They do not invent curation. They make missing curation visible faster.
Immature programmes measure tokens in the window.
Mature programmes measure governed, fresh, authorised context per decision.
The five database questions context still needs
I use the same questions I would ask of any operational store. Swap "table" for "corpus and window" and the discipline holds.
1. Ownership
Who is accountable when a bad chunk ships into an answer?
"The AI team" is not an owner for pricing policy. "Support" is not an owner for security incident notes. Without a named steward per corpus class, every incident becomes a meeting about blame (Architecture Files #10).
Shared context without ownership is a shared database with better marketing.
2. Freshness
What is the SLA for invalidation when the source changes?
Nightly rebuilds are a choice, not a default. Pricing, inventory, and entitlement data that can drive action need tighter loops or live reads, not a hopeful index. If the model cites last Tuesday while finance closed a credit today, you have eventual consistency with a confident voice (Production Notes #09).
Name the freshness contract per source. Pending is allowed. Silent stale certainty is not.
3. Access control
Which identity may retrieve which classes of document?
Row-level and document-level rules do not disappear because you embedded the text. If a user cannot open the file in the CMS, the agent should not summarise it into their session. Context stores that flatten ACLs into "everyone on the agent" recreate the worst shared-drive mistakes at generation speed.
4. Schema and meaning
What does this chunk claim to be?
A speculative CRM note, a signed policy, a draft FAQ, and a vendor marketing PDF are not the same type. Without labels for authority, audience, and status, retrieval ranks by similarity and the model narrates them as equal. That is contract drift in prose form (Architecture Files #13).
Treat chunk metadata as schema. Version it. Test it when upstream labels move.
5. Failure modes
What happens when retrieval is empty, conflicting, or timed out?
Empty should become pending or abstain, not invented certainty (AI Reality Checks #21). Conflicting sources should surface disagreement, not a blended paragraph that sounds settled. Timeouts should not fall back to an older cache without marking it.
Databases have nulls, locks, and error codes. Context paths need the equivalent.
A contrast teams blur
| What teams ship | What a context store needs |
|---|---|
| "We indexed the wiki" | Owned corpora with stewards |
| Bigger windows | Freshness SLAs per source class |
| One shared vector DB | ACLs that survive embedding |
| Similarity scores | Authority and status metadata |
| Prompt that says "use context carefully" | Eval cases for stale, empty, and conflict |
Prompts do not replace the right-hand column any more than a comment replaces a foreign key.
How this shows up in delivery
When we build AI app delivery for agencies and product teams, context design shows up before model selection.
On the Arkreach PR analytics build, the hard problem was not stuffing more articles into a window. It was deciding which measurements were authoritative enough for agencies to quote in client reviews, how freshness was defined, and how disagreement between sources stayed visible instead of smoothed into a fluent score. The store had to be defensible, not just large (case study).
That is the same shape every serious RAG programme eventually faces. The corpus is part of the product. Treat it like infrastructure with owners, or it will behave like a junk drawer that can write email.
Window pressure creates silent deletion
Long contexts feel like safety.
They also create a new failure: important facts fall out of the working set while the answer still sounds complete. Summarisation layers compress away hedges. "Might churn" becomes "will churn." Tool results from an earlier turn age inside the session without a timestamp the model respects.
I treat window contents like a cache with eviction:
- Pin authoritative facts that mutations depend on
- Re-fetch entitlement and money fields near the write, do not trust mid-session memory alone
- Keep source IDs on claims that reach customers
- Log what was actually in context for the turn that acted
If you cannot reconstruct the working set, you cannot defend the decision (AI Reality Checks #20).
Governance for the marketplace of chunks
Once multiple teams can publish into the agent corpus, you have a marketplace.
Who may add a source? Who reviews PII classes? Can marketing dump campaign decks next to finance policy in the same index? Can an experiment corpus register beside production without a namespace?
Minimal governance that survives audit:
| Question | Why it matters |
|---|---|
| Who publishes to production context? | Prevents shadow sources |
| What freshness and ACL apply per class? | Embeddings do not inherit CMS rules alone |
| How are versions pinned and rolled back? | Bad chunks travel farther than bad prompts |
| What is logged per retrieval? | "The model read something" is not a record |
| When do we read live vs retrieve? | Money and entitlement often need live |
This is boring work.
It is also the difference between a demo corpus and something legal will allow near a customer.
A checklist after "we added RAG"
Use this when a stakeholder says context is solved:
- Name a steward per corpus class. Pricing, security, support macros, CRM notes.
- Write freshness SLAs and invalidation paths when sources change.
- Enforce ACLs at retrieval, not only at original file share.
- Label authority and status on chunks (draft, approved, speculative, deprecated).
- Separate live reads from retrieved prose for fields that drive money or access.
- Define empty and conflict behaviour before the model speaks.
- Log retrieved IDs and versions on turns that can act.
- Add evals for stale, restricted, and conflicting sources, not only happy citations.
If items 2 through 8 are TBD, you added retrieval.
You did not yet build a store you can operate.
What good looks like six months in
Calm teams treat context like any other operational data plane:
- Corpora are versioned products with owners, not dump folders
- Freshness and ACL breaches page someone named
- Metadata contracts are tested when upstream labels move
- Dashboards show retrieval quality, staleness, and access denials, not only answer thumbs
- Incidents produce corpus and sync changes, not only prompt edits
They do not argue about whether the window is large enough.
They argue about which facts are allowed into the decision, under whose authority, with which failure modes when the store is wrong.
Related reading
- Why AI Needs Better Data Than Humans Do: models need decidable meaning, not tribal notes
- Data Contracts Matter More Than APIs: schema and semantics are load-bearing for retrieval too
- The Problem With Shared Databases: shared context without ownership recreates the same failure
- The Hidden Cost of AI Hallucinations: empty or stale context becomes invented action
The honest close
Context feels soft. Screenshots of chat windows make it look like prompting.
Operationally it is a store: copied, cached, ranked, and sometimes written back through the agent's actions.
Treat it with the same seriousness you treat a database (ownership, freshness, access, meaning, failure), or the agent will keep answering from a junk drawer that sounds like truth.
The window is not memory. It is a working set. Govern it like one.
