Architecture Files #29 · Part of Binary and Beyond. LinkedIn newsletter edition follows.
When an integration fails, the postmortem often starts with a protocol.
Timeouts. Retries. Schema mismatch. "They sent the wrong status." Someone opens Wireshark energy in a Slack thread and the room feels technical again.
Most of the expensive failures I see are not protocol failures. The bytes arrived. The JSON parsed. The systems disagreed about meaning, timing, or responsibility, and nobody had been named to resolve that disagreement before production forced the issue.
Enterprise integrations fail for human reasons: unclear ownership, misaligned incentives, and ambiguity about who gets paged when two correct-looking systems tell different stories.
The sync that worked until it mattered
Picture a mid-market manufacturer connecting ERP, warehouse, and a customer portal.
Orders flow ERP to warehouse. Shipments flow warehouse to portal. Invoices flow back to ERP. Each team owns its box. The project plan has arrows and dates. Go-live is green.
Then a partial shipment lands. Warehouse marks lines shipped. Portal shows the order complete because it keyed off header status. ERP still expects remaining lines. Finance books partial revenue one way. Customer support quotes another. Three dashboards, three truths, one angry account manager.
Engineering finds no bug in the sense of a thrown exception. Each system applied a local rule that made sense inside its walls. The failure was that partial shipment meaning was never owned across the seam, and nobody knew whose pager rang when the meanings collided.
Every Integration Is a Distributed System is the technical frame. This essay is the organisational one: distributed systems need distributed accountability, and most programmes only fund the happy-path diagram.
Protocols are easy to blame because they feel fixable
Choosing REST vs events, sync vs async, shared DB vs API: these debates are concrete. You can time-box them. You can buy a tool. You can put a decision in Confluence.
Ownership debates feel political. Who decides what shipped means? Who pays when a partner needs a migration window? Who is on call when finance and warehouse disagree about an order that both "accepted"?
So teams optimise the protocol and leave the social contract implicit. The implicit contract works until the first ugly case: partials, retries, late cancellations, duplicate webhooks, daylight-saving cutovers, a weekend batch that runs twice.
The Problem with Shared Databases shows one version of this: sharing tables looks like integration because it avoids an explicit API. It also avoids an explicit owner. Everyone can write. Nobody owns the meaning when writes collide.
Why Race Conditions Are Business Problems makes the same point on timing: two correct updates in the wrong order are not a "tech edge case." They are a missing rule about which business outcome wins. Missing rules are human gaps wearing technical symptoms.
The three human gaps that produce "integration debt"
Unclear ownership of the seam. Teams own services. Rarely does anyone own the contract between services as a product: field meanings, failure modes, escalation paths. When the seam is unowned, every incident becomes a negotiation about whose backlog should absorb the fix.
Incentives that punish coordination. Producer teams are rewarded for shipping features in their domain. Consumer teams are rewarded for their SLAs. The person who slows a release to notify partners looks like friction. The person who silently changes a status enum looks like velocity until the war room.
Paging without a decision right. On-call can restart jobs and replay messages. On-call cannot invent a cross-company definition of "fulfilled" at 2am. If the only escalation path is "page whoever wrote the last connector," you will get heroic patches and no durable rule. The next partial shipment will page someone else.
These gaps compound. Each incident teaches individuals local workarounds. The workarounds become load-bearing folklore. New hires inherit rituals instead of contracts. The integration looks mature because it has been alive for years. It is fragile because its rules live in people who might leave.
A mental model: integrations as agreements under stress
An integration is not a pipe. It is an agreement about three questions that only get asked under stress:
- What must stay true when both sides are healthy?
- What is allowed to be temporarily false when one side is slow, down, or replaying?
- Who decides when two healthy sides disagree about the business outcome?
Protocols answer how messages move. Agreements answer what the messages commit you to. Teams that only design the pipe discover the agreement in production, which is the most expensive design review format available.
Data Contracts Matter More Than APIs covers the payload side of that agreement. The human side is staffing the agreement: a named steward for the seam, a change process that matches blast radius, and an escalation that can make a meaning decision without rewriting history in Slack.
How this shows up when you modernise
On legacy application modernisation programmes, the hard inventory is rarely "list the endpoints." It is "list the seams where two teams can both be right and still produce a wrong business outcome," then name who owns each seam after cutover.
Replatforming without re-homing ownership just moves the war rooms to a new runtime. New queues, same undefined status. New gateway, same incentive to ship without notifying the other side.
The useful modernisation artefact is often a seam map: producer, consumer, meaning owner, failure owner, and the ugly cases that have already bitten once. Protocols come after. Otherwise you buy better plumbing for the same ambiguity.
Implications for how you fund and staff integrations
Fund the seam as a product, not as a ticket between two roadmaps.
That means:
- A steward who can say no to silent meaning changes.
- Consumer notification as a release gate for fields that move money, stock, or customer promises.
- Explicit "who pages whom" for disagreement, not only for downtime.
- Budget for migration windows, not only for build sprints.
- Incident reviews that ask which ownership gap the protocol symptom revealed.
Also stop treating "alignment meetings" as optional ceremony. If two systems share a business word, that word needs a definition meeting with teeth: a written outcome, an owner, and a date when the old informal meaning is retired.
You do not need to boil the ocean. Start with the seams that already generate tickets between departments. Those are the human failures announcing themselves.
Tooling will not substitute for that staffing choice. A prettier schema registry, an event bus, or a shared integration platform can make good ownership cheaper to enforce. None of them create ownership. Organisations that buy a platform and skip the steward role usually recreate the same Slack archaeology with better dashboards. The human gaps travel with the architecture unless someone is accountable for closing them.
There is also a hiring tell. If every integration hire is framed as "connector developer" and never as "contract steward" or "seam owner," you will keep getting people skilled at moving payloads and under-skilled at defending meaning. Hire and promote for the agreement, not only for the pipe. Otherwise the next programme will rediscover the same war rooms with a new protocol preference.
A framework for the next integration kickoff
Use this before choosing tools, and again before go-live.
- Name the seam steward. One role accountable for the contract across both sides, not "both teams jointly."
- List the business words that cross the wire. Status, total, available, shipped, paid, cancelled. Define each under partial and failure cases.
- Map incentives. Who is rewarded for shipping a producer change without a consumer window? Make that costly on purpose.
- Write the disagreement path. When two systems accept the same event and compute different outcomes, who decides, in what forum, within what SLA?
- Separate downtime paging from meaning paging. Restarts and replays are ops. Competing truths are product. Do not collapse them into one rotation without authority.
- Inventory the folklore. Ask who still knows the weekend batch quirks. Convert at least the top three into written rules or scheduled deletion.
- Fund the first ugly case in the plan. Partials, duplicates, late cancels. If the plan only shows the happy path, the plan is fiction.
If a kickoff cannot answer these, you are not ready to argue REST vs events. You are ready to argue who owns the argument.
Fingerprint
Failed integrations usually look technical in the ticket title and organisational in the root cause.
Protocols matter. Ownership, incentives, and decision rights matter more, because they decide whether the protocol is defending a shared meaning or laundering a permanent argument between teams.
The integrations that survive contact with partial shipments and month-end are the ones where someone was paid to keep the agreement coherent, not only to keep the messages flowing.
Bytes move for technical reasons.
Integrations hold for human ones: named owners, aligned incentives, and a place to decide when meanings collide.
Related reading
- Every Integration Is a Distributed System
- The Problem with Shared Databases
- Why Race Conditions Are Business Problems
- Data Contracts Matter More Than APIs
Architecture Files #29 · Part of Binary and Beyond. LinkedIn newsletter edition follows. Naming ownership and decision rights across brittle integration seams? Start a conversation.
