Architecture Files #34 · Part of Binary and Beyond. LinkedIn newsletter edition follows.
Teams adopt feature flags to ship without fear.
That sentence is half true and completely incomplete.
Flags do reduce the drama of a Friday deploy. They also decide which customers see which promise, which regions absorb which experiment, and which sales motion gets which pricing path. When you treat that surface as deploy plumbing, you get flag debt: orphaned toggles, unclear owners, and behaviour nobody can explain in a war room.
Feature flags are about business control. Deployment is only one of the jobs they do.
The toggle that "kept the release green"
Picture a mid-market SaaS platform that sells tiered billing and partner white-label.
Engineering introduced a flag service so releases could land dark. The first flags were honest: new_checkout_v2 for internal dogfood, partner_portal_beta for three design partners. Release notes stayed quiet. Incidents dropped. Leadership celebrated continuous delivery.
Eighteen months later the catalogue of flags had three hundred entries. Half still evaluated on every request. Names like temp_fix_pricing, enable_legacy_tax, and rollout_invoice_pdf_v3 lived in production with no owner, no expiry, and no documented audience. Support could not say why two similarly priced accounts saw different invoice layouts. Sales closed a deal on a feature that only existed behind a flag for "enterprise_pilot_q3," which nobody had cleaned up after Q3.
The deploy pipeline was healthy.
The product surface was opaque.
That is the production problem. Flags that begin as release safety become an unowned configuration layer for who gets which product, under which commercial rule, with which kill switch. When something breaks, the first hour is archaeology: not "is the code wrong," but "which combination of flags is this tenant, region, and plan actually running?"
Why deploy-first thinking creates flag debt
A flag invented for CI has a short mental model: on or off until the code is safe, then delete.
A flag that survives contact with product reality has a longer list of meanings:
- Audience. Internal only, design partners, a percentage of traffic, a named account list, a geography, a plan tier.
- Business rule. Who is allowed to see this pricing path, this workflow step, this data export.
- Risk control. What can we disable in sixty seconds if the new path corrupts invoices or spikes support?
- Experiment. What are we measuring, and when does the experiment end?
- Migration. Are we running old and new side by side until cutover is complete?
Those are product and risk questions. They look like boolean environment variables only if you squint.
Teams that treat flags as deploy tech optimise for merging code. They under-invest in ownership, expiry, naming, and the distinction between a temporary rollout switch and a permanent entitlement. The result is familiar: every new feature adds a flag "just in case," nobody owns removal, and the live behaviour of the system becomes a matrix nobody can draw (The Economics of Technical Debt, Why Every Workflow Is Really a State Machine).
The interest rate shows up as:
- Support tickets that cannot be reproduced without cloning a customer's flag state
- Race-like incidents where two paths disagree because both are "live" under different flags (Why Race Conditions Are Business Problems)
- Audit and security questions about who can enable what for which tenant
- Hiring drag: new engineers ship carefully around a minefield of toggles they do not understand
Flags are cheap to create. They are expensive to leave unexplained.
The mental model: control surface, not merge helper
Treat the flag system as a business control surface with four axes.
1. Intent. Is this a rollout, an entitlement, an experiment, a kill switch, or a migration bridge? If you cannot pick one primary intent, you are about to overload a boolean with politics.
2. Subject. Who is evaluated: user, account, plan, region, partner, request attribute? "Percentage of traffic" is a subject. "All enterprise accounts in APAC" is a subject. "Whoever had the cookie last Tuesday" is not a subject. It is an accident.
3. Owner and expiry. Product owns entitlements. Engineering owns rollouts and kill switches for paths they still control. Experiments need a named end date and a decision owner. Migration bridges need a cutover trigger. No owner and no expiry means the flag is already technical debt wearing a friendly name.
4. Failure mode. What happens when the flag service is down? Fail closed for money and data paths. Fail open only when the default is the safer legacy behaviour you still trust. Recovery design matters here as much as it does for workers and queues (Building Systems That Recover Instead of Restart).
This model is a tradeoff, not a purity test. Temporary rollout flags are valid. Permanent plan entitlements are valid. Mixing them in one undifferentiated bag is what produces debt.
A rollout flag that lasts eighteen months is no longer a rollout. It is an undeclared product SKU.
An entitlement encoded only as a flag with no billing or admin surface is a shadow catalogue. Sales will sell it. Finance will not recognise it. Support will invent folklore about how to turn it on.
How this shows up in delivery
On legacy modernisation programmes, we often find flags sitting on top of two coexisting implementations of the same workflow: old path and new path, selected by tenant or by "temporary" toggle. The modernisation work is not finished when the new path exists. It is unfinished until the control surface is explicit: which tenants are on which path, who may flip them, what kill switch exists, and what deletes the bridge after cutover. Treating that as "just a deploy flag" leaves two products in production with one roadmap.
A checklist before you create or keep a flag
- What is the primary intent: rollout, entitlement, experiment, kill switch, or migration bridge?
- Who is the subject of evaluation (user, account, plan, region, partner), and who must not see this?
- Who owns this flag, and what date or trigger forces review or removal?
- What is the documented kill switch behaviour if this path fails in production?
- How will support and sales answer "what does this customer have?" without reading code?
- If this flag still exists in six months, what product truth does it encode, and should that live in billing, admin, or config instead?
If you cannot answer ownership, subject, and expiry, do not create the flag yet. You are about to borrow clarity from the future.
Implications for product and engineering
Separate catalogues mentally even if one tool stores them. Rollout flags should be noisy about age. Entitlements should be boring and auditable. Experiments should die on schedule. Kill switches should be few, named, and tested. One flat list of three hundred booleans trains everyone to treat all of them as temporary.
Name for the business question, not the ticket. enable_invoice_pdf_v3 tells you a version. invoice_pdf_layout_for_eu_vat tells you a rule. Future you will thank the second name when an auditor asks why two PDFs differ.
Put flag state in the incident story. When behaviour diverges by customer, the timeline must include which flags evaluated how. Otherwise you debug code that is "correct" for a path the customer never hit.
Budget removal like feature work. Cleaning flags is not chores for a quiet Friday. It is reducing the branching factor of the live product. Roadmaps that only add flags and never retire them are building a second configuration product nobody staffed.
Do not use flags to avoid deciding ownership. If two teams disagree about whether a path is ready, a flag postpones the disagreement. It does not resolve who owns the customer promise when both paths are half-true.
None of this argues against flags. Flags are one of the best risk tools in modern delivery when they are treated as deliberate controls. The failure mode is treating them as a guilt-free way to leave unfinished product decisions in production forever.
Deploy without fear is a good slogan.
Ship with a known audience, a known rule, and a known kill switch is a better operating principle.
The fingerprint
If your release process is calm but nobody can answer which customers are on which behaviour without opening the flag console, you do not have a delivery problem.
You have an unowned business control surface wearing the costume of CI.
Related reading
- Why Every Workflow Is Really a State Machine
- The Economics of Technical Debt
- Building Systems That Recover Instead of Restart
- Why Race Conditions Are Business Problems
Architecture Files #34 · Part of Binary and Beyond. LinkedIn newsletter edition follows. Making release controls match product ownership and kill-switch reality? Start a conversation.
