Live · Sat, Oct 10, 2026 · 12:02 UTC Block 843,917 Fees 14 sat/vB Fear & Greed 72 · Greed
Newsletter Pro Terminal Sign in
ITop Field News.
Subscribe →
Live · 12:02 UTC Block 843,917 F&G 72
Software development Software development desk

Feature toggles at scale: when the flag list becomes the problem

Feature toggles are one of the most practical tools in a dev team's kit, but they accumulate quietly until the flag list itself becomes a source of bugs, confusion, and deployment risk.

Young software developer typing code on a laptop in a modern office setting, focused on programming.

Photo by Varun Bhatheja on Pexels

Feature toggles promise a cleaner way to ship. Push code to production, keep the feature dark, flip a flag when you're ready. It works beautifully for the first dozen flags. Then the list hits 80, three engineers have left the team, and nobody is sure which flags are still load-bearing and which ones haven't been touched since 2024.

This is the feature toggle debt problem, and it's more common than most Australian dev teams admit. The conversation usually starts because feature flags decouple deployment from release, which is genuinely useful. The problem isn't the pattern. It's what happens after the pattern succeeds.

How flag sprawl actually happens

Flags accumulate in two distinct ways. The first is intentional: a team adopts feature toggles properly, creates flags for every new release, and ships faster as a result. The second is opportunistic: someone adds a flag to hotfix a behaviour in production without a plan for removal. Both create flags. Only the first type has a clear owner.

The realistic lifecycle of a flag looks like this. A flag gets created for a new checkout flow. It rolls out, succeeds, and the team moves on. Six months later, a new developer assumes the flag is still gating something important and doesn't touch it. A year later, the flag is evaluated on every request, but the code behind the false branch hasn't run in production for eight months. Nobody removes it because removing it feels riskier than leaving it.

That pattern, repeated across 60 or 80 flags, creates a codebase where conditional logic has compounded into something unreadable. Testing becomes harder because the number of flag combinations grows exponentially with every new flag added.

The combinatorial testing trap

Two flags give you four possible states. Five flags give you 32. Ten flags give you 1,024. At 20 flags, you're looking at over a million combinations, most of which nobody has tested and many of which can't coexist in practice.

In reality, not all combinations are meaningful. But without documented constraints on which flags interact with which others, your test suite can't know that. Most teams end up testing a handful of common configurations and hoping the edge cases don't matter in production. Sometimes they're right. Sometimes a user in an unusual flag state triggers a code path that hasn't run since a test environment last year.

This is closely related to the problem of test suites that appear to pass but miss critical branches. High coverage numbers don't tell you which flag combinations your tests actually exercise.

Organisational signals that flag sprawl has set in

There are four concrete signals to watch for:

  • Engineers ask "is this flag still needed?" before touching any conditional block.
  • The flag configuration for staging diverges from production and nobody is sure when it happened.
  • A bug is traced back to two flags interacting in an undocumented way.
  • Onboarding new developers requires a verbal explanation of which flags are "real" vs "safe to ignore".

Any one of these is a warning. All four together means the flag system is already a maintenance liability.

A practical approach to flag lifecycle management

The single most effective change most teams can make is treating flags as temporary by default. A flag should have an expiry date or a removal condition set at creation. Not aspirationally. In the ticket. In the code comment. The flag's death is as planned as its birth.

Concretely, this means a few things. First, categorise flags by type. Release flags (hiding unfinished features) should be short-lived, removed within days or weeks of full rollout. Experiment flags (A/B tests) have a defined analysis window. Kill switches (emergency circuit breakers) are permanent but should be few in number and clearly documented as such. Ops flags (tuning behaviour without a deployment) sit in a separate management bucket entirely.

Second, build a flag registry. Not a Confluence page that goes stale. A structured record, ideally attached to your feature flag platform, that tracks the flag name, owner, creation date, type, and removal condition. If you're using LaunchDarkly, Unleash, or Flagsmith, most have built-in stale flag detection. Use it. Don't rely on memory or documentation that lives outside the tool.

Third, make flag removal a first-class task. It needs to sit in the same sprint as the feature work it closes out. Teams that treat removal as optional will never do it.

What flag evaluation in hot paths actually costs

Every flag evaluation has a cost. For flags backed by a local in-memory cache, that cost is negligible. For flags that make a remote call to a flag service on every request, in a hot code path, the cost adds up. A 5ms flag evaluation in a service handling 2,000 requests per second is adding 10 seconds of cumulative latency per second of wall time. That's not hypothetical. It's a common finding in performance reviews of services that adopted feature flags without thinking about evaluation overhead.

The fix is usually simple: cache aggressively, use a local SDK that streams flag updates rather than polling, and never put a flag evaluation inside a database loop. But the fix requires knowing the problem exists, which requires someone looking at the flag evaluation code critically rather than just at the feature it wraps.

When the flag system itself needs a refactor

There's a point where incremental cleanup isn't enough. If the flag list has grown past 100 active flags, if the categorisation is inconsistent, and if several flags have been "temporarily" in place for over 18 months, a structured purge is the right call.

Run it like any other refactor. Inventory every flag. Classify each one. Assign ownership. Set removal dates. Then work through the list in priority order: flags that interact with security-sensitive code first, flags in critical user paths second, everything else third. Block out sprint capacity. Don't treat it as background work.

One flag removed cleanly is worth more than ten flags documented carefully. The goal is fewer conditionals in the codebase, not better-organised conditionals.

Governance without bureaucracy

The resistance to flag lifecycle governance usually sounds like: "we don't want to slow down shipping." That's a reasonable concern. The answer isn't to add approval gates for every flag. It's to make the right behaviour the default.

If your flag creation template includes a removal condition field that engineers have to fill in, most engineers fill it in. If your CI pipeline fails when a flag has been stale for more than 90 days without a documented extension, most teams deal with it before the deadline. Defaults and friction shape behaviour more reliably than policies that require someone to remember to enforce them.

Feature toggles are a sound pattern. The teams that benefit most from them are the ones that treat the flag list as a living part of the codebase, not a configuration afterthought.

→ The Confirmations · Daily newsletter

One email at 06:00 UTC. Six minutes. The only digest written for desks, not for retail.