Est.

Alert Noise as an Engineering Productivity Tax

Unaddressed alert noise costs engineering teams millions in lost productivity and reliability.

Editor at Large · · 11 min read
Cover illustration for “Alert Noise as an Engineering Productivity Tax”
Cost of Observability Gaps · September 21, 2026 · 11 min read · 2,574 words

Alert noise gets treated as a tuning problem: adjust a threshold, mute a flapping check, silence a channel, move on. That framing is wrong. Alert noise is a structural tax that compounds with every new service, every unowned threshold, and every page an engineer learns to ignore, and treating a structural problem as a configuration problem is why the bill never stops growing.

A configuration fix doesn't scale, because it treats a symptom instead of the system generating it. A structural problem needs a discipline instead: defined ownership, measurable signal-quality standards, and a feedback loop that catches drift before it becomes the norm. Most teams never build that discipline, because the failure modes producing noise look like isolated incidents rather than what they actually are. Static thresholds get set once during a service's early days and never revisited as traffic patterns shift. Overlapping tools cover the same service, each firing its own version of the same event, a pattern common enough that tool-sprawl duplication is widely treated as the default state rather than the exception. Over-sensitive triggers fire on deviations too minor to warrant action, often with no remediation path attached even when they do matter.

None of these failure modes happens by accident. Each is what monitoring surface area looks like when it grows continuously while no one owns the aggregate signal-to-noise ratio, only the individual alert. Alert noise behaves like a tax, and like any tax, it compounds quietly until the bill comes due.

What the numbers say about how much noise engineering teams live with

Start with volume. A 2025 survey of hundreds of DevOps and SRE teams found that teams receive over 2,000 alerts a week, and only 3% need immediate action. Healthy alerting systems are generally expected to run at a substantially higher actionable rate. A system running at 3% is a fundamentally different system, one built on a different logic than a healthy alerting system. It's a different system entirely, one where the page has stopped meaning anything by default.

Other findings from the same research point the same direction. Eighty-five percent of teams report that most of their alerts are false positives. Seventy-four percent say their on-call engineers feel overwhelmed by volume. Sixty-seven percent admit to ignoring or dismissing alerts without investigating first, two out of three engineers waving pages through without a second look. Call that a rational adaptation to an environment where most of what fires doesn't matter. Training doesn't fix an incentive structure that rewards ignoring the pager.

The reliability consequence follows close behind. Splunk's State of Observability report, surveying 1,855 organizations, found that 73% experienced outages linked to alerts that were ignored or suppressed, a productivity tax turning into a reliability tax. That's where a productivity tax turns into a reliability tax. The inverse failure runs just as often: the State of Production Reliability and AI Adoption Report found that 78% of organizations experienced incidents where no alert fired. Noise and silence are the same calibration failure: miscalibrated thresholds either fire too often, producing noise, or fail to fire at all, producing silence, on opposite sides of the threshold.

How the productivity tax accumulates: context switching, desensitization, and ignored pages

The tax doesn't arrive as one bill. It accrues through three mechanisms that build on each other, and the first one is the most underpriced.

Time-to-acknowledge metrics capture how fast someone dismisses an alert, but that's not where the cost lives. The real cost is the deep work an engineer loses re-entering a task after the interruption, and recovery estimates for that run around 23 minutes per switch. Dismissing a page takes seconds. Getting back to where you were before it fired does not, and that asymmetry is the part no dashboard tracks.

Desensitization comes next, and it's a cognitive response. Research into alert fatigue has found that repeated exposure to alerts that require no action trains the brain to reduce attention to that stimulus over time. That's habituation, and it's how cognition works under repeated low-stakes stimuli. The trouble is that a genuine incident looks, on the surface, exactly like the noise that trained the response. So when something real does fire, the reaction runs slower, because the system spent months teaching the engineer that pages rarely deserve full attention.

Suppression is where habituation ends up. At a 67% rate of engineers admitting to routine dismissal, that's not an individual lapse anymore. It's an emergent team policy nobody voted on, and the cost of that policy shows up in Splunk's 73% outage-linkage figure: suppressed and ignored alerts don't just waste time, they eventually let something real slip through. By the time an engineer engages with an actual incident, they're doing it with depleted attention, no context, and a learned skepticism toward the tooling meant to help them. The scramble to reconstruct context in an incident channel is downstream of an instrument that trained its own users to stop trusting it.

Diagram: Alert Volume vs. Actionable Rate: The 3% Problem. Visualizes: Visualize the stark contrast between the volume of alerts engineering teams receive and the tiny fraction that actually require action.

The dollar cost of noise: toil hours, incident severity, and retention

Diagram: The Toil Threshold: From Manageable to Reactive. Visualizes: Show toil as a percentage of engineering time across three reference points: the Google SRE Workbook ceiling of 20–25%, Catchpoint's 2020 baseline of 25%, and Catchpoint's 2025…

Toil is the more controllable number, so start there. Catchpoint's SRE Report put median operational toil at 34% of working time, above the ceiling the Google SRE Workbook sets at 20 to 25%. Past 30%, teams shift into reactive mode and feature delivery suffers in a way leadership eventually notices without being told.

The climb from 25% to 30% is what makes the number sting. Catchpoint's 2025 report found toil had risen to 30% from 25%, the first increase measured in five years, during a period of heavy investment in monitoring tooling. More tooling did not buy less toil. That single fact should end the assumption that observability spend automatically produces better outcomes, because the data says otherwise.

Translate 30% lost time into dollars and the number stops being abstract. A simplified model using a $125,000 average engineering salary puts organizations with 250 or more engineers at a multi-million-dollar figure a year in lost productivity from toil alone. Exact totals shift with geography and role mix, but the order of magnitude holds across most mid-to-large engineering shops.

Downtime is the tail risk produced by that toil: sustained alert fatigue degrades response quality, which is how toil converts into downtime risk. Industry data puts the average cost of unplanned downtime at $5,600 per minute. More than 54% of significant outages cost over $100,000, and around 16% cross the $1 million mark. Noise doesn't just waste hours day to day. It raises the odds that the incident nobody caught in time becomes the expensive kind.

Then there's a cost that rarely makes it into an incident budget: people leaving. Catchpoint's report found nearly 70% of SREs say on-call stress contributed to burnout and attrition, and a separate industry survey found 22% of engineering leaders facing critical burnout, with another 24% at moderate levels. Turnover is real money, recruiting cost, ramp time, institutional knowledge walking out the door, none of it filed anywhere near the postmortem where the noise actually originated. Reducing alert noise is cost avoidance, plain and simple, and that framing is what makes the investment defensible to a budget owner who has no patience for on-call morale treated as an abstraction.

Why more monitoring coverage makes the problem worse before it makes it better

The instinct after a missed incident is almost always to add more monitoring. That instinct is wrong in the near term, because more surface area generates more events, and without something upstream filtering for actionability, the signal-to-noise ratio gets worse, not better.

The clearest evidence sits in the AI investment paradox from the previous section. Toil rose from 25% to 30% in Catchpoint's 2025 report during the same window that AI adoption in observability tooling climbed sharply. Roughly half of surveyed SREs say AI reduced their toil. Roughly half report no change or more work. Adoption alone doesn't predict the outcome. What the tool gets pointed at does.

Part of this is mechanical. Every new monitoring layer brings its own alert vocabulary, its own thresholds, its own escalation logic. Stacking four or five of those without deliberate deduplication and correlation causes overlapping tools to start producing overlapping pages for the same underlying event. The 2026 State of Production Reliability and AI Adoption Report found that 83% of organizations juggle four or more separate tools during a live incident, each with its own view of what's happening and none of them talking to the others.

One assumption deserves to be retired outright: that pre-production testing is closing the gap driving this volume. The 2025 DORA State of AI-Assisted Software Development report, surveying nearly 5,000 tech professionals, found only 8.5% of organizations hit elite-level change failure rates of 0 to 2%, while 39.5% still see failure rates above 16%. Staging environments aren't catching what production catches, and no amount of added monitoring coverage changes that on its own.

What matters is whether every alert fired is actionable, right now, by someone with the context to act on it. Coverage answers a different question entirely, and mistaking one for the other is how organizations end up spending more on observability while toil keeps climbing.

What a well-designed alerting system looks like as an engineering target

Two to three actionable incidents per on-call shift, with operational work capped at 20 to 25% of total engineering time: that's the concrete target the Google SRE Workbook sets. It's a load-bearing number, the one that keeps feature delivery from collapsing under reactive work.

Actionability has to be the design criterion, full stop. An alert should fire only when there's a defined action the recipient can take immediately. If the honest answer to "what do I do with this" is "wait and see," the alert is a dashboard metric that got miscategorized as urgent. It should get demoted, not tuned.

Five structural properties separate low-noise systems from the rest, and each is an engineering decision rather than a knob to turn. Dynamic baselines, calibrated against historical traffic patterns, replace static thresholds, which remain the single largest source of false positives because nobody revisits them once traffic patterns shift, and that neglect is what produces the false positives. Deduplication and correlation group related signals by service, host, or time window into one incident instead of a dozen separate pages for the same root cause, the way Cambia Health Solutions, running AIOps in its network operations center, achieved critical alerts surfaced within 30 seconds and 95% SLA compliance. Tiered severity routes non-urgent findings to a Slack channel with a four-hour SLA and genuinely critical issues to an immediate page, instead of treating every alert as equally urgent, a design failure that hides as a safety margin. Clear ownership assigns every alert to someone accountable not just for responding but for maintaining its signal quality over time, since ownership without accountability for the false positive rate is only half the job. A continuous audit loop reviews signal-to-noise KPIs on a set cadence, because alerting configurations left untouched drift toward noise as the systems underneath them change shape.

None of this aims for silence. It aims for a state where every page represents a fact about production that genuinely requires a human decision in that moment. Anything short of that is the system outsourcing work to an engineer that the system should have handled on its own. Production telemetry is the only ground truth available here: an alert that never triggers in CI and never surfaces in staging only gets validated by real traffic. Alerting discipline and post-deploy observability are, in practice, the same discipline wearing two names.

How AI changes the alert correlation and triage layer

Correlation, not remediation, is the defensible entry point for AI in this stack. Read-only analysis that reduces operator workload before any agent touches production is where trust gets earned. Staring down 200 discrete alerts versus reading three coherent incident narratives is a gap substantial enough to matter on its own.

Correlation does what rule-based systems structurally can't. It pulls signals across logs, metrics, traces, and topology into a single investigation path instead of firing one alert per affected component. It generates root-cause hypotheses from live telemetry and incident history rather than pattern-matching against a static signature library. And it adapts to failure modes no runbook ever anticipated, since it isn't limited to matching against what someone already wrote down.

The resolution-time numbers coming out of early deployments are striking. WGU's SRE team, using the AWS DevOps Agent, cut total resolution time from an estimated two hours to 28 minutes, a 77% improvement. At SREcon25 EMEA, Solo.io reported infrastructure incident resolution dropping from four hours to eight minutes using specialized AI agents. Augment's internal Cosmos Incident Investigator Expert reportedly cut human on-call investigation effort by around 81%.

Those numbers need an honest counterweight, and IBM supplies one. IBM's ITBench evaluation tested current AI models against 42 real-world SRE scenarios and found they resolved 13.8% of them. Setting the vendor-reported wins in the double-digit-minute range against that 13.8% independent resolution rate reveals a gap wide enough that it should govern how much autonomy anyone hands these systems today, not how much marketing copy gets built around them.

That gap is why the trust ladder matters: read-only correlation first, advised recommendations second, human-approved execution third, and bounded autonomous remediation only after a track record earns it. Irreversibility, not confidence, should decide when human oversight stays mandatory. None of this replaces the discipline covered above, either. AI correlation performs best when the underlying signal has already been shaped by dynamic baselines, tiered severity, and clear ownership. Pointed at raw, undisciplined noise, it's still analyzing garbage, just faster.

Agentic development creates a new class of alert noise that existing tooling cannot see

Gartner projects that 40% of enterprise applications will include task-specific AI agents by the end of 2026, up from under 5% in 2025. That's a step change, not a gradual shift, and every deployed agent is a new source of production changes. Each new source of production changes is a new source of potential alerts that existing tooling was never built to catch.

The failure modes are the real problem here. An agent can return an HTTP 200 while producing a wrong answer, burning through its token budget, or drifting silently off-task, and none of that trips an error code or crosses a threshold, because nothing in the traditional monitoring vocabulary was built to measure task correctness. Non-deterministic systems can't be reliably reproduced in staging, so production behavior under real traffic becomes the only ground truth available, the same principle from earlier in this piece, now applied to a system that behaves differently every time it runs.

The early data on how this actually plays out is unsettled, and it deserves to be treated that way rather than smoothed into a clean trend line. Gravitee's State of AI Agent Security report found that 88% of organizations running AI agents reported some form of incident. Confirmed incident rates dropped between December 2025 and April, which could mean genuine improvement, or could just as plausibly mean detection is maturing faster than the underlying agent behavior. Either reading leaves the same gap standing. The existing alerting discipline, dynamic baselines, correlation, tiered severity, clear ownership, was built for systems that fail in recognizable ways. Agentic systems don't yet fail in ways that discipline was designed to catch, and that gap is where the next round of noise, and the next round of missed incidents, is most likely to start.

Sources

  1. Alert Noise Reduction: A Complete Guide to Improving On-Call Performance (2025)
  2. AI SRE: The 2026 Guide to AI-Powered Site Reliability Engineering
  3. AIOps for SRE — Using AI to Reduce On-Call Fatigue and Improve Reliability - DevOps.com
  4. Alert Fatigue Drags Down IT Production Environments, Leads to Costly Outages - Carrier Management
  5. dev.to
  6. augmentcode.com
  7. businesswire.com

More in Cost of Observability Gaps