What Is an Alarm Flood — and How to Get Out of One

An alarm flood is not “a lot of alarms.” It is a defined, measurable condition in which the alarm system stops doing the one job it exists to do — telling an operator what needs attention now. This article covers what counts as a flood, why floods happen, and the sequence that actually gets a plant out of one.

The Definition Is a Number, Not a Feeling

Both EEMUA Publication 191 and ANSI/ISA-18.2 (adopted internationally as IEC 62682) use the same operating measure: alarms presented per operator console per 10 minutes. The thresholds are not opinions — they come from studies of what a human operator can actually process while also running a plant.

Alarm rate (per console, per 10 min)ClassificationWhat it means in practice
~1–2AcceptableThe operator can read, judge and act on each alarm
~5ManageableWorkload is high but recoverable
~10Maximum manageableAt the edge; judgement starts to degrade
>10Alarm floodThe operator can acknowledge but cannot evaluate

The distinction that matters is between acknowledging and evaluating. Above roughly ten alarms in ten minutes, an operator is still clicking — but the clicking has become a clearing action, not a decision. The alarm system is technically running and functionally dead.

Why Floods Happen

Floods are rarely caused by one bad alarm. They are the product of three things that accumulate quietly over years of operation:

1. Correlated cascades

One process upset propagates. A pump trips, and within seconds the system reports low flow, falling discharge pressure, rising suction level, a downstream temperature deviation, a compressor recycle action and three interlock pre-warnings. Every one of those alarms is technically true. Collectively they tell the operator nothing he could not have inferred from the first one.

2. Bad actors that were never cleaned up

In most plants a very small number of tags generate a very large share of the alarm load. It is common to find that the top ten tags produce more than half of all alarm occurrences — usually chattering instruments sitting on a setpoint, a failed transmitter nobody removed from service, or a valve limit switch that flutters. These do not cause the flood, but they raise the baseline so that a normal upset tips straight into flood territory.

3. Priority inflation

When every alarm is configured as “High” because no one wanted to be the engineer who down-prioritised the one that mattered, priority stops carrying information. A properly distributed alarm system has roughly 80% low, 15% medium and 5% high priority. Plants that have never rationalised often sit at 60–90% high.

Why This Is a Safety Issue, Not an Optimisation Issue

Alarm management became a formal engineering discipline because of incidents, not efficiency studies.

The 1994 explosion at the Texaco refinery in Milford Haven, Wales is the case most often cited. The investigation found that in the final eleven minutes before the explosion, two operators were presented with roughly 275 alarms. The information needed to prevent the event was on the screen. It was indistinguishable from everything else on the screen.

An alarm that cannot be acted upon is not a safeguard. It is noise wearing the uniform of a safeguard.

The 2005 BP Texas City incident raised related findings about alarm and instrumentation effectiveness. Together these are why owner specifications now routinely require EEMUA 191 or ISA-18.2 conformance as a contractual deliverable, with evidence, rather than as a design aspiration.

The Sequence That Actually Works

The instinct when facing a flood is to start suppressing alarms. That is the wrong first move — it hides the evidence you need and creates a new hazard. The sequence below is the one that holds up under audit.

Step 1 — Measure before you change anything

Collect at least 30 days of alarm history and establish the baseline: average alarm rate, peak 10-minute rate, flood duration and frequency, standing alarm count, priority distribution, and the ranked contribution of individual tags. Without this you cannot prove improvement later, and you will not know which changes did the work.

Step 2 — Remove the bad actors

This is the highest-yield step and it touches the fewest alarms. Identify chattering, fleeting and stale alarms, then correct them at source: deadband and on-delay tuning, repair or replacement of failed instruments, and removal of alarms on equipment that is out of service. It is common for this step alone to cut total alarm load by 40–60% without a single rationalisation meeting.

Step 3 — Rationalise

Every remaining alarm must answer three questions: who responds, how long do they have, and what do they do? An alarm that cannot answer all three is not an alarm — it is an event, a log entry or a trend, and it should be reclassified as one. Rationalisation also sets priority from consequence and available response time, not from habit.

Step 4 — Design for the flood case

Cascades cannot be fixed tag by tag. They need state-based alarming (different alarm sets for startup, normal operation and shutdown), first-out and cause-effect grouping so the root alarm is presented rather than its twenty consequences, and designed suppression that is documented, approved and auditable — never ad-hoc shelving that quietly becomes permanent.

Step 5 — Monitor so it does not come back

Alarm systems drift. Every modification, every new unit, every “temporary” alarm added during a turnaround pushes the rate back up. Continuous KPI monitoring with a named owner is what separates a plant that fixed its alarms once from a plant whose alarm system stays fixed.

What Good Looks Like

MetricTargetTypical starting point in an unmanaged plant
Average alarms per operator per 10 min≤ 210–60
Peak 10-minute alarm rate≤ 10100+
Time in flood condition< 1% of operating time5–20%
Standing (long-duration) alarms< 10 at any time50–200
Priority distribution (low/med/high)~80 / 15 / 5Often 20 / 20 / 60
Contribution of top 10 tags< 5% of load40–60%

These are not aspirational numbers. They are the published benchmarks, and plants reach them routinely once the work is done in the order above.

Where Software Fits — and Where It Does Not

Software does not rationalise alarms. Engineers and operators do that, in a room, with the P&IDs open. What software does is make the work possible at scale: it collects and normalises alarm history from the DCS, identifies bad actors and correlated sequences automatically, holds the master alarm database with its rationalisation record, and reports the KPIs continuously so drift is visible before it becomes a flood again.

That is the boundary we design to. NEO AEGIS covers the analysis, the governance record and the continuous KPI monitoring; the engineering judgement stays with the people who own the process.


Assessing Your Own Alarm System

If you want a baseline before committing to anything: export 30 days of alarm and event history from your DCS or historian, and we will return the six KPIs above, a ranked bad-actor list, and an estimate of how much of your alarm load is removable without any rationalisation workshop. No obligation, and the conversation is with an engineer.

Related reading: EEMUA 191 vs ISA-18.2 · Chattering alarms and bad actors · How alarm rationalisation is actually done

Software Consultation

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top