How Alarm Rationalisation Is Actually Done

Rationalisation is the part of alarm management that cannot be automated and cannot be skipped. It is a structured review in which every alarm is justified, prioritised and documented — or removed. This article describes how the work is actually run: the test each alarm must pass, how priority is decided, how to structure the sessions, and the failure modes that stall the effort.

The Test Every Alarm Must Pass

An alarm exists to prompt a human action. That gives a single, unforgiving test. For each alarm, three questions must have answers:

  1. Who responds? A named role — board operator, field operator, shift supervisor. Not “someone”.
  2. How long do they have? An allowable response time before the consequence occurs. If the answer is “no particular deadline”, it is not an alarm.
  3. What do they do? A specific, available corrective action. If the only honest answer is “note it” or “call the engineer in the morning”, it is not an alarm.

If any of the three has no answer, the item is reclassified — as an event, a log entry, a trend, or an operator display indication — and removed from the alarm system. In a first-pass rationalisation on an unmanaged plant, it is normal for 20–40% of configured alarms to fail this test.

The most valuable output of rationalisation is not better alarms. It is fewer alarms.

Setting Priority: Consequence × Time

Priority is not a measure of importance in the abstract. It is a scheduling instruction telling the operator what to do first when two alarms arrive together. It is derived from two factors: the severity of the consequence if no action is taken, and the time available to act.

Consequence if no action> 30 min to respond5–30 min< 5 min
Severe — safety, environmental, major lossMediumHighHigh
Significant — equipment damage, off-spec productionLowMediumHigh
Minor — efficiency, minor quality deviationLowLowMedium

The matrix is calibrated in the alarm philosophy before rationalisation starts. Its purpose is to stop priority being argued case by case — the matrix decides, and the session records the two inputs rather than debating the output.

A properly rationalised system lands close to 80% low, 15% medium, 5% high. If your distribution comes out at 50% high, the matrix was not applied — it was overridden.

The Master Alarm Database

Rationalisation produces a record, and that record is the deliverable. The master alarm database becomes the authoritative definition of the alarm system; the DCS configuration is expected to conform to it.

The minimum field set per alarm:

FieldPurpose
Tag, description, alarm typeIdentity
Setpoint, deadband, on/off delayConfiguration of record
Consequence of inactionThe justification input
Allowable response timeThe second priority input
Priority (derived)Output of the matrix
Alarm classDrives testing, training and documentation obligations
CauseWhat makes this alarm occur
Operator actionWhat the responder does — the basis of the response procedure
Rationalisation date, participants, approverAudit trail
MOC referenceChange history

If the database and the DCS disagree, that is a finding — and the reconciliation should be automated and reported, not discovered during an audit.

How to Run the Sessions

Who is in the room

A process engineer who knows the consequences, an experienced operator who knows what actually happens on shift, a control systems engineer who knows what is configurable, and a facilitator who keeps the pace and owns the record. Four to six people. Larger groups halve the throughput.

Pace and scope

Plan on 20–40 alarms per hour once the team finds its rhythm — slower at the start, faster once precedents accumulate. Sessions longer than three hours produce visibly worse decisions. Work unit by unit, not plant-wide: a finished unit is a visible win and keeps sponsorship alive.

Sequence within a unit

  1. Pre-populate the database from the current DCS configuration and 30–90 days of alarm history — occurrence counts in front of the team change the conversation.
  2. Deal with the bad actors first; many resolve without discussion.
  3. Work through the remainder by system, applying the three-question test.
  4. Record consequence and response time; let the matrix set priority.
  5. Capture the operator action text — this becomes the response procedure.
  6. Approve, then implement through management of change.

Bring the alarm history into the room. A team looking at “this alarm fired 1,240 times last month and nobody did anything” reaches a decision in seconds that would otherwise take twenty minutes of theory.

Five Failure Modes

Failure modeWhat it looks likePrevention
No philosophy firstEvery alarm re-argues the rules from scratchWrite and approve the philosophy before session one
Engineering-only roomDecisions that do not survive contact with a night shiftAn experienced operator in every session, without exception
Priority inflationEverything ends up High “to be safe”Record consequence and response time; let the matrix decide
Rationalise but never implementA beautiful database and an unchanged DCSMOC and implementation scheduled in the same plan, not “later”
No monitoring afterwardsRates creep back within 18 monthsNamed owner, monthly KPI reporting, MOC on additions

How Long It Takes

For a mid-sized unit with roughly 2,000 configured alarms: expect two to three weeks of preparation, six to ten workshop sessions, and two to four weeks of implementation and verification. Total elapsed time of two to three months is realistic when the sessions are protected in people’s calendars, and considerably longer when they are not.

The preparation is where tooling earns its place. Arriving at session one with the current configuration, 90 days of occurrence data, chatter and stale classifications and a pre-populated database turns rationalisation from an archaeology exercise into a decision exercise.


Planning a Rationalisation

If you are scoping this work — whether for an owner requirement or after an audit finding — send us the alarm count, the units in scope and the applicable specification. We will come back with a session plan, a realistic duration, and the preparation data you would want in the room. NEO AEGIS holds the master alarm database and reports conformance between the database and the live system continuously.

Related reading: What is an alarm flood · EEMUA 191 vs ISA-18.2 · Chattering alarms and bad actors

Software Consultation

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top