Introduction
Alarms are the cheapest protection layer on the facility — and that is precisely the problem. Adding an alarm costs a configuration entry and a moment's optimism, so every project, every trip review, and every incident action adds a few more. Twenty years later the control room receives thousands of alarms a day, the operators have learned to acknowledge without reading, and the one alarm that mattered scrolls past in a flood of five hundred others.
Every major alarm-related incident investigation tells the same story: the information was present, and unusable. The discipline that fixes it is alarm rationalisation, sitting inside the alarm management lifecycle of ISA-18.2 (and IEC 62682, its international twin), with the performance benchmarks of EEMUA 191. This post covers how the process works and where it goes wrong.
What Good Looks Like: The Benchmarks
EEMUA 191's targets have become the industry's common language for alarm-system health, per operator console:
- Steady state: on the order of one alarm per ten minutes — roughly 150 a day. Above that, alarms are being managed instead of the process.
- Upset: a flood is conventionally defined as more than ten alarms in ten minutes; the design intent is that even during an upset the rate stays within what a person can read, diagnose, and act on.
- Standing alarms: approaching zero. A wall of stale standing alarms is camouflage for the new one.
- Priority distribution: roughly 80% low, 15% medium, 5% high. When a system shows 40% high-priority, the priorities are decoration.
Few mature facilities meet these numbers without deliberate work — that is what the rationalisation programme is for.
The ISA-18.2 Lifecycle
ISA-18.2 treats alarms like any other engineered system, with a lifecycle rather than a one-off clean-up:
- Philosophy — the governing document: what an alarm is (an event requiring operator action — not a status, not a log entry), the priority-setting method, performance targets, roles, and management of change.
- Identification — candidate alarms from HAZOP actions, C&E development, operating procedures, and incident findings.
- Rationalisation — the systematic review of every candidate against the philosophy (below).
- Detailed design — setpoints, deadbands, delays, priorities, HMI presentation.
- Implementation, operation, maintenance — training, shelving discipline, periodic testing.
- Monitoring, MOC, and periodic audit — KPIs against the benchmarks, no alarm added or changed outside change control — the leak that refills the system — and a periodic audit of the whole system against the philosophy.
Rationalisation: The Four Questions
In the rationalisation session, every alarm must answer four questions or be deleted:
- What causes it? The specific process condition — not "level high" but which scenario drives it.
- What is the consequence if the operator does nothing?
- What is the operator's corrective action? If there is no action, it is not an alarm — it is an event for the historian. This single test removes a remarkable fraction of most legacy databases.
- How much time does the operator have? The gap between annunciation and the consequence (or the trip) — which drives both the setpoint and the priority.
Priority then falls out of a matrix of consequence severity against time to respond — set by rule, not by negotiation. Documenting the answers in a master alarm database turns the session's output into the permanent basis for the settings, the HMI response guidance, and every future audit.
Alarms as Protection Layers
Where a LOPA takes credit for an alarm as an independent protection layer, the rationalisation carries extra duties: the alarm's initiator must be independent of the initiating cause and of the SIS it backs up, the operator response must be a defined, drilled procedure, and the available response time must genuinely support the credit — with a risk-reduction credit normally capped at a factor of ten (PFD 0.1); the deeper 0.01 credit is defensible only in the rare well-proceduralised, drilled case with more than 40 minutes of available response time. An IPL alarm also inherits testing and MOC obligations closer to instrumented-function discipline than to ordinary alarm housekeeping — highest priority class, restricted suppression, periodic proving of the full path from sensor to annunciation.
Engineering the Flood Away
Rationalisation sets which alarms exist; a handful of design techniques control when they speak:
- Deadbands and on/off delays sized to the signal's noise and dynamics kill chattering alarms — routinely the top of the bad-actor list.
- State-based alarming suppresses alarms that are meaningless in the current mode: the low-flow alarm during a planned shutdown announces nothing. Designed suppression, engineered and documented, is the difference between a quiet system and a blind one.
- First-out and grouping on trip events: the cause-and-effect logic knows which initiator tripped the unit — present that, not the eighty consequential alarms that follow.
- Shelving with governance — operators need a sanctioned way to silence a nuisance alarm, with automatic return and visible accounting. Without it they will find unsanctioned ways, which do not return automatically.
Common Failures
- Priority inflation — everything important becomes high-priority until nothing is.
- The suppression graveyard — alarms disabled years ago, undocumented, discovered during the incident investigation.
- Copy-paste setpoints — identical alarm settings across non-identical equipment, inherited from a template nobody remembers.
- Alarms as procedure reminders — using the annunciator as a to-do list, in defiance of the philosophy's own definition.
- Rationalising once and walking away — without KPI monitoring and MOC, the database regrows; the lifecycle exists because entropy is patient.
Conclusion
An alarm system is an operator-performance system: its output is not annunciations but correct human action under pressure. Rationalisation to ISA-18.2 gives every alarm a cause, a consequence, an action, and a time; EEMUA 191's benchmarks tell you whether the result works at the rate a human can use; and the lifecycle keeps it that way after the project team disbands.
The test worth designing to is simple to state: on the facility's worst day, the alarm system should be the thing that helps.
