How did we get into this mess?
You've seen it in movies. Probably over-dramatized, but not by much. A big problem has come up in some control room (it doesn't matter what kind). It could be because of a malfunction or, perhaps, aliens. The computer control screens light up and start flashing. Horns are blaring and rotating beacons activate. Everyone is shocked and confused — it's a chaotic scene.
Sound far-fetched? It's not. In the minutes leading up to actual, major process accidents, operators can face hundreds to thousands of alarms occurring within just a few minutes. They experience alarm listings scrolling too fast to read, process computer graphics covered in bright flashing symbols and loud horns that recur as soon as they are silenced. The alarm system becomes a nuisance distraction to the operator instead of a useful tool to help deal with abnormal conditions. The likelihood of a successful outcome to the situation diminishes, and process outages or even damage can result. So, how did our alarm systems get this bad?
A short history
Like many problems, this one began with the best of intentions. In the good old days (say, before the 1990s), a typical control room had a wall full of individual process indicators, lights, switches and moving-pen charts. These items took up a lot of space, which was always in short supply. The alarm system was simple – a rectangular array of a few dozen (at most) labeled windows that individually lit up and flashed according to their process connections. This lightbox also incorporated a horn that would sound and an acknowledge button to silence the horn and change the flashing light to a steady light. (The operators often equipped this acknowledge button with a wedge of paper or coin to hold it in and keep the infernal noise from happening in the first place. This user modification would almost certainly be in place on the night shift, but it might get removed during the day.)
The control wall concept had many positive things going for it. Considerable thought went into instrument placement and grouping. Normal ranges on the instruments were marked. Trends were always visible if the paper and ink were replaced. The overall health of the process could usually be determined at a glance. The alarm display would often produce repeatable patterns depending on the type of upset.
These systems also had many disadvantages. Inter-controller connectivity was almost on-existent. Complex control schemes were difficult and expensive to create. The introduction of new controls involved either a costly relocation of adjacent elements or sacrificing their logical placement. Communication of control system information to other systems was generally impractical. And data analysis? Forget about it.
The addition of a new alarm on the lightbox was expensive. Their total number was limited by space availability and cost, so each was evaluated and justified individually (which was a good thing).
This was the situation prior to the digital revolution and the introduction of modern controls, such as distributed control systems (DCSs) and supervisory control and data acquisition (SCADA) systems. These systems provide substantial operational and business advantages, including expandability, ease of reconfiguring control strategies and process data history/analysis. Almost everything in the control system can be changed without much trouble. (All these attributes can bring with them their own problems.) For these advantages, organizations converted older-style control systems such as the one pictured to DCSs and SCADA systems beginning in the 1990s.
The situation for alarms is far different in a DCS than in an older system. Since alarms are displayed in computerized scrolling lists and in graphics, they have unlimited space. And since every "point" in the DCS is essentially a software construct, alarms became free. Most point types in a DCS have several pre-programmed alarms just waiting for the control engineer or other user to configure and activate them by touching a few keys. No justifying, no wiring, no tubing, no plastic engraving – just click, click, click and you have a new alarm.
And create them we did. With no consistent guidelines for alarm configuration, massive over-configuration of DCS alarms became common. After all, if the manufacturer supplied the functionality of a high, high-high and even HHH alarm, well then, they must be there for a good reason, so let's use them all.
With no guidelines or cost for creating alarms, poor practices arose – such as all alarms being enabled by default, settings made by inconsistent rules of thumb or settings by an individual's preference. Consistency was low; similar process systems implemented by different teams would have different alarm configurations and behavior. (Engineers love to be creative when we aren't given any guidelines). Alarms were often used as an easy method to indicate status (something is on or off) rather than indicating an actual abnormal situation (something is off, but it is supposed to be on).
The result? For the operator, this meant that while their former control wall likely had less than 100 possible alarms, their new DCS console likely had 2,000 to 4,000 configured alarms producing hundreds to thousands of alarm annunciations every day. Even during steady-state process operation, the alarm system is activating almost constantly, creating far more alarms than the operator can possibly understand and act upon. During a process upset, there is an order-of-magnitude increase in the number and speed of alarms, rendering the alarm system useless and creating an active hindrance to the operator's ability to deal with the situation. Time and time again, investigative reports following major industrial accidents have shown that overloaded, bypassed or ignored alarm systems have played a significant role in worsening the situation.
Major accidents are only the beginning. An ineffective alarm system can make an ordinary process upset worse and last longer. Such upsets can cost companies a lot of money.
The ease of modifying alarms in a DCS made this bad situation even worse. Not only could engineers modify the alarm configuration, but so could operators, maintenance technicians, contractors, managers and even college interns. Anyone can change an alarm from a console keyboard and at many installations, such changes had little security or oversight for years.
Since the 1990s, manufacturing sites have had rigorous management of change (MOC) policies to address almost any physical change in the facility itself, but these policies often did not apply to changes to alarms. For decades, many alarm systems have had settings that change from day to day because they are at the individual whim of various people. Imagine if pilots boarding an airliner had no idea where the previous pilots had left the aircraft's alarm settings. For many years, MOC policies and practices have often failed to adequately address the configuration, alteration and bypassing of alarms in a DCS.
The result? Widespread cases of overloaded and ineffective alarm systems.
Where are we now?
Industry experts began identifying and writing about the alarm problem. Investigations into some major accidents cited DCS alarm systems as significant contributing factors.
An example from the UK's Health and Safety Executive (HSE) report on a major refinery accident:
• There were too many alarms and they were poorly prioritized
• The control room displays did not help the operators understand what was happening
• In the last 11 minutes before the explosion, the two operators had to recognize, acknowledge and act on 275 alarms
Industry professionals wrote articles on alarm management and several companies began to offer various products and services to address the issue, including software designed to analyze alarm occurrences. The industry developed the concept of alarm rationalization to improve existing systems and introduced dynamic and real-time alarm management software.
The Abnormal Situation Management (ASM®) Consortium formed in 1994 and began studying aspects of the problem and acted to greatly increase awareness of it. In 1999, the UK's Engineering Equipment and Materials Users Association (EEMUA) produced a seminal reference document (their Publication 191) on the topic.
Octave published The Alarm Management Handbook, based on hundreds of successful alarm improvement projects. The book has been widely regarded as having the best and most practical knowledge on making alarm systems effective.
That book was followed by The High Performance HMI Handbook, which discussed ways to make process graphics effective. Among many other topics, the HMI book thoroughly details how to effectively display alarms in process graphics.
Octave co-authored the Electric Power Research Institute's recommended practice for alarm management and participated in the American Petroleum Institute's creation of a similar recommended practice for the pipeline industry. And Octave helped write the ANSI/ISA-18.2 Alarm Management standard, which was a major development with regulatory implications. Alarm management is now a thoroughly documented topic and the knowledge needed to fix an alarm system is widely available.
Regulatory agencies and alarm management
The regulatory environment concerning alarm management is complex. The mandatory statements in standards such as ISA 18.2 (and the IEC 62682 international version of 18.2) regulatory agencies worldwide can and do enforce. This is generally via a "general duty" clause in the regulations, such as "The employer shall document that equipment complies with recognized and generally accepted good engineering practices," (RAGAGEP).
A rigorous consensus process develops standards such as ISA 18.2 and IEC 62682. They are RAGAREP and so they become enforceable because of these general duty clauses. Regulators have assessed fines and penalties for noncompliance with standards. You can volunteer to be on an ISA standard development or review committee and help shape its direction. Please contact us for more information or if you have questions.
In this Taming the Wild Alarm System blog series, we will cover several practical ways to improve your alarm system. We will start by identifying and fixing your worst nuisance alarms using straightforward methods to achieve significant improvement with minimal effort.
Ready to tame your alarm system? Let's talk!
Review other Taming the Wild Alarm System topics in this 7 part blog series:
Part one - How did we get in this mess?
Part two - The most important alarm improvement technique in existence
Part three - Silence the noise: fixing chattering and fleeting alarms
Part five - What alarm rationalization really uncovers — and why it matters
Part seven - Beyond alarm management – doing more with a powerful tool