Taming the wild alarm system - part two

Engineer Operating Control Panel At Power Plant Generating Electricity For City.

The most important alarm improvement technique in existence

There is a single method that has more effect, at lower cost and with lower effort, than any other technique at improving an existing, poorly performing alarm system. But what do we mean by "poorly performing?" Here are examples from some of the worst-performing alarm systems we have encountered - all of which were solvable.

  • Many different control systems with individual alarms that occur over 100,000 times per month

  • An alarm system with over 70% of all alarm occurrences (about a thousand a day) caused by instruments that were not working and needed maintenance

  • A system so dominated by a few nuisance alarms that 98% of all alarm occurrences came from just seven alarms, averaging over 600 a day

  • A system without good management of change, where uncontrolled and untracked manual alarm suppression eliminated 98% of all alarm occurrences (about 18,000 a day) from the operator's view. This included the suppression of some very important alarms

  • Many systems have over 25,000 alarms per day on average, with some exceeding 100,000 – that is one alarm every 3 seconds, to more than one alarm per second

  • A system that was in continuous alarm flood, averaging almost 40 alarms per minute for over four days

  • A single alarm that occurred over 200,000 times in one day

  • A large networked multi-site facility that generated over a billion alarms per year – 2.7 million a day

At first glance, problems such as these seem overwhelming. How do you possibly deal with 50,000 alarms a day? We assure you that, with some smart, focused effort, cases like these can be vastly improved in just a few days to a few weeks.

Seven steps

There is a seven-step process for improving existing alarm systems that is simple and proven effective in over a thousand alarm improvement projects.

  1. Develop an Alarm Philosophy document.

  2. Analyze your existing alarm data to establish a baseline and identify your problem areas.

  3. Perform "bad actor" alarm resolution.

  4. Perform alarm documentation and rationalization (D&R) and create a master alarm database.

  5. Implement alarm audit and enforcement technology for management of change.

  6. Implement real-time alarm management techniques, such as state-based alarming.

  7. Control and maintain your improved system, with ongoing analyses and work processes.

The first three steps are often initiated simultaneously. These three steps are easy, fast, low-cost and do not require a lot of internal resources, with powerful results at the start.

The alarm philosophy is important but is not a "prerequisite" for finding and fixing your most frequent alarms. The step of alarm analysis also involves setting up to monitor alarm system performance going forward. Both of these are mandatory requirements of the ISA 18.2 alarm management standard. But even the initial baseline by itself can point you to the critical step three – finding and fixing your most frequent and nuisance alarms – the "bad actors." We will cover the other steps in future blog posts.

This bad actor resolution step can cut your alarm rates by 60% to 80% - and even more. It can be accomplished in as little as a few days or weeks of part-time effort and does not require consultants. While it does not solve many problems (such as poor alarm-priority selection), it is a great way to make an impressive start and lend credibility to the alarm-improvement effort, helping you gain buy-in and build momentum.

There are several categories of bad actor (nuisance) alarms and several methods for dealing with them. With enough bad actors, an alarm system is rendered useless. This may lead to hazardous plant conditions when important or critical alarms are lost in the "sea" of bad actor alarms.

Experience shows that a comparatively small number of configured alarms cause most alarm occurrences, which drive all high-alarm-rate issues. Few means 20 to 50 individual configured alarms. No one ever intentionally designed an alarm to occur 20,000+ times a month, but they exist and they can be fixed.

The top 20 most frequent alarms typically comprise anywhere from 25% to 95% of the entire system load. If those alarms are managed successfully, then a major system improvement occurs. It is amazing that such high numbers of nuisance alarms exist, because it is doubtful the best control engineer in a company could intentionally design alarms to behave in the ways we will discuss. Yet they do exist; all varieties are in almost every system we analyze.

Consider one real-world example: An analysis of eight weeks of alarm data found that just 10 alarms accounted for 96% of the total alarm load, with several occurring more than 100,000 times. Five of those 10 alarms were “BADPV” alarms indicating malfunctioning instruments. Addressing only those few alarms and instrument issues could dramatically reduce the overall alarm load.

The impact of bad actor resolution can be significant. Across 15 different control systems, fewer than 50 alarms per system were analyzed using these techniques. The result was an average alarm-rate reduction of more than 65%. In other words, identifying and addressing a relatively small number of problem alarms can substantially improve overall alarm system performance.

Here are the major types of nuisance alarms:

  • Chattering alarms (quickly clear, then immediately repeat)

  • Fleeting alarms (last only a few seconds before clearing, and might repeat later)

  • Stale alarms (are in effect continuously for days, weeks, or months)

  • Suppressed alarms (the operator does not see when they occur, but whose suppression is not controlled and tracked)

  • Duplicate alarms (dynamic, where one condition causes multiple but different points to alarm)

  • Duplicate alarms (configured, where multiple linked points all alarm if any of them has an alarm)

  • Nuisance instrument diagnostic alarms (such as "bad measurement" types)

The first two — chattering and fleeting alarms — are often the biggest contributors to high alarm rates. Fixing them requires a calculation technique that’s a bit too involved to tackle here, so we’ll give them the attention they deserve in the next blog in this series. (If you can’t wait, see the references at the end).

Stale (long-standing) alarms

Stale alarms come in and remain in alarm for extended periods. A good starting point is to look for alarms that have been in effect for more than 24 hours. We have found alarms that have been in effect for months and even years. They clutter the alarm screens and devalue the importance of all alarms.

Are there truly many abnormal conditions lasting more than a day that require operator action to avoid a consequence? Or for months? Such alarms often reflect stable unit conditions, such as equipment that is intentionally shut down. They generally indicate alarms that were not configured in accordance with the principles contained in The Alarm Management Handbook.

Stale alarms are dealt with by understanding the process states and hardware involved. They are usually eliminated by reconfiguring them, so they comply with the very definition of an alarm. Alarms that go stale are often not alarms at all – they are merely status indications. They indicate if an item is on or off. One should not create an alarm that is based on some item just being on or off. There are always valid circumstances that an item should be off. Instead, the alarm should indicate that this item is supposed to be on but is off (or vice versa). Such a situation is abnormal and requires operator action. Design of that alarm may require some imagination, or implementation of some logic or a simple state-based alarm method.

Suppressed alarms

An initial analysis of a system used for determining the bad actor resolution list must also identify any configured alarms that are suppressed. This means the alarm is still configured, but an override has been selected to prevent its annunciation to the operator. Almost all control systems have this capability and it is often misused. Alarm suppression is often uncontrolled. We have found very important alarms that were suppressed for months with no one being aware. At the end of the bad-actor resolution step, there should be no remaining suppressed alarms. Alarms are often suppressed because of nuisance behavior, such as chattering, which can be fixed. Suppression must be rigorously controlled, visible and tracked. This is a technique called alarm shelving.

Two types of duplicate alarms.

1. Dynamic duplicate alarms

These consistently occur within a short time period of other specific alarms. If you use your alarm analysis software to list the alarms that always occur within, for example, one second of each other, you will likely find a good list to work on. Such alarms are likely to be multiple annunciations of the same process event, in different ways. For example, if a pump stops, one might immediately get low discharge pressure, low flow and low amps alarms. Those others could be valid alarms when the pump is running, but not when it is intentionally stopped and those values are expected.

The individual situation will determine which alarms are kept and which are not, or what logic adjustments must be made.

2. Configured duplicate alarms

Interconnections between points in a DCS can create cases of duplicate alarm configuration. For example, a process measurement sensor point may be connected to a selector point, to a totalizer point, to a logic point, to a controller point and so forth. Often a bad measurement type of alarm is configured on each point (usually by default), and thus if the sensor point goes into that condition, several simultaneous alarms will result. These distract the operator by annunciating multiple alarms caused from a single event (the one bad sensor). There should only be one such alarm, configured on the point where the operator is most likely to take the action. If the sensor point feeds a separate controller point, the controller would be the proper point to alarm on the bad measurement. This is because the operator action to be taken from a bad reading is likely to put the controller in manual mode and adjust the output manually. The controller point itself will show that the input measurement has gone bad.

Nuisance instrument diagnostic alarms

It is common, but still surprising, to see large numbers of alarms indicating a bad measurement or a similar instrument problem. In some systems, these occurrences can number in the hundreds or thousands, with individual instruments generating hundreds of bad measurement alarms in a single week.

When a loop was designed, did someone tell the control engineer, “Oh, and by the way, I want this sensor to go into bad measurement frequently”? Probably not, but we find these situations on almost every system we look at.

Since no instrument was designed to be in such a state, every one of these situations can be fixed. They are misconfigured in range, in measurement clamping or there is an installation problem (e.g., impulse leads filling up). The original justification for installing a flow meter probably did not include a specification that it was okay if it didn't work half of the time. But people put up with it. We wouldn't put up with a broken speedometer on our car.

These situations must be addressed. An instrument malfunction removes a process indicator from the operator's view. The time operators spend confirming the instrument problem reduces their attention to other duties. If a non-working instrument is not needed, it should be removed, following a management-of-change (MOC) procedure. An indefinitely broken instrument could be considered to be an MOC violation.

Decades ago, the available analog instrument sensors had a significant tradeoff between accuracy (significant digits) and range; you could obtain high accuracy only over a small range, probably less than the possible variation of the process. Control engineers were well aware of this tradeoff and were accustomed to designing within those constraints. But when such sensors with constrained ranges are implemented in a DCS, bad-measurement alarms occur frequently and do not represent an abnormality.

The digital electronic revolution that gave us the DCS also gave us much-improved measurement sensors. Modern sensors can generally provide all of the accuracy needed over the entire range the process is likely to vary. But some installations continue to follow older configuration practices and do not consider the consequences of generating numerous bad measurement alarms under conditions such as startup and shutdown. Controller points will usually have shed modes. These are predetermined actions taken when an input measurement goes bad, such as go full output, go zero output or maintain the last output. These should be chosen with care, but minimize the possibility of the measurement going bad in the first place.

The default should now be to configure the instrument to cover the full range of possible process values (including shutdown or ambient conditions) and then check the accuracy. If not (rarely, with modern transmitters), buy a better transmitter. But don't configure the range where you know you will get a bad measurement state at expected conditions.

Differential pressure flows are often the worst offenders. If, at zero flow, there is a slight imbalance in the leads, the meter attempts to report a backward or negative flow. The flow range might not be configured for a negative, so the bad measurement condition and alarms occur. Such points should be configured to handle the zero case. A cutoff can be configured and clamped to zero, so a small negative flow number is not produced, which could affect downstream calculations.

Most DCSs can clamp an analog value at the ends of the range rather than enter a bad-measurement state. This ability should be fully understood and used properly.

Ongoing work process

Nuisance alarms don't stay fixed. Processes change, sensors age and new issues emerge — which means ongoing alarm analysis isn't enough on its own. Someone has to own the response and act on what it surfaces. The good news: once operators experience a clean alarm environment, they expect it to stay that way. That accountability is exactly what drives continuous improvement.

Ready to tame your alarm system? Contact us — or dive deeper with our white paper: Making a big dent in nuisance alarms.

Review other Taming the Wild Alarm System topics in this 7 part blog series:

  1. Part one - How did we get in this mess?

  2. Part two - The most important alarm improvement technique in existence

  3. Part three - Silence the noise: fixing chattering and fleeting alarms

  4. Part four - Just how bad is your alarm system?

  5. Part five - What alarm rationalization really uncovers — and why it matters

  6. Part six - Why did they have to call it philosophy?

  7. Part seven - Beyond alarm management – doing more with a powerful tool