Just how bad is your alarm system?
And how can you tell? Alarm analysis has become commonplace. You have to analyze something to improve it. There are many analyses that provide great information for improving your alarm system. The most important one, though? It is all about alarm rate. How many alarms are being generated, in how much time, and can the operator receiving them deal with that number?
Looking at the alarm rate is independent of the type of process under control. It doesn't matter if you are making gasoline or aspirin or megawatts. All processes involve an operator monitoring and controlling factors like flow, temperature, pressure and composition. All alarm rate analyses are normalized by looking at the alarms presented to a single human – the operator responsible for dealing with them.
At any given time, a single operator is responsible. For example, looking at alarms per day for a continuously staffed process will involve either two or three humans in sequence, depending on if there are two shifts or three in the 24-hour period. We call this staffing combination an operating position. It is common that a process may have more than one operating position, with different operators controlling different parts of the process and coordinating when necessary. It is a best practice, though, that each operator is sent only the alarms relevant to their span-of-control responsibility.
All alarms are a source of human-machine interaction. The human must detect the alarm, understand it, examine the process to determine why it is occurring, determine the appropriate response, take that action and continue monitoring the process to ensure the chosen action is successful. This sequence takes time and thought.
An operator could accomplish those steps if the alarm rate is one alarm per hour. An operator cannot deal with one alarm per second (but we commonly see alarm rates far higher than one per second). Remember, an alarm is about an abnormal condition that requires operator action to avoid a consequence. If an operator misses or cannot get to an alarm in time, then the related consequence will occur.
The best overall measure, one that tells the story of your alarm performance to operators, engineers and managers, is the number of alarms per day for a single operating position. At least a month’s worth of data provides a useful picture of alarm system performance.
For an unimproved system, alarm rates of 10,000 to 20,000 per day are unfortunately common. Compare that with long-established guidelines of approximately 150 alarms per day as an acceptable rate and 300 as a maximum manageable rate for a single operator managing a typical process. On an hourly basis, that equates to roughly six to 12 alarms.
Some people may think that is a pretty low rate to shoot for, but consider what those numbers mean. Would you be happy if your control system were operating so poorly that every five minutes the operator had to stop, analyze a situation and take action to avoid a significant consequence? Or would you rather have operators focused on more valuable tasks, such as monitoring and adjusting the process to improve efficiency and profitability?
Unimproved alarm systems are often full of nuisance alarms and other unnecessary alarms that drive these rates higher. Nuisance alarms are solvable, and alarm rationalization can eliminate alarms that do not meet the criteria for being an alarm.
One way to justify an alarm improvement project is to ask a simple question: How many alarms were likely missed by the operator last week because of excessive alarm rates? Assuming the operator saw and responded to every important alarm while ignoring only the less important ones is not a sound strategy for success or safety.
Here’s an example of how alarm analysis can reveal several problems at once.
A distributed control system (DCS) allows operators to suppress a configured alarm. Alarm occurrences can still be saved in the log for analysis, but they are not annunciated to the operator. When the use of suppression is uncontrolled, there may also be limited tracking or visibility into which alarms are suppressed and for how long.
In one system, the alarms actually presented to operators were mostly within the desired range of fewer than 300 per day. Based on that number alone, the alarm system would appear to be performing well.
A deeper analysis told a different story.
It identified 147 tags with almost 500 configured alarms that had been suppressed but were still generating large numbers of occurrences. Over time, operators had dealt with nuisance alarms by suppressing them, but the practice was uncontrolled and included some important alarms.
This is not the way to solve an alarm problem. The findings pointed to broader issues with operating discipline and management of change within the control system. Rationalization must be applied to determine which alarms are truly needed, necessary alarms must be restored and engineering controls should govern the use of alarm suppression.
Average alarm rates can be misleading and do not tell the whole story. Alarm floods are periods of high alarm activity, defined as more than 10 alarms in 10 minutes. During a severe flood, the alarm system can become a nuisance distraction rather than an effective tool, impeding the operator’s ability to manage an upset.
Consider a system that analysis showed was in alarm flood 96% of the time. Imagine trying to solve a major process upset while the alarm system activates every few seconds, sometimes in bursts of dozens of alarms. Instead of helping the operator understand what is happening and take corrective action, the alarm system adds to the workload.
Alarm floods have preceded major industrial incidents, making flood frequency, duration and severity important measures of alarm system performance.
A comprehensive alarm management approach can also include analyses such as:
Most frequent alarms
Stale alarms (that have been in effect continuously for days or weeks)
Alarm priority distribution (compared to best practices)
Alarm flood analysis (duration and quantity)
Breakdown by type of alarm (process value alarms, instrument malfunction alarms, etc.)
Correlated alarms (that always occur close together)
Alarm configuration changes (that should have been made and documented in the management of change (MOC) procedures)
Alarm management standards
Alarm management standards provide recommended performance metrics for monitoring the health of alarm systems. These metrics are important, but they should not be applied without context. Appropriate targets depend on factors such as process type, operator skill, the human machine interface (HMI), degree of automation, operating environment and the types and significance of alarms being generated.
Maximum acceptable alarm rates may be lower or, in some cases, slightly higher depending on the operating environment. Alarm rate alone is not enough to determine whether an alarm system is performing effectively. Organizations should establish and monitor KPIs that reflect the characteristics and risks of their specific operations.
Let's consider a couple of lesser-known but useful analyses. Control systems log a lot more than just alarms. They record a variety of operator actions. If you have Octave Tempo Alarm Management (formerly PAS AlarmManagement) there are many other useful things you can analyze and report automatically. These can give you insight into the challenges faced by your operators.
Controller mode analysis
Process controllers have multiple modes, with the most common being auto and manual. You invested in a controller because you wanted it to run in auto. A mode analysis can show what percentage of the time each controller operates in each mode.
You will likely find controllers that operate in manual mode much of the time. Why? They may perform better in manual than in auto or operators may believe they do. Either way, the analysis has identified an opportunity for improvement.
In one week of data from a single operating position, three controllers each experienced more than 100 mode changes. Is it likely they were designed to operate that way? Probably not.
Operators do not change controller modes without a reason. They perceive, rightly or wrongly, a need to intervene. Frequent mode changes or controllers that spend significant time outside their intended operating mode can therefore indicate control performance issues that warrant investigation.
The objective is not simply to identify the behavior. It is to understand why it is occurring and address the underlying problem so the controllers can operate as designed.
Operator Action Analysis:
DCSs capture all the operator's interactions with the control system. Some of those directly affect the process and others do not. The ones that always do are:
Adjusting a controller setpoint
Changing a controller mode
Directly controlling the output of a controller when placed in manual
Manually initiating an on-off or similar discrete action (using a digital output point)
These actions represent direct manipulation of the process by the operator. A stable, optimally performing process should generally require fewer operator interventions. Tracking how often operators make these changes can therefore provide another indication of control system performance and operator workload.
High rates of operator action may indicate that the process is unstable, controls are not performing as intended or operators are being asked to intervene too frequently. It may also mean there is not enough time to fully consider each change before another action is required.
Looking at these patterns over time can help identify operating positions that may be overloaded and point to areas where control improvements could reduce unnecessary intervention.
Summary
Once nuisance alarm behaviors, such as chattering, have been eliminated and rationalization has removed alarms that do not meet the criteria for being alarms, what remains should represent abnormal situations requiring operator response.
At that point, persistently high rates of valid alarms can indicate that the control system is not keeping the process within appropriate operating boundaries. The answer is no longer simply to improve the alarm system. It is to improve control of the process itself.
Alarm analysis helps reveal where improvement is needed most. With the right software, these analyses and reports can also be automated, giving teams ongoing visibility into alarm system performance and opportunities for improvement.
Ready to tame your alarm system? Contact us — or dive deeper with our white paper: Making a big dent in nuisance alarms.
Review other Taming the Wild Alarm System topics in this 7 part blog series:
Part two - The most important alarm improvement technique in existence
Part three - Silence the noise: fixing chattering and fleeting alarms
Part four - Just how bad is your alarm system?
Part five - What alarm rationalization really uncovers — and why it matters
Part seven - Beyond alarm management – doing more with a powerful tool