Taming the wild alarm system - part seven

Security worker watches plant at night. Man controls process on monitors in control room. Engineer looks at industrial manufacture complex. Safety inspector monitors factory, operation with

Beyond alarm management - doing more with a powerful tool

In this blog series, we've mostly discussed the alarm system. It is a small but important part of the overall control system. To accomplish alarm management and comply with standards, we need to continuously analyze alarm system performance and monitor alarm settings for inappropriate changes. We need to document all our alarms, including their causes, consequences and corrective actions and make that valuable information available to the operator.

Doing these things is not complicated and can be highly automated. They involve some good software, a connection to the control system to collect alarm data for analysis, and a master alarm database for detecting and managing alarm change and providing useful information to the operator. Now, this is a powerful toolset! And there is nothing that engineers (and business managers) love more than getting more capability and performance out of a tool that you already have.

Octave’s capabilities have continued to evolve based on the operational challenges customers need to solve. With Octave Tempo Control System Effectiveness (formerly PAS PlantState Integrity) the same infrastructure used for alarm management can support a broader approach to managing control system performance and operational risk.

From alarm management to operational risk management for automation systems

The energy, process, power and similar industries rely on operators, control systems and independent safety systems to manage operational risk. Alarm management is an important part of that strategy, but it is only one part.

A more structured approach brings these elements together so that each layer reinforces the others, building on the same underlying control system data and infrastructure. Let’s look at several areas where this approach can extend beyond alarm management.

Control loop performance monitoring

The control system is at the center. We depend on it to keep our process within designed boundaries to make product safely, efficiently and profitably. But – does it? There are numerous problems throughout industry with control loop basics and the many problems they develop over the years. Problems such as loops not working as designed, loops that must be run sub-optimally in manual, high loop variability, tuning problems, valve hysteresis and stiction, poor control strategies and an industry-wide shortage of capable engineers to deal with all these issues. The real question is not: "How do I justify improving control loop performance?" It is: "Where is the justification for poor control loop performance?"

Octave Tempo Control Loop Performance provides automated monitoring of control loop performance across modern control systems. It continuously analyzes loop behavior across multiple performance categories, applying expert knowledge to help identify and prioritize issues. Automated reports provide engineers with detailed performance metrics and insights, helping control and operations teams diagnose control loop problems and focus improvement efforts where they can have the greatest impact.

High performance HMI

Process control graphics evolved during the transition to digital control systems such as supervisory control and data acquisition (SCADA) and distributed control systems (DCS). Without established guidelines for effective HMI design, many early graphics became crowded with numbers, excessive color and detailed process information that made it difficult for operators to quickly distinguish normal from abnormal conditions.

High Performance HMI applies human-factors principles to address those limitations. Rather than presenting operators with busy process diagrams filled with individual values, effective graphics organize information to emphasize operating conditions, trends, abnormal situations and alarms. This provides context for the hundreds or thousands of measurements an operator may be expected to monitor and helps improve overall situation awareness.

Effective HMI design works hand in hand with alarm management. While alarms alert operators to conditions requiring attention, the HMI provides the context needed to understand what is happening and determine the appropriate response. Together, they help operators recognize abnormal situations earlier and respond more effectively.

Beyond the operator

We've now provided the operator with the tools they need. But major industrial accidents still occur far too frequently. Most such accidents involve operating some part of the process outside the prescribed and designed safe or acceptable boundaries. In many companies, management is highly concerned with verifying at all times that processes remain within such boundaries, including those of safety, environmental, quality, efficiency and profitability. The obvious answer is the automated, continuous monitoring of a plant's conditions relative to such boundaries. The infrastructure from accomplishing alarm management makes this easy.

The important boundaries of a process are typically stored in a hodgepodge of different procedures, design documents and reports. These are often contradictory, out of date or even lost. Control system interlock settings, for instance, have been found to differ from the correct values in design documents. This is often a management of change (MOC) issue that we will get to shortly.

Visualization of boundaries is essential to safe and profitable operations. When optimum ranges are depicted, they can be achieved. Visible safety boundaries become avoidable. Boundary excursions from the optimum may go unnoticed by operators or managers. Processes may run for considerable periods outside desirable ranges. This may be safe, but maximum efficiency and profitability cannot be achieved under such circumstances.

Accomplishing boundary management

A one-time research effort identifies and agrees upon all relevant process boundary information. Sources of the data include process design documents, P&IDs, equipment specifications, process hazard analyses, operating procedures and similar documentation. The correct data is placed into a new section of the existing, secure and controlled master alarm database (MADB), becoming the single source of truth.

Then we use the data connection to continuously monitor the process. This feeds real-time analysis, depiction and automated reporting of key boundary information. Useful, automated reports include:

  • Most frequent boundary excursions

  • Time duration of each excursion

  • Excursions per processing unit and boundary type

  • Excursions ranked by importance and time

  • Automatically calculated financial opportunity cost or losses per excursion.

Once those boundaries are established, continuous monitoring can show when the process approaches or moves outside optimum operating ranges. Teams can track proximity to quality, productivity, safety and environmental boundaries in real time, while automated alerts can draw attention to conditions that require investigation. This gives operators, engineers and managers greater visibility into where the process is operating relative to its intended boundaries.

Management of safety systems and risk

Plants are protected by safety instrumented systems (SISs). Their design involves a complex body of knowledge that supports several international standards, such as IEC61511, Functional Safety – Safety Instrumented Systems for the Process Industry Sector. Much of the work and knowledge in this area revolves around system design rather than operations. But operations is where accidents occur.

These systems are often designed by functional safety experts and then turned over to an operating facility without such expertise. There are many tasks and checks that should be made on safety systems during operation. Safety systems can be subject to the same lack of coherent documentation as was found in documenting process boundaries - no single source of the truth. An operating process may mention one setpoint for safety function activation, a design document specifies another and a maintenance test procedure has yet another. It is common for disparate systems to accumulate errors this way. Safety interlocks are often bypassed (such as for testing) and bypass may be controlled by procedures susceptible to human error.

These factors mean that the effectiveness of the operating SIS can differ significantly from the design. Risks may be well hidden and not obvious.

A safety system requires performance monitoring, maintenance, bypass management, MOC, periodic proof testing and suitability verification. Safety function demand rates and response times assumed during the design phase are expected to be verified by actual operational performance numbers. This task is often overlooked. If an assumption is wrong, the function may be under-designed, over-designed, overly complex or tested more often than needed, wasting money. These tasks require the attention, time and effort of engineers and technicians – and are often accomplished using a variety of inconsistent, error-prone and unreliable methods, such as uncontrolled spreadsheets, notes and manually marked-up drawings and sketches. Management of safety systems often varies by site. Automation can greatly improve the execution of these tasks.

As in boundary management, a one-time documentation research effort consolidates the correct settings for all safety instrumented functions into a new section of the MADB.

These settings are then automatically monitored. Changes are automatically reported to ensure proper MOC. The needed calculations are done automatically. Every safety function activation is analyzed, tracked and a summary report is made automatically. Bypasses are displayed, tracked and controlled. Testing costs can be minimized. Key performance indicators for the safety system are now under control, requiring far less staff effort than in the past.

Risk monitoring

Operational risk increases when a safety function is bypassed, when testing is underway or overdue and when other control system performance factors cause problems. In the past, it has been impractical or even impossible to determine the real-time risk level of a process resulting from combinations of these factors. The plant may be operating at a higher risk level than it was ever intended to. The familiar safety model showing that "the holes in the slices of Swiss cheese are lining up" applies and accidents are more likely to occur.

By monitoring all these factors, the risk profile can be displayed on a dashboard, appear in automatic reporting and generate immediate notifications. Management can make operational decisions taking risk into account.

Automation system management of change

MOC of automation systems is essential. It is easy to dismiss MOC as simply filling out forms and obtaining signatures, so that only authorized people make changes. But that ignores a significant underlying need of the "authorized" people who work with these systems. And that is to know this:

"With my best intentions, is this change I am about to make going to mess something up?"

The answer is not simple. Modern automation systems are exceedingly complex, and their inner workings are not easily examined. A small change in the control system can have major, hidden ramifications. As an example, a control engineer that simply changes a tagname can cause function loss in:

  • Several different process control graphics and trends

  • Compensation calculations for a flowmeter used for billing purposes

  • Logic points or interlock points used for protective purposes

  • Process historization, tracking and inaccurate calculations of efficiency reports in other applications

The engineer needs to learn all of this in advance of a "simple" change, rather than by picking up the pieces afterwards. The problem is compounded when a single site has control systems of several different types, each with their own idiosyncrasies.

What appears to be a simple control system change can therefore have consequences across multiple systems and applications. Understanding those dependencies before making the change is critical. Otherwise, engineers may discover the impact only after functionality has been disrupted.

The challenge becomes even greater at sites with multiple types of control systems, each with its own configurations, relationships and dependencies.

Multi-platform automation system configuration management

Octave Tempo Automation Integrity automatically and regularly imports configuration information from more than 75 types of control systems and connected devices into a single engineering environment. The information is aggregated and given context, helping engineers understand the relationships and dependencies across the automation environment.

What initially appears to be overwhelming complexity reflects how interconnected these systems have become. By making those relationships visible, engineers can better understand what a proposed change may affect before it is implemented.

Connections and references for each component or data entity can be identified, automated management of change (MOC) reports can be generated and the MOC history of individual components or entities can be tracked. Unauthorized changes can also be detected automatically. With greater visibility into these dependencies, engineers can account for the potential impact of planned changes during design and execute changes with greater confidence.

Engineering and managerial visualization

Information from across these areas can also provide engineers and managers with a broader view of operational performance and risk. KPIs related to control loop and alarm performance, process boundary excursions, production, unauthorized changes and safety system status can be monitored together, including across multiple sites.

Bringing these indicators together helps engineering and management teams identify emerging issues, compare performance and better understand where operational attention is needed. This broader visibility can support improvements in safety, production and profitability.

Bringing it all together

Throughout this seven-part series, we’ve explored what it takes to bring an alarm system under control, from understanding how alarm problems develop and measuring performance to establishing an alarm philosophy, rationalizing alarms and uncovering issues that may otherwise remain hidden.

But effective alarm management is not the end of the story. The same infrastructure, data and analysis that improve alarm system performance can provide insight into broader control system and operational challenges. Control loop performance, HMI design, process boundaries, safety systems and configuration changes all influence how effectively operators can understand and respond to what is happening in the plant.

Bringing this information together can provide a more complete view of plant performance and emerging risk, helping operators, engineers and managers identify potential issues earlier and make better-informed decisions. Ultimately, the opportunity is to move beyond managing alarms in isolation toward a broader approach to control system performance and operational risk.

If you missed any installments in the Taming the Wild Alarm System series, explore Parts 1–6 below.

Ready to tame your alarm system? Let's talk about we can help!

Review other Taming the Wild Alarm System topics in this 7 part blog series:

  1. Part one - How did we get in this mess?

  2. Part two - The most important alarm improvement technique in existence

  3. Part three - Silence the noise: fixing chattering and fleeting alarms

  4. Part four - Just how bad is your alarm system?

  5. Part five - What alarm rationalization really uncovers — and why it matters

  6. Part six - Why did they have to call it philosophy?

  7. Part seven - Beyond alarm management – doing more with a powerful tool