NEO AEGIS — Industrial Alarm Management and Intelligent Analytics System
Help operators see the alarms that require action. AEGIS converts raw alarm events into governed priorities, explainable patterns, accountable improvement work, and auditable lifecycle evidence.
AEGIS combines alarm data acquisition, performance monitoring, alarm flood and bad-actor analytics, rationalisation, shelving and suppression governance, management of change, and continuous KPI review into one alarm management lifecycle.
| SEE Real-time load, floods and critical exceptions | UNDERSTAND Patterns, sequences, context and likely causes | GOVERN Rationalisation, changes, ownership and audit |
The Problem: Alarm Quantity Is Not the Same as Operator Protection
An alarm system can be technically available and still fail to support timely, correct operator response. When nuisance alarms, poor priorities, recurring floods and undocumented changes compete for attention, operators can miss the few conditions that matter.
| Failure mode | What happens |
|---|---|
| Alarm flood | A process upset triggers more alarms than an operator can interpret and act on within the available time. |
| Bad actors | A small group of tags produces a large share of events through chatter, fleeting behaviour or poor settings. |
| Priority inflation | Too many high-priority alarms weaken differentiation and create inconsistent response expectations. |
| Standing alarms | Long-duration alarms become part of the background and may represent equipment, process or maintenance risk. |
| Weak rationalisation | Cause, consequence, required response, response time and priority basis are incomplete or inconsistent. |
| Uncontrolled change | Setpoints, priorities, suppression, disabled states and logic changes lack complete approval and audit evidence. |
Positioning: A Lifecycle Governance Platform, Not Another Alarm List
AEGIS complements control systems by adding cross-system analytics, engineering governance, collaboration and audit. The DCS or SCADA remains the real-time alarm source and operator interface unless the project explicitly defines otherwise. AEGIS provides lifecycle governance and analytics without weakening established safety and control responsibilities.
Standards Alignment
| Reference | How it informs the solution |
|---|---|
| ANSI/ISA-18.2-2016 | Alarm management lifecycle, philosophy, rationalisation, operation, monitoring and management of change |
| IEC 62682:2023 | Lifecycle concepts and requirements for alarm systems in process industries |
| EEMUA Publication 191 | Practical design, management, performance and benchmark guidance |
| OPC AE/DA or UA | Integration patterns where supported and approved by the project architecture |
Standards boundary. The cited documents guide solution design; they do not by themselves certify a site or guarantee compliance. Applicable editions, owner standards, alarm philosophy, safety lifecycle, operating procedures and acceptance criteria must be confirmed for each project.
Architecture and Acquisition Boundary
A layered architecture protects control-system boundaries while allowing normalised data, context, analytics, workflows and management reporting to evolve independently.
Acquisition principle. Read-only alarm and event acquisition is the normal starting point. Any write-back, suppression command, shelving control, setpoint change or logic change requires explicit authorisation, engineering design, test evidence, rollback and audit.
【建站提示,做完删掉这段】 这里放方案文档里的 Figure 1 参考架构图。技术买家最先看的就是架构图 —— 它一眼说明你没有动他的控制系统。这是本页优先级最高的一张图。
The Alarm Management Lifecycle
Monitoring is only one phase; sustainable performance requires a closed lifecycle. It begins with the alarm philosophy, identifies candidate alarms, rationalises intent and priority, implements approved logic and documentation, manages daily operation and maintenance, monitors performance, and feeds improvement decisions back into philosophy and design.
【建站提示,做完删掉这段】 这里放 Figure 2 报警管理生命周期图。
Integration: Normalise Sources Without Losing Provenance
Every event retains its source, timestamp, transition, priority, state and relevant system context.
| Source system | Data purpose | Integration note |
|---|---|---|
| DCS / SCADA | Alarm transitions, priorities, acknowledgements, suppression and configuration context | Vendor-supported or approved interface |
| Historian | Process trends, unit state, sequence context and KPI enrichment | Read-only query or interface |
| MES / production | Grade, campaign, operating mode, batch and production context | Project-defined tags or API |
| CMMS / EAM | Maintenance work, equipment status, failure evidence and action closure | Reference or synchronised workflow |
| Identity / directory | Users, roles, organisation and account lifecycle | Least privilege and approved mapping |
Time integrity. Source clocks, time zones, daylight-saving behaviour, acquisition latency, duplicate events, missing transitions and event ordering must be tested before sequence or causal conclusions are accepted.
Analytics and KPI: More Than One Number
A balanced KPI set combines rate, peak demand, persistence, nuisance behaviour, priority quality, response and action completion.
| Indicator | Question answered | Required context |
|---|---|---|
| Alarm rate | How much demand reaches the operator? | Unit state and time window |
| Peak 10-minute load | When does demand exceed practical response capacity? | Event sequence |
| Top contributors | Which tags dominate load or recurrence? | Tag classification |
| Standing duration | Which conditions persist without resolution? | Owner and risk |
| Priority distribution | Does priority differentiate urgency? | Rationalisation basis |
| Action closure | Is improvement work completed and verified? | Workflow evidence |
Alarm Quality Patterns
| Pattern | Definition |
|---|---|
| Alarm flood | A defined count or rate is exceeded in a rolling interval; onset, duration, sequence and contributors are analysed. |
| Chattering alarm | Repeated transitions within a short period, often indicating noise, poor deadband or unstable process behaviour. |
| Fleeting alarm | An alarm appears and clears too quickly to support a meaningful operator response. |
| Stale alarm | An alarm remains active beyond a defined duration and requires ownership, risk review and action. |
| Duplicate or consequential | Multiple alarms describe the same condition or arise as predictable consequences of one initiating event. |
| Bad actor | A tag repeatedly contributes disproportionate load, recurrence, standing time or action effort. |
Interpretation boundary. A pattern is a diagnostic lead, not automatic proof of poor design. Process conditions, safety intent, equipment state, maintenance and rationalisation evidence must be reviewed before a change is approved.
From Finding to Verified Improvement
Improvement fails when analytics produces a report without an owner, due time, approval path or verification step.
| Work element | Required content | Governance value |
|---|---|---|
| Finding | Pattern, time range, tags, evidence and operational impact | Reproducible problem statement |
| Priority | Risk, operator burden, recurrence, unit criticality and urgency | Consistent work order |
| Owner | Responsible discipline, reviewer, approver and due time | Clear accountability |
| Action | Analysis, rationalisation, procedure, maintenance or logic proposal | Controlled response |
| Verification | Post-change KPI, recurrence, operator feedback and evidence period | Effectiveness check |
| Escalation | Overdue rules, critical exception and management visibility | Timely intervention |
Notification principle. Alerts and reminders are role-based, deduplicated, rate-limited and linked to the exact evidence and action record. Notification volume must not create a second alarm problem.
Change Control and Audit
| Control | Scope |
|---|---|
| Rationalisation record | Alarm purpose, cause, consequence, operator action, response time, priority basis and classification |
| Management of change | Requested change, reason, risk review, affected tags, approvals, test and effective date |
| Shelving governance | Authorised users, maximum duration, reason, reminder, extension approval and return-to-service |
| Suppression by design | Documented state logic, permissives, maintenance mode, testing and visibility |
| Configuration comparison | Detect and review differences in priority, limits, deadband, delay, class or enable state |
| Audit trail | Who viewed, created, changed, approved, exported or closed an item, and when |
Safety boundary. Safety-instrumented functions, trips, interlocks, permissives and protective logic remain subject to their dedicated safety and management-of-change processes. AEGIS does not authorise changes outside approved engineering governance.
Implementation: Baseline First, Govern Changes, Then Scale Analytics
| Phase | Main work | Exit criteria |
|---|---|---|
| 1. Discover | Confirm philosophy, systems, sources, units, users, pain points and baseline | Approved scope and KPI baseline |
| 2. Connect | Acquire and normalise alarms, events, priorities, states and process context | Traceable and ordered event data |
| 3. Diagnose | Configure KPI, flood, pattern, standing alarm and bad-actor analysis | Findings reproduced by users |
| 4. Govern | Launch rationalisation, actions, shelving, change and audit workflows | Governance scenarios accepted |
| 5. Improve | Verify changes, review KPIs, tune rules and expand units or sites | Improvement cadence operating |
Typical inputs: alarm philosophy, system inventory, tag and priority export, event history, operator feedback, standing alarm list, suppression states, network architecture, role matrix, management-of-change procedure and acceptance baseline.
Deployment and Security
| Domain | Baseline control | Verification evidence |
|---|---|---|
| Network | Zones, conduits, allowlisted flows and controlled remote access | Architecture and firewall review |
| Identity | Roles, authentication, account lifecycle and privileged access | Role and access test |
| Data | Integrity, retention, backup, restore and time synchronisation | Recovery and time test |
| Application | Secure configuration, session control, input validation and audit logging | Security checklist |
| Integration | Read-only default, protocol controls, error handling and rate limits | Interface test record |
| Operations | Monitoring, patch policy, incident response and support escalation | Operating procedure |
Project boundary. Exact cybersecurity controls are determined by customer policy, applicable regulation, site architecture, threat assessment, data classification and approved detailed design.
Business Value
| ATTENTION Less avoidable demand on operators | RESPONSE Clearer priorities and expected actions | CONTROL Approved, tested and traceable changes | LEARNING Verified actions and reduced recurrence |
Value boundary. Benefits depend on baseline alarm performance, operating discipline, system coverage, action authority, engineering resources, shutdown opportunities and sustained governance. No fixed reduction or financial saving is promised without a customer-specific measurement method.
Request an AEGIS Assessment
Provide the alarm philosophy, representative event history, system and unit scope, known flood or bad-actor examples, current KPI reports, user roles, and the improvement decisions that are hardest to complete. Our project team will prepare a scenario-based review and pilot scope.
