NEO AEGIS — Industrial Alarm Management and Intelligent Analytics System

Help operators see the alarms that require action. AEGIS converts raw alarm events into governed priorities, explainable patterns, accountable improvement work, and auditable lifecycle evidence.

AEGIS combines alarm data acquisition, performance monitoring, alarm flood and bad-actor analytics, rationalisation, shelving and suppression governance, management of change, and continuous KPI review into one alarm management lifecycle.

SEE
Real-time load, floods and critical exceptions
UNDERSTAND
Patterns, sequences, context and likely causes
GOVERN
Rationalisation, changes, ownership and audit

The Problem: Alarm Quantity Is Not the Same as Operator Protection

An alarm system can be technically available and still fail to support timely, correct operator response. When nuisance alarms, poor priorities, recurring floods and undocumented changes compete for attention, operators can miss the few conditions that matter.

Failure modeWhat happens
Alarm floodA process upset triggers more alarms than an operator can interpret and act on within the available time.
Bad actorsA small group of tags produces a large share of events through chatter, fleeting behaviour or poor settings.
Priority inflationToo many high-priority alarms weaken differentiation and create inconsistent response expectations.
Standing alarmsLong-duration alarms become part of the background and may represent equipment, process or maintenance risk.
Weak rationalisationCause, consequence, required response, response time and priority basis are incomplete or inconsistent.
Uncontrolled changeSetpoints, priorities, suppression, disabled states and logic changes lack complete approval and audit evidence.

Positioning: A Lifecycle Governance Platform, Not Another Alarm List

AEGIS complements control systems by adding cross-system analytics, engineering governance, collaboration and audit. The DCS or SCADA remains the real-time alarm source and operator interface unless the project explicitly defines otherwise. AEGIS provides lifecycle governance and analytics without weakening established safety and control responsibilities.

Standards Alignment

ReferenceHow it informs the solution
ANSI/ISA-18.2-2016Alarm management lifecycle, philosophy, rationalisation, operation, monitoring and management of change
IEC 62682:2023Lifecycle concepts and requirements for alarm systems in process industries
EEMUA Publication 191Practical design, management, performance and benchmark guidance
OPC AE/DA or UAIntegration patterns where supported and approved by the project architecture

Standards boundary. The cited documents guide solution design; they do not by themselves certify a site or guarantee compliance. Applicable editions, owner standards, alarm philosophy, safety lifecycle, operating procedures and acceptance criteria must be confirmed for each project.


Architecture and Acquisition Boundary

A layered architecture protects control-system boundaries while allowing normalised data, context, analytics, workflows and management reporting to evolve independently.

Acquisition principle. Read-only alarm and event acquisition is the normal starting point. Any write-back, suppression command, shelving control, setpoint change or logic change requires explicit authorisation, engineering design, test evidence, rollback and audit.

【建站提示,做完删掉这段】 这里放方案文档里的 Figure 1 参考架构图。技术买家最先看的就是架构图 —— 它一眼说明你没有动他的控制系统。这是本页优先级最高的一张图。


The Alarm Management Lifecycle

Monitoring is only one phase; sustainable performance requires a closed lifecycle. It begins with the alarm philosophy, identifies candidate alarms, rationalises intent and priority, implements approved logic and documentation, manages daily operation and maintenance, monitors performance, and feeds improvement decisions back into philosophy and design.

【建站提示,做完删掉这段】 这里放 Figure 2 报警管理生命周期图。


Integration: Normalise Sources Without Losing Provenance

Every event retains its source, timestamp, transition, priority, state and relevant system context.

Source systemData purposeIntegration note
DCS / SCADAAlarm transitions, priorities, acknowledgements, suppression and configuration contextVendor-supported or approved interface
HistorianProcess trends, unit state, sequence context and KPI enrichmentRead-only query or interface
MES / productionGrade, campaign, operating mode, batch and production contextProject-defined tags or API
CMMS / EAMMaintenance work, equipment status, failure evidence and action closureReference or synchronised workflow
Identity / directoryUsers, roles, organisation and account lifecycleLeast privilege and approved mapping

Time integrity. Source clocks, time zones, daylight-saving behaviour, acquisition latency, duplicate events, missing transitions and event ordering must be tested before sequence or causal conclusions are accepted.


Analytics and KPI: More Than One Number

A balanced KPI set combines rate, peak demand, persistence, nuisance behaviour, priority quality, response and action completion.

IndicatorQuestion answeredRequired context
Alarm rateHow much demand reaches the operator?Unit state and time window
Peak 10-minute loadWhen does demand exceed practical response capacity?Event sequence
Top contributorsWhich tags dominate load or recurrence?Tag classification
Standing durationWhich conditions persist without resolution?Owner and risk
Priority distributionDoes priority differentiate urgency?Rationalisation basis
Action closureIs improvement work completed and verified?Workflow evidence

Alarm Quality Patterns

PatternDefinition
Alarm floodA defined count or rate is exceeded in a rolling interval; onset, duration, sequence and contributors are analysed.
Chattering alarmRepeated transitions within a short period, often indicating noise, poor deadband or unstable process behaviour.
Fleeting alarmAn alarm appears and clears too quickly to support a meaningful operator response.
Stale alarmAn alarm remains active beyond a defined duration and requires ownership, risk review and action.
Duplicate or consequentialMultiple alarms describe the same condition or arise as predictable consequences of one initiating event.
Bad actorA tag repeatedly contributes disproportionate load, recurrence, standing time or action effort.

Interpretation boundary. A pattern is a diagnostic lead, not automatic proof of poor design. Process conditions, safety intent, equipment state, maintenance and rationalisation evidence must be reviewed before a change is approved.


From Finding to Verified Improvement

Improvement fails when analytics produces a report without an owner, due time, approval path or verification step.

Work elementRequired contentGovernance value
FindingPattern, time range, tags, evidence and operational impactReproducible problem statement
PriorityRisk, operator burden, recurrence, unit criticality and urgencyConsistent work order
OwnerResponsible discipline, reviewer, approver and due timeClear accountability
ActionAnalysis, rationalisation, procedure, maintenance or logic proposalControlled response
VerificationPost-change KPI, recurrence, operator feedback and evidence periodEffectiveness check
EscalationOverdue rules, critical exception and management visibilityTimely intervention

Notification principle. Alerts and reminders are role-based, deduplicated, rate-limited and linked to the exact evidence and action record. Notification volume must not create a second alarm problem.


Change Control and Audit

ControlScope
Rationalisation recordAlarm purpose, cause, consequence, operator action, response time, priority basis and classification
Management of changeRequested change, reason, risk review, affected tags, approvals, test and effective date
Shelving governanceAuthorised users, maximum duration, reason, reminder, extension approval and return-to-service
Suppression by designDocumented state logic, permissives, maintenance mode, testing and visibility
Configuration comparisonDetect and review differences in priority, limits, deadband, delay, class or enable state
Audit trailWho viewed, created, changed, approved, exported or closed an item, and when

Safety boundary. Safety-instrumented functions, trips, interlocks, permissives and protective logic remain subject to their dedicated safety and management-of-change processes. AEGIS does not authorise changes outside approved engineering governance.


Implementation: Baseline First, Govern Changes, Then Scale Analytics

PhaseMain workExit criteria
1. DiscoverConfirm philosophy, systems, sources, units, users, pain points and baselineApproved scope and KPI baseline
2. ConnectAcquire and normalise alarms, events, priorities, states and process contextTraceable and ordered event data
3. DiagnoseConfigure KPI, flood, pattern, standing alarm and bad-actor analysisFindings reproduced by users
4. GovernLaunch rationalisation, actions, shelving, change and audit workflowsGovernance scenarios accepted
5. ImproveVerify changes, review KPIs, tune rules and expand units or sitesImprovement cadence operating

Typical inputs: alarm philosophy, system inventory, tag and priority export, event history, operator feedback, standing alarm list, suppression states, network architecture, role matrix, management-of-change procedure and acceptance baseline.


Deployment and Security

DomainBaseline controlVerification evidence
NetworkZones, conduits, allowlisted flows and controlled remote accessArchitecture and firewall review
IdentityRoles, authentication, account lifecycle and privileged accessRole and access test
DataIntegrity, retention, backup, restore and time synchronisationRecovery and time test
ApplicationSecure configuration, session control, input validation and audit loggingSecurity checklist
IntegrationRead-only default, protocol controls, error handling and rate limitsInterface test record
OperationsMonitoring, patch policy, incident response and support escalationOperating procedure

Project boundary. Exact cybersecurity controls are determined by customer policy, applicable regulation, site architecture, threat assessment, data classification and approved detailed design.


Business Value

ATTENTION
Less avoidable demand on operators
RESPONSE
Clearer priorities and expected actions
CONTROL
Approved, tested and traceable changes
LEARNING
Verified actions and reduced recurrence

Value boundary. Benefits depend on baseline alarm performance, operating discipline, system coverage, action authority, engineering resources, shutdown opportunities and sustained governance. No fixed reduction or financial saving is promised without a customer-specific measurement method.


Request an AEGIS Assessment

Provide the alarm philosophy, representative event history, system and unit scope, known flood or bad-actor examples, current KPI reports, user roles, and the improvement decisions that are hardest to complete. Our project team will prepare a scenario-based review and pilot scope.

Software Consultation
Scroll to Top