Massive incident management
Managing incidents that span many assets — automatic correlation, the incident lifecycle, the operational dashboard and post-incident review.
A Massive Incident is a set of interrelated Operational Issues that share a common cause.
This allows large operational events to be analyzed as a single management object.
Incident Correlation
OHM automatically groups Operational Issues into a Massive Incident when several factors match:
- time of occurrence;
- region;
- Root Cause;
- communication infrastructure;
- software version;
- an external event.
Massive Incident Lifecycle
flowchart TD
A["Detection"] --> B["Correlation"]
B --> C["Massive Incident Created"]
C --> D["Investigation"]
D --> E["Mitigation"]
E --> F["Recovery"]
F --> G["Verification"]
G --> H["Post Incident Review"]
H --> I["Closure"]
Operational Dashboard
For each Massive Incident, the system displays:
- the number of related Operational Issues;
- the number of assets;
- the affected regions;
- the presumed Root Cause;
- the investigation status;
- the SLA;
- the business impact;
- the current recovery progress.
Post Incident Review
After closure, the incident undergoes analysis.
The report includes:
- the timeline;
- the confirmed Root Cause;
- response effectiveness;
- SLA compliance;
- Lessons Learned;
- recommendations for preventing recurrence.
All conclusions can be automatically converted into Knowledge Base articles.
Súvisiace témy
Bola táto stránka užitočná?
Ďakujeme za vašu spätnú väzbu!