Incident, Problem, and Change Management are three distinct ITSM practices. Incident restores service; Problem finds root cause; Change controls how production evolves.
12.3 Incident, Problem, and Change Management
Three practices, three goals
Practice
Goal
Incident
Restore service quickly
Problem
Eliminate recurrence
Change
Keep changes safe and reversible
Incident lifecycle
Step
Activity
Detect
Monitoring + user reports
Log
Ticket with severity + impact
Classify
Category + assignment
Diagnose
Initial triage
Resolve
Restore service (workaround OK)
Close
Confirm with user; capture learning
Severity examples
SEV
Definition
Communication
1
Total outage; war room
CEO informed; status page
2
Major degradation
Senior engineers
3
Minor; standard team
Internal channels
4
Request or low impact
Ticket only
Change categories
Category
Risk
Approval path
Standard
Low
Pre-approved
Normal
Medium
CAB review
Emergency
Critical
Bypass + post-review
Blameless post-mortem template
Section
Content
What happened (timeline)
Hour-by-hour or minute-by-minute
Impact
Users, money, data
Root cause
Technical + organisational
What we did well
Reinforce these
What we should improve
Honest list
Action items
Owner + date
Public summary
If customers affected
Worked example - bank payment outage
Item
Detail
Incident
Payment fails 22 minutes; SEV2
Action
Failover to standby; queue drains in 8 min
Problem
Root cause: connection pool exhaustion
Fix
Pool sizing + circuit breaker via Normal change
Improve
Chaos test; alert on pool > 80 %
Mentor’s tip: Incident restores; Problem prevents; Change controls. Blameless post-mortems convert pain into improvement. Tier your changes; let standard be standard, normal be normal, emergency be rare.
Discussion
Loading…