About this run
Run
- Id
- ca115d77-0be6-47f1-a517-8ac1272dcad1
- Prompt
- Incident management process
- Configuration
- cheap 3x2
- Loaded from
- slow-thinker.keys.json ← slow-thinker.base.json ← slow-thinker.cheap.json
- Made with
- Process executor 0.1.2
- This page
- Report generator 0.1.16 · Analysis generator 0.1.2
- Status
- completed
- Started
- 2026-09-21 00:37
- Duration
- 7 min 45 s
Council
- Proposers
- 3
- Refinement rounds
- 2
- Voters
- 3
- Selected plan
- deepseek-flash_refine_2 · 2 of 3 votes · 24 steps
Plan cost
Analysis LLM analysis
Models of the council
| Model | Thinking |
|---|---|
| claudeHaiku4.5 · anthropic/claude-haiku-4-5 | extended thinking, 16.0k tokens · temp 1 |
| deepseek-flash · deepseek/deepseek-flash | thinking on, effort high |
| qwen3.8-flash · alibaba/qwen3.8-flash | thinking on, budget 16.0k tokens |
Analyses of this run
| # | Date | Analyst | Schema | Analysis generator | Calls | Cost | Verdict on the vote | Report |
|---|---|---|---|---|---|---|---|---|
| 1 | 2026-09-21 00:51 | anthropic/claude-opus-5 |
v5 | 0.1.2 slow-thinker 0.1.0 |
7 | 2.26 USD | agrees | this page |
Task
A B2B payments platform in New York (260 engineers in 28 teams, 2,100 customers, $4B processed a month, a 99.95% SLA with service credits) runs about 180 services on Kubernetes across two AWS regions, with a shared PostgreSQL cluster for the ledger. In the last 12 months: 31 customer-impacting incidents, median time to detect 22 minutes (customers detected 40% of them first), median time to mitigate 3 h 10 min, $1.3M paid in SLA credits, two incidents in which nobody was sure who was in charge for over an hour, and a CEO email about "outages we hear about from clients". On-call exists in 12 of the 28 teams, unpaid, fed by alerts from six different tools — 3,400 alerts a month, 85% of them noise; there is no severity scale, status-page updates are written by whoever is around, and postmortems happen for some incidents, in various formats, with action items that are rarely tracked (11 of 64 closed). A SOC 2 Type II audit in eight months will test incident response. Engineers push back against "carrying a pager for other teams' code".
Define the incident management process: severity levels and what each one triggers; roles (incident commander, communications, scribe, subject-matter responders) and how they are staffed 24x7 across 28 teams; the on-call structure, rotations, compensation and the rules for alert quality; detection and escalation paths; internal and customer communications (status page, account managers, regulators when required) with their timings; postmortems (when mandatory, format, blameless review, ownership and tracking of actions); the metrics and reviews that show whether it works; and how it is introduced across the teams without waiting for the audit.
Process
Round 0 — initial proposals
All three agents converge on the same skeleton — severity matrix → roles → paid on-call → tool consolidation → comms SLAs → blameless postmortems with tracked actions → metrics → phased rollout — but differ sharply in depth and in the on-call model. P2 is the most operationally concrete (baseline evidence pack, SLO-based detection, federated ownership, SOC 2 control mapping plus dry run); P1 is exhaustive but serializes training and drills after full rollout; P3 is the compact version and leaves audit evidence and detection engineering thin.
What the proposals share
- Four-tier severity scale as the keystone: every plan keys paging, comms cadence and postmortem obligation off SEV1–SEV4 (P1 S1, P2 S3, P3 S2).
- Same role set — IC who commands but does not debug, comms lead, scribe, SME responders — with explicit decision rights (P1 S2, P2 S4, P3 S4).
- Collapse the six alerting tools into one platform and attack the 85% noise with actionable-alert standards, dedup and per-team caps (P1 S5–S6, P2 S7/S14, P3 S5).
- Paid on-call plus mandatory blameless postmortems for SEV1/SEV2 with action items tracked in Jira and reviewed by leadership; pilot first, then waves, not a big bang (P1 S4/S12/S13/S19, P2 S9/S12/S13/S18, P3 S3/S8/S9/S11).
What separates them
- On-call architecture: P2 S8 is federated — one owning team per service, "own your code, own your pager", minimum rotation of six, 16 teams must build a rotation or transfer ownership, plus a central 24x7 IC roster. P3 S3 does the opposite, merging 12 rotations into one unified pool covering all 28 teams, which directly recreates the "pager for other teams' code" grievance it claims to solve. P1 S3 sits in between: dedicated IC pool of 4–6 plus per-team SME on-call.
- Detection engineering: only P2 S6 defines SLIs/SLOs for the top 20 customer journeys, symptom-based alerting, external synthetics in three locations, ledger-specific PostgreSQL signals (replication lag, TXID exhaustion) and a separate blast-radius resilience track. P3 S6 has synthetics only; P1 treats detection mostly as alert routing and never defines SLOs.
- SOC 2 rigour: P2 devotes S19–S20 to Trust Services Criteria mapping, a named owner and evidence artifact per control, retention rules and a dry run six weeks out. P1 S16 has a control mapping plus a month-6 mock audit. P3 has one bullet in S11 ("simulate auditor questions") — far too light for an eight-month Type II window.
- Ordering flaws differ: P1 S17 depends on all sixteen prior steps and pushes training (S18), rollout (S19) and the first drill (S20) to the end, so nobody practises before going live. P2 front-loads a 12-month baseline register (S2) that every metric later hangs on. P3 S11 defers comp and role rollout to months 3–4 while alert cleanup starts month 1, and P3 S9 ties action-item completion to performance reviews, which cuts against its own blameless charter (S8).
The calls of this round
21 steps covering severity tiers, roles, a hybrid IC-pool/per-team SME rotation, concrete comp numbers ($500–1,000/week stipend, 1.5x callback, comp day), single alert tool with suppression rules, per-severity playbooks, status-page timings (SEV1 in 3 minutes), postmortem and action-item tracking, metrics and governance cadence. Rollout is a single mega-step (S17) depending on all sixteen predecessors, followed by training, waves and drills. Targets are the most aggressive of the three: MTTD <8 min, credits <$100k, 0 findings.
The plan produced
1. Define severity levels and decision criteria
Create a four-tier severity framework (SEV-1 through SEV-4) that guides all downstream decisions about response, escalation, and communications.
Each level must specify: customer impact (revenue at risk, customers affected, data loss risk); financial threshold triggering service credits; whether an incident commander is required; response time SLA (e.g., SEV-1 < 5 min notification, SEV-4 < 2 hours); and the go-live decision tree (when to declare and when to resolve).
- SEV-1: Complete service down or critical path broken for >5% of customers; every minute costs money; IC required; 99.99% uptime threatened
- SEV-2: Significant degradation, features unavailable, affecting 1–5% of customers; IC typically required
- SEV-3: Minor impact, limited customer footprint or workaround exists; escalation path but not automatic IC
- SEV-4: Observations or minor issues; alert-driven, no escalation unless pattern emerges
2. Define incident roles and responsibilities (after 1)
Create the organizational roles that operate during an incident: who is in charge, who talks to customers, who writes down what happened, who fixes the system, and how decisions are made under pressure.
Each role must have a single-sentence mission, decision authority, and escalation upward.
- Incident Commander: owns decision-making and timeline; declares severity; resolves conflicts; may or may not be technical
- Deputy IC: shadow to IC and takes over if IC becomes unavailable
- Communications Lead: writes status page, notifies account teams, manages customer perception
- Scribe: records decisions, who did what, key timestamps; not responsible for fixing
- SME Responders: engineers with context on the failing service(s); take IC's direction without debate
3. Design 24x7 on-call rotation structure (after 2)
Build a rotation model that covers all 28 teams with primary and backup on-call engineers every hour across weekdays, evenings, weekends, and holidays; addresses the pager-carrying resistance.
Key design decisions: Is coverage per-team (each team owns its services) or pooled (shared responder pool handles anything)? How many people per rotation? How long are shifts (one week, two weeks)? When can engineers opt out without leaving the team exposed? Which roles are on-call (IC, communications, SME)?
- Recommend: dedicated IC pool (4–6 people in fast rotation) + per-team SME on-call for each team's own services
- Recommend: two-week rotation blocks to reduce handoff friction
- Recommend: one primary, one secondary per slot; secondary handles during primary's escalation
- Provide swaps, blackout dates, and a rule that no engineer is on-call more than 2 weeks per quarter
4. Define on-call compensation and incentives (after 3)
Create a pay model that makes on-call acceptable and rewards engineers who carry the pager; ties compensation to real business risk.
- Base on-call stipend: e.g., $500–1,000 per week while on-call (regardless of incidents)
- Callback pay: 1.5× hourly rate for time spent mitigating incidents during off-hours
- Incident bonus: $50–100 extra per SEV-1 or SEV-2 incident mitigated (recognition)
- Comp time: full business day off after an incident that required >2 hours mitigation during night/weekend
- Annual bonus tie-in: 10–20% bonus multiplier for flawless on-call reviews
- Communicate: position as investment in reliability, not punishment for being online
5. Consolidate alert routing infrastructure
Replace six alert tools with a single ingestion and routing system; stop engineers from being woken by duplicate alerts, and make escalation automated instead of manual.
Evaluate existing tools (likely candidates: PagerDuty, Opsgenie, or Incident.io) or build a lightweight wrapper. The system must: accept alerts from all 180 services; deduplicate and correlate (same outage, different monitoring source); route to correct on-call engineer; expose an API for playbook automation; log every alert for postmortem analysis.
- Choose tool by week 2 of S5
- Migrate alerting endpoints from 6 sources to 1 by week 4
- Set up audit trail and retention
- Ensure mobile app works (on-call engineers need to engage from phone)
6. Build alert quality rules to cut noise (after 1, 5)
Implement rules that automatically suppress the 85% of alerts that are noise (flapping, transient errors, auto-recovered conditions). Target: <400 actionable alerts per month.
Rules to implement: suppress alerts if service auto-recovered within 30 seconds; deduplicate same alert from multiple monitoring sources; suppress alerts for known maintenance windows; group flapping alerts (same service, >5 occurrences in 2 minutes) into a single page to on-call; rate-limit alerts from noisy services (e.g., max 1 alert per 5 minutes per service until silence clears).
- Audit existing 3,400 alerts per month: which are true signals, which are noise
- Tag each alert source with severity level (S1, S2, S3, S4 from S1)
- Create exceptions list: services known to be noisy, require different rules
- Weekly review: alert teams that trigger >50 alerts per week for reduction strategies
7. Implement automated detection and escalation paths (after 1, 5, 6)
Wire the alert system to automatically escalate based on time or severity; removes the need for manual judgment calls during chaos.
Logic: SEV-1 alert arrives → IC notified instantly via phone call + SMS + Slack + mobile; if IC does not acknowledge within 2 minutes, page deputy IC; Communications Lead pinged simultaneously. SEV-2: on-call SME for that service + IC notify via Slack and mobile, escalate to IC's manager if not acknowledged in 10 min. SEV-3/4: on-call SME only, escalate after 30 min.
- Implement in alert routing system (S5)
- Test all paths weekly via synthetic page to on-call
- Track escalation metrics: how many pages reach secondary, how many hit manager
- Adjust timing based on first month of operations
8. Build incident dashboard and status tracking (after 5)
Create a single source of truth during an incident that every responder sees in real time: who is on-call, incident timeline, who said what, current status, next steps.
Dashboard displays: active incidents and their severity; who is the IC and communications lead; timeline of all events (alert fired, IC assigned, customer notified, mitigation started, resolved); Slack channel and mobile notification status; on-call rosters (who is on-call right now for each team); postmortem link as soon as incident closes.
- Integrate with alert tool (S5) to auto-populate incident creation and initial severity
- Push updates to status page and customer account managers automatically
- Log all timeline entries for audit and postmortem completeness
- Mobile-optimized so IC can work from any device
9. Write incident playbooks for each severity (after 1, 2, 3)
Create a one-page (or one-screen) reference for the IC and SMEs during an incident; sequences the steps and removes ambiguity.
Each severity level gets its own playbook: who gets paged (roles, order); first questions to ask (is it real, how big, who knows); what the IC should declare in first message (status page text, account manager notification, regulatory trigger); how long before escalating to executive team; decision rules for going dark vs. continuing to update customers.
- SEV-1 playbook: immediate IC + comms + CTO notification; customer status every 5 minutes
- SEV-2 playbook: IC + comms + tech lead notification; status every 15 minutes
- SEV-3 playbook: on-call SME + comms if customer-visible; status every 30 min or as resolved
- SEV-4 playbook: on-call SME only; update customers only if promised SLA is at risk
- Include decision trees: is this SEV-1 or SEV-2? Is it our code or dependency? Escalate or containment?
10. Define internal communication workflows (after 2, 3)
Specify who informs whom, in what order, via what channel (call, Slack, email) during an incident; prevents gaps like "nobody knew who was in charge for an hour."
Workflow for SEV-1: IC assigned → IC calls CTO/VP Eng and incident channel lead within 1 minute; incident declared in #incidents Slack channel with severity, IC name, service affected; SME on-call for that service joins call automatically; IC pushes updates to #incidents every 5 minutes or when material change occurs. For SEV-2: IC notifies team leads via Slack, updates #incidents every 15 min. Define escalation: if IC is unreachable, deputy IC takes over and announces it.
- Create a phone tree or on-call list accessible to responders
- Set expectations: "If you don't hear from IC in 2 minutes, call them"
- Use a single incident Slack channel per incident (auto-created by incident tool)
- Log all comms in the incident dashboard for postmortem review
11. Design customer communication and status page process (after 1, 2)
Plan when and how to inform customers, account managers, and regulators; ensure 2,100 customers are not learning about outages from Twitter before you tell them.
Rules by severity: SEV-1 detected → status page updated within 3 minutes (even if root cause unknown; post "Investigating"); account managers of affected customers called within 5 minutes; regulatory notification (if payment processing down) queued for approval; customer email within 10 minutes with ETA for next update. SEV-2: status page within 10 min, account managers called within 15 min, email if affecting >10 customers. SEV-3/4: no customer communication unless SLA at risk.
- Empower Communications Lead to update status page without IC approval if delay >3 min
- Prepare templated messages for common scenarios (database failover, data pipeline stuck, service crashed)
- Route regulatory notifications through legal/compliance; don't wait for perfect root cause
- Track customer impact in real time: how many customers affected by severity
12. Establish blameless postmortem process and format (after 1, 2)
Build a systematic way to learn from incidents so the same failure does not happen twice; counter the fear that admitting a mistake leads to being blamed.
Mandatory postmortems: all SEV-1 and SEV-2 incidents, within 48 hours of resolution. Optional but encouraged: SEV-3 if interesting or if >3 of same type in 30 days. Format: what was the user-visible impact and for how long; what was the root cause (not "human error" but the system condition that made error possible); timeline of discovery and response; action items with owner and deadline; blameless tone (focus on process and system design, not individual mistakes).
- Assign a facilitator (not the on-call IC) to run postmortem
- Attendees: IC, comms lead, SMEs involved, team lead, customer success if customer-facing
- Write postmortem in shared doc; make it findable (searchable, linked from incident)
- No discussion of "who screwed up"; only "why did the system allow this to happen"
13. Build action item tracking and accountability (after 12)
Create a system that tracks postmortem action items so they are not forgotten; currently 11 of 64 (17%) are being tracked, leaving 53 unfinished improvements.
System: each postmortem generates action items (e.g., "add monitoring for X," "update runbook for Y," "write test for Z"). Each item gets: clear description, owner (engineer's name), due date (1–4 weeks based on priority), severity (critical = must do before similar incident happens again; important = improve next month; nice-to-have = backlog). Action items live in a dedicated Jira project visible to all teams; owners are accountable (their manager reviews quarterly). Weekly: incident commander reviews open items due that week. Monthly: each team's postmortem items reviewed in their standup.
- Export action items from postmortem document to tracking system automatically
- Require IC to sign off that an action is complete before closing
- Report on completion rate as a metric (target: >80% by month 3)
14. Define incident metrics and KPIs
Establish what "good" looks like; measure so you can improve. Target metrics for 12 months out: mean time to detect 8 minutes (vs. 22 now), customers detect first <5% of incidents (vs. 40%), MTTR 45 minutes for SEV-1 (vs. 190), SLA credits <$100k/year.
Metrics to track: (1) MTTD = time from incident start to first alert/report; disaggregate: external report vs. internal detection. (2) MTTR = time from first report to full mitigation; track by severity and by service. (3) Customer-reported incidents per month (should drop to <2 per month). (4) Alert signal-to-noise ratio (goal: <5% false positive after S6 rules). (5) On-call satisfaction (survey: would you do this again?). (6) Postmortem action completion rate. (7) Incident commander and responder utilization (hours per week per person).
- Dashboard: auto-populated from incident tool, updated daily
- Disaggregate by team and service: which teams have bad MTTR? Which service is most incident-prone?
15. Create review cadence and governance process (after 14)
Establish regular rhythm to inspect the metrics, spot trends, and adjust the process itself; prevent the system from calcifying.
Weekly: incident commander and on-call lead review prior week—number of incidents, any escalations, any communication gaps. Monthly: director-level incident review—trends by service, top causes of incidents, action item status, whether severity classification is working. Quarterly: full leadership review—MTTD, MTTR, customer impact, SLA credit spend, on-call satisfaction score, any systemic changes needed. Annually: audit the entire process for SOC 2 compliance.
- Assign meeting owners: weekly = on-call lead; monthly = director of reliability; quarterly = VP Eng + CFO (SLA cost) + customer success
- Use same data dashboard (S14) for all reviews
- Publish a monthly "incident newsletter" to all engineers: what happened, what we learned, what's improving
16. Prepare SOC 2 Type II audit checklist (after 1, 2, 9, 12, 13, 14, 15)
Document that the incident management system meets the control requirements for a SOC 2 audit; audit happens in 8 months, so this work builds confidence in coverage.
Audit will test: (1) Is there a defined incident response process? (2) Are roles and responsibilities clear? (3) Are incidents logged and tracked? (4) Is root cause analysis performed? (5) Are action items tracked and completed? (6) Is on-call staffing adequate? (7) Are communications timely? (8) Are postmortems documented and blameless? Create a control mapping document that links each SOC 2 requirement to your process (S1–S15). Collect evidence: incident logs, postmortem documents, action item tickets, metrics reports, training records.
- Designate a compliance owner (often a reliability lead or security engineer)
- Run a mock audit at month 6 to identify gaps
- Ensure all postmortems and incidents are retained and searchable for auditor review
17. Develop implementation and rollout plan (after 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16)
Create a phased timeline to roll out the incident management system across all 28 teams; avoids big-bang failure and builds credibility.
Recommended structure: Phase 1 (weeks 1–4): build and test infrastructure (S5, S8, alerting); deploy severity levels and roles (S1, S2); pick pilot teams (2–3 high-traffic teams). Phase 2 (weeks 5–12): train pilot teams, run incident drills, refine playbooks based on learning; expand to half of remaining teams. Phase 3 (weeks 13–20): full rollout to all 28 teams; continue drills; track metrics. Phase 4 (weeks 21–28): stabilize, iterate on metrics, prepare for audit.
- Assign a release manager to coordinate across teams
- Create a detailed Gantt chart with swim lanes (infra, process, training, rollout)
- Identify risks: competing priorities, engineers worried about pager burden, tool adoption friction
- Plan stakeholder engagement: weekly updates to eng leadership, monthly town halls for all engineers
18. Build training and documentation (after 17)
Create role-specific education so engineers understand the new system and are confident executing during an incident.
Training tracks: (1) All engineers: 30-min video on severity levels, communication expectations, postmortem format, action items (mandatory); (2) On-call responders: 2-hour workshop on playbooks, alert tool, escalation paths, when to call manager, case studies of real incidents, mobile app walkthrough; (3) Incident commanders: 4-hour IC bootcamp covering leadership under pressure, decision-making, communicating with executives, status page discipline, postmortem facilitation, practiced drills; (4) Communications leads: templates, when to update, how to talk to customers, regulatory notification rules.
- Record videos so async teams can learn on their schedule
- Create runbooks and quick-reference cards for each role (print + digital)
- Pair new on-call engineers with experienced responder for first week
- Require IC certification before anyone joins IC rotation (pass a practical drill)
19. Execute staged rollout across teams (after 18)
Progressively activate the incident management system with feedback loops at each stage; reduces risk of system-wide failure.
Wave 1 (week 6–8): 3–4 pilot teams begin on-call rotations and incident response using new system; capture feedback daily. Wave 2 (week 10–14): 8–10 additional teams, incorporating lessons from Wave 1; ensure diversity of team types (payment processing, monitoring, data pipeline, auth, etc.). Wave 3 (week 15–20): remaining teams; by now, the system is proven and less hand-holding needed.
- Daily retros with Wave 1 teams: what worked, what was confusing, what broke
- Each wave produces a "lessons learned" document that informs the next
- Track adoption metrics: how many incidents reported per team, alert quality, MTTD/MTTR
- Address resistance: engineers who are skeptical of the system, on-call burden, tool friction; assign a "change champion" in each team
20. Run incident response drills and simulations (after 19)
Practice incidents in a controlled setting so responders gain confidence and gaps are found before a real crisis.
Drill 1 (week 8, pilot teams only): simulated database failover; test alert routing, IC assignment, on-call notification, status page update, postmortem. Drill 2 (week 12, all teams): full incident spanning multiple services; test escalation, communications, decision-making. Drill 3 (week 18): chaos drill with intentional confusion—what if IC is unreachable, what if comms lead is on PTO, what if an alert tool goes down? Quarterly thereafter: repeat drills with new on-call engineers.
- Schedule drills at a time when key people are available (not Friday afternoon)
- Run drill with live observers (reliability lead, manager) who provide feedback
- Measure drill success: all steps completed on time, communications clear, postmortem written
- Debrief immediately: what did we learn, what do we change?
21. Establish continuous improvement process (after 15, 20)
Build a feedback loop so the incident management system evolves based on what you learn.
Monthly improvement cycle: in the monthly incident review (S15), leadership identifies one or two friction points from recent incidents (e.g., "alert tool response is slow," "postmortem meetings are boring," "action items keep slipping"). Assign an owner to investigate and propose a fix (change to playbook, tweak to tool config, additional training, process change). Test the fix with pilot team, gather feedback, roll out if successful. Annual review: can we reduce on-call burden, simplify communication, improve MTTD further? Do any teams need special handling (new services, difficult customers)?
- Keep a public backlog of improvement ideas (Jira, shared doc) so any engineer can suggest
- Quarterly all-hands: share wins ("MTTD improved 50%," "on-call satisfaction up," "zero customer-detected incidents") to build momentum
- Link improvements back to metrics (S14): show that changes actually worked
- Median time to detect reduced from 22 minutes to <8 minutes within 6 months
- Customer-detected incidents drop from 40% to <5% within 6 months
- Median time to mitigate (MTTR) reduced from 3h 10min to <45 minutes for SEV-1 incidents within 6 months
- Annual SLA credits decrease from $1.3M to <$100k within 12 months
- Alert noise reduced from 3,400 per month (85% false positive) to <400 per month (>95% signal) within 3 months
- Zero incidents with command-and-control ambiguity (>1 hour without clear IC) within 2 months
- Postmortem action item completion rate reaches >80% (from 17%) within 4 months
- On-call satisfaction score (survey) reaches >7/10 for on-call engineers within 3 months
- All 28 teams integrated into incident management system with active on-call rotations by week 20
- SOC 2 Type II audit passes incident response controls with no findings 8 months from start
- Incident commander certification: 100% of active ICs trained and drilled within 2 months
- Monthly incident review meeting established and attended by leadership; trends documented
- New incident system integration complete: single alert tool, single dashboard, all 180 services feeding in, <5 min deployment
[SYSTEM] You are an expert assistant in complex project planning. Your task is to generate a detailed and structured action plan to reach the main objective in the most professional and most detailed way, no matter how much work or steps will be needed to perform. Use your internal reasoning processes to think deeply about the problem and create the most comprehensive plan possible. Take as much time and space as you need to think through all aspects of the problem. After your thorough analysis, answer with the plan in the requested structure. Every text field you write will be read by a busy person, so write for them: short sentences, one idea each, in short paragraphs. Text fields accept Markdown: separate paragraphs with a blank line, use a bullet list when you enumerate things and plain paragraphs when you explain or argue, and bold at most one key phrase per paragraph. A text of more than three sentences must be split into paragraphs; never deliver a single unbroken block of text. [HUMAN] Main Objective: "A B2B payments platform in New York (260 engineers in 28 teams, 2,100 customers, $4B processed a month, a 99.95% SLA with service credits) runs about 180 services on Kubernetes across two AWS regions, with a shared PostgreSQL cluster for the ledger. In the last 12 months: 31 customer-impacting incidents, median time to detect 22 minutes (customers detected 40% of them first), median time to mitigate 3 h 10 min, $1.3M paid in SLA credits, two incidents in which nobody was sure who was in charge for over an hour, and a CEO email about "outages we hear about from clients". On-call exists in 12 of the 28 teams, unpaid, fed by alerts from six different tools — 3,400 alerts a month, 85% of them noise; there is no severity scale, status-page updates are written by whoever is around, and postmortems happen for some incidents, in various formats, with action items that are rarely tracked (11 of 64 closed). A SOC 2 Type II audit in eight months will test incident response. Engineers push back against "carrying a pager for other teams' code". Define the incident management process: severity levels and what each one triggers; roles (incident commander, communications, scribe, subject-matter responders) and how they are staffed 24x7 across 28 teams; the on-call structure, rotations, compensation and the rules for alert quality; detection and escalation paths; internal and customer communications (status page, account managers, regulators when required) with their timings; postmortems (when mandatory, format, blameless review, ownership and tracking of actions); the metrics and reviews that show whether it works; and how it is introduced across the teams without waiting for the audit." Generate an initial, detailed, and well-structured plan proposal. The plan has these parts: - "steps": the list of steps, each with "step_id" (a short label unique within this proposal, such as S1, S2, S3), "title" (a short, descriptive title), "description" (what the step does and how, in Markdown: a short opening paragraph and then, if there are several concrete things to say, a bullet list) and "dependencies" (the step_ids of the other steps of THIS proposal that must be completed first; empty if none). - "estimated_complexity": "low", "medium" or "high". - "success_metrics": clear and measurable success metrics, one per line as a bullet list. Think deeply about the problem, considering all aspects and implications. Develop a comprehensive understanding of what needs to be done to achieve the objective, then produce the plan.
{"steps": [{"step_id": "S1", "title": "Define severity levels and decision criteria", "description": "Create a four-tier severity framework (SEV-1 through SEV-4) that guides all downstream decisions about response, escalation, and communications.\n\nEach level must specify: customer impact (revenue at risk, customers affected, data loss risk); financial threshold triggering service credits; whether an incident commander is required; response time SLA (e.g., SEV-1 < 5 min notification, SEV-4 < 2 hours); and the go-live decision tree (when to declare and when to resolve).\n\n- SEV-1: Complete service down or critical path broken for >5% of customers; every minute costs money; IC required; 99.99% uptime threatened\n- SEV-2: Significant degradation, features unavailable, affecting 1–5% of customers; IC typically required\n- SEV-3: Minor impact, limited customer footprint or workaround exists; escalation path but not automatic IC\n- SEV-4: Observations or minor issues; alert-driven, no escalation unless pattern emerges", "dependencies": []}, {"step_id": "S2", "title": "Define incident roles and responsibilities", "description": "Create the organizational roles that operate during an incident: who is in charge, who talks to customers, who writes down what happened, who fixes the system, and how decisions are made under pressure.\n\nEach role must have a single-sentence mission, decision authority, and escalation upward.\n\n- Incident Commander: owns decision-making and timeline; declares severity; resolves conflicts; may or may not be technical\n- Deputy IC: shadow to IC and takes over if IC becomes unavailable\n- Communications Lead: writes status page, notifies account teams, manages customer perception\n- Scribe: records decisions, who did what, key timestamps; not responsible for fixing\n- SME Responders: engineers with context on the failing service(s); take IC's direction without debate", "dependencies": ["S1"]}, {"step_id": "S3", "title": "Design 24x7 on-call rotation structure", "description": "Build a rotation model that covers all 28 teams with primary and backup on-call engineers every hour across weekdays, evenings, weekends, and holidays; addresses the pager-carrying resistance.\n\nKey design decisions: Is coverage per-team (each team owns its services) or pooled (shared responder pool handles anything)? How many people per rotation? How long are shifts (one week, two weeks)? When can engineers opt out without leaving the team exposed? Which roles are on-call (IC, communications, SME)?\n\n- Recommend: dedicated IC pool (4–6 people in fast rotation) + per-team SME on-call for each team's own services\n- Recommend: two-week rotation blocks to reduce handoff friction\n- Recommend: one primary, one secondary per slot; secondary handles during primary's escalation\n- Provide swaps, blackout dates, and a rule that no engineer is on-call more than 2 weeks per quarter", "dependencies": ["S2"]}, {"step_id": "S4", "title": "Define on-call compensation and incentives", "description": "Create a pay model that makes on-call acceptable and rewards engineers who carry the pager; ties compensation to real business risk.\n\n- Base on-call stipend: e.g., $500–1,000 per week while on-call (regardless of incidents)\n- Callback pay: 1.5× hourly rate for time spent mitigating incidents during off-hours\n- Incident bonus: $50–100 extra per SEV-1 or SEV-2 incident mitigated (recognition)\n- Comp time: full business day off after an incident that required >2 hours mitigation during night/weekend\n- Annual bonus tie-in: 10–20% bonus multiplier for flawless on-call reviews\n- Communicate: position as investment in reliability, not punishment for being online", "dependencies": ["S3"]}, {"step_id": "S5", "title": "Consolidate alert routing infrastructure", "description": "Replace six alert tools with a single ingestion and routing system; stop engineers from being woken by duplicate alerts, and make escalation automated instead of manual.\n\nEvaluate existing tools (likely candidates: PagerDuty, Opsgenie, or Incident.io) or build a lightweight wrapper. The system must: accept alerts from all 180 services; deduplicate and correlate (same outage, different monitoring source); route to correct on-call engineer; expose an API for playbook automation; log every alert for postmortem analysis.\n\n- Choose tool by week 2 of S5\n- Migrate alerting endpoints from 6 sources to 1 by week 4\n- Set up audit trail and retention\n- Ensure mobile app works (on-call engineers need to engage from phone)", "dependencies": []}, {"step_id": "S6", "title": "Build alert quality rules to cut noise", "description": "Implement rules that automatically suppress the 85% of alerts that are noise (flapping, transient errors, auto-recovered conditions). Target: <400 actionable alerts per month.\n\nRules to implement: suppress alerts if service auto-recovered within 30 seconds; deduplicate same alert from multiple monitoring sources; suppress alerts for known maintenance windows; group flapping alerts (same service, >5 occurrences in 2 minutes) into a single page to on-call; rate-limit alerts from noisy services (e.g., max 1 alert per 5 minutes per service until silence clears).\n\n- Audit existing 3,400 alerts per month: which are true signals, which are noise\n- Tag each alert source with severity level (S1, S2, S3, S4 from S1)\n- Create exceptions list: services known to be noisy, require different rules\n- Weekly review: alert teams that trigger >50 alerts per week for reduction strategies", "dependencies": ["S1", "S5"]}, {"step_id": "S7", "title": "Implement automated detection and escalation paths", "description": "Wire the alert system to automatically escalate based on time or severity; removes the need for manual judgment calls during chaos.\n\nLogic: SEV-1 alert arrives → IC notified instantly via phone call + SMS + Slack + mobile; if IC does not acknowledge within 2 minutes, page deputy IC; Communications Lead pinged simultaneously. SEV-2: on-call SME for that service + IC notify via Slack and mobile, escalate to IC's manager if not acknowledged in 10 min. SEV-3/4: on-call SME only, escalate after 30 min.\n\n- Implement in alert routing system (S5)\n- Test all paths weekly via synthetic page to on-call\n- Track escalation metrics: how many pages reach secondary, how many hit manager\n- Adjust timing based on first month of operations", "dependencies": ["S1", "S5", "S6"]}, {"step_id": "S8", "title": "Build incident dashboard and status tracking", "description": "Create a single source of truth during an incident that every responder sees in real time: who is on-call, incident timeline, who said what, current status, next steps.\n\nDashboard displays: active incidents and their severity; who is the IC and communications lead; timeline of all events (alert fired, IC assigned, customer notified, mitigation started, resolved); Slack channel and mobile notification status; on-call rosters (who is on-call right now for each team); postmortem link as soon as incident closes.\n\n- Integrate with alert tool (S5) to auto-populate incident creation and initial severity\n- Push updates to status page and customer account managers automatically\n- Log all timeline entries for audit and postmortem completeness\n- Mobile-optimized so IC can work from any device", "dependencies": ["S5"]}, {"step_id": "S9", "title": "Write incident playbooks for each severity", "description": "Create a one-page (or one-screen) reference for the IC and SMEs during an incident; sequences the steps and removes ambiguity.\n\nEach severity level gets its own playbook: who gets paged (roles, order); first questions to ask (is it real, how big, who knows); what the IC should declare in first message (status page text, account manager notification, regulatory trigger); how long before escalating to executive team; decision rules for going dark vs. continuing to update customers.\n\n- SEV-1 playbook: immediate IC + comms + CTO notification; customer status every 5 minutes\n- SEV-2 playbook: IC + comms + tech lead notification; status every 15 minutes\n- SEV-3 playbook: on-call SME + comms if customer-visible; status every 30 min or as resolved\n- SEV-4 playbook: on-call SME only; update customers only if promised SLA is at risk\n- Include decision trees: is this SEV-1 or SEV-2? Is it our code or dependency? Escalate or containment?", "dependencies": ["S1", "S2", "S3"]}, {"step_id": "S10", "title": "Define internal communication workflows", "description": "Specify who informs whom, in what order, via what channel (call, Slack, email) during an incident; prevents gaps like \"nobody knew who was in charge for an hour.\"\n\nWorkflow for SEV-1: IC assigned → IC calls CTO/VP Eng and incident channel lead within 1 minute; incident declared in #incidents Slack channel with severity, IC name, service affected; SME on-call for that service joins call automatically; IC pushes updates to #incidents every 5 minutes or when material change occurs. For SEV-2: IC notifies team leads via Slack, updates #incidents every 15 min. Define escalation: if IC is unreachable, deputy IC takes over and announces it.\n\n- Create a phone tree or on-call list accessible to responders\n- Set expectations: \"If you don't hear from IC in 2 minutes, call them\"\n- Use a single incident Slack channel per incident (auto-created by incident tool)\n- Log all comms in the incident dashboard for postmortem review", "dependencies": ["S2", "S3"]}, {"step_id": "S11", "title": "Design customer communication and status page process", "description": "Plan when and how to inform customers, account managers, and regulators; ensure 2,100 customers are not learning about outages from Twitter before you tell them.\n\nRules by severity: SEV-1 detected → status page updated within 3 minutes (even if root cause unknown; post \"Investigating\"); account managers of affected customers called within 5 minutes; regulatory notification (if payment processing down) queued for approval; customer email within 10 minutes with ETA for next update. SEV-2: status page within 10 min, account managers called within 15 min, email if affecting >10 customers. SEV-3/4: no customer communication unless SLA at risk.\n\n- Empower Communications Lead to update status page without IC approval if delay >3 min\n- Prepare templated messages for common scenarios (database failover, data pipeline stuck, service crashed)\n- Route regulatory notifications through legal/compliance; don't wait for perfect root cause\n- Track customer impact in real time: how many customers affected by severity", "dependencies": ["S1", "S2"]}, {"step_id": "S12", "title": "Establish blameless postmortem process and format", "description": "Build a systematic way to learn from incidents so the same failure does not happen twice; counter the fear that admitting a mistake leads to being blamed.\n\nMandatory postmortems: all SEV-1 and SEV-2 incidents, within 48 hours of resolution. Optional but encouraged: SEV-3 if interesting or if >3 of same type in 30 days. Format: what was the user-visible impact and for how long; what was the root cause (not \"human error\" but the system condition that made error possible); timeline of discovery and response; action items with owner and deadline; blameless tone (focus on process and system design, not individual mistakes).\n\n- Assign a facilitator (not the on-call IC) to run postmortem\n- Attendees: IC, comms lead, SMEs involved, team lead, customer success if customer-facing\n- Write postmortem in shared doc; make it findable (searchable, linked from incident)\n- No discussion of \"who screwed up\"; only \"why did the system allow this to happen\"", "dependencies": ["S1", "S2"]}, {"step_id": "S13", "title": "Build action item tracking and accountability", "description": "Create a system that tracks postmortem action items so they are not forgotten; currently 11 of 64 (17%) are being tracked, leaving 53 unfinished improvements.\n\nSystem: each postmortem generates action items (e.g., \"add monitoring for X,\" \"update runbook for Y,\" \"write test for Z\"). Each item gets: clear description, owner (engineer's name), due date (1–4 weeks based on priority), severity (critical = must do before similar incident happens again; important = improve next month; nice-to-have = backlog). Action items live in a dedicated Jira project visible to all teams; owners are accountable (their manager reviews quarterly). Weekly: incident commander reviews open items due that week. Monthly: each team's postmortem items reviewed in their standup.\n\n- Export action items from postmortem document to tracking system automatically\n- Require IC to sign off that an action is complete before closing\n- Report on completion rate as a metric (target: >80% by month 3)", "dependencies": ["S12"]}, {"step_id": "S14", "title": "Define incident metrics and KPIs", "description": "Establish what \"good\" looks like; measure so you can improve. Target metrics for 12 months out: mean time to detect 8 minutes (vs. 22 now), customers detect first <5% of incidents (vs. 40%), MTTR 45 minutes for SEV-1 (vs. 190), SLA credits <$100k/year.\n\nMetrics to track: (1) MTTD = time from incident start to first alert/report; disaggregate: external report vs. internal detection. (2) MTTR = time from first report to full mitigation; track by severity and by service. (3) Customer-reported incidents per month (should drop to <2 per month). (4) Alert signal-to-noise ratio (goal: <5% false positive after S6 rules). (5) On-call satisfaction (survey: would you do this again?). (6) Postmortem action completion rate. (7) Incident commander and responder utilization (hours per week per person).\n\n- Dashboard: auto-populated from incident tool, updated daily\n- Disaggregate by team and service: which teams have bad MTTR? Which service is most incident-prone?", "dependencies": []}, {"step_id": "S15", "title": "Create review cadence and governance process", "description": "Establish regular rhythm to inspect the metrics, spot trends, and adjust the process itself; prevent the system from calcifying.\n\nWeekly: incident commander and on-call lead review prior week—number of incidents, any escalations, any communication gaps. Monthly: director-level incident review—trends by service, top causes of incidents, action item status, whether severity classification is working. Quarterly: full leadership review—MTTD, MTTR, customer impact, SLA credit spend, on-call satisfaction score, any systemic changes needed. Annually: audit the entire process for SOC 2 compliance.\n\n- Assign meeting owners: weekly = on-call lead; monthly = director of reliability; quarterly = VP Eng + CFO (SLA cost) + customer success\n- Use same data dashboard (S14) for all reviews\n- Publish a monthly \"incident newsletter\" to all engineers: what happened, what we learned, what's improving", "dependencies": ["S14"]}, {"step_id": "S16", "title": "Prepare SOC 2 Type II audit checklist", "description": "Document that the incident management system meets the control requirements for a SOC 2 audit; audit happens in 8 months, so this work builds confidence in coverage.\n\nAudit will test: (1) Is there a defined incident response process? (2) Are roles and responsibilities clear? (3) Are incidents logged and tracked? (4) Is root cause analysis performed? (5) Are action items tracked and completed? (6) Is on-call staffing adequate? (7) Are communications timely? (8) Are postmortems documented and blameless? Create a control mapping document that links each SOC 2 requirement to your process (S1–S15). Collect evidence: incident logs, postmortem documents, action item tickets, metrics reports, training records.\n\n- Designate a compliance owner (often a reliability lead or security engineer)\n- Run a mock audit at month 6 to identify gaps\n- Ensure all postmortems and incidents are retained and searchable for auditor review", "dependencies": ["S1", "S2", "S9", "S12", "S13", "S14", "S15"]}, {"step_id": "S17", "title": "Develop implementation and rollout plan", "description": "Create a phased timeline to roll out the incident management system across all 28 teams; avoids big-bang failure and builds credibility.\n\nRecommended structure: Phase 1 (weeks 1–4): build and test infrastructure (S5, S8, alerting); deploy severity levels and roles (S1, S2); pick pilot teams (2–3 high-traffic teams). Phase 2 (weeks 5–12): train pilot teams, run incident drills, refine playbooks based on learning; expand to half of remaining teams. Phase 3 (weeks 13–20): full rollout to all 28 teams; continue drills; track metrics. Phase 4 (weeks 21–28): stabilize, iterate on metrics, prepare for audit.\n\n- Assign a release manager to coordinate across teams\n- Create a detailed Gantt chart with swim lanes (infra, process, training, rollout)\n- Identify risks: competing priorities, engineers worried about pager burden, tool adoption friction\n- Plan stakeholder engagement: weekly updates to eng leadership, monthly town halls for all engineers", "dependencies": ["S1", "S2", "S3", "S4", "S5", "S6", "S7", "S8", "S9", "S10", "S11", "S12", "S13", "S14", "S15", "S16"]}, {"step_id": "S18", "title": "Build training and documentation", "description": "Create role-specific education so engineers understand the new system and are confident executing during an incident.\n\nTraining tracks: (1) All engineers: 30-min video on severity levels, communication expectations, postmortem format, action items (mandatory); (2) On-call responders: 2-hour workshop on playbooks, alert tool, escalation paths, when to call manager, case studies of real incidents, mobile app walkthrough; (3) Incident commanders: 4-hour IC bootcamp covering leadership under pressure, decision-making, communicating with executives, status page discipline, postmortem facilitation, practiced drills; (4) Communications leads: templates, when to update, how to talk to customers, regulatory notification rules.\n\n- Record videos so async teams can learn on their schedule\n- Create runbooks and quick-reference cards for each role (print + digital)\n- Pair new on-call engineers with experienced responder for first week\n- Require IC certification before anyone joins IC rotation (pass a practical drill)", "dependencies": ["S17"]}, {"step_id": "S19", "title": "Execute staged rollout across teams", "description": "Progressively activate the incident management system with feedback loops at each stage; reduces risk of system-wide failure.\n\nWave 1 (week 6–8): 3–4 pilot teams begin on-call rotations and incident response using new system; capture feedback daily. Wave 2 (week 10–14): 8–10 additional teams, incorporating lessons from Wave 1; ensure diversity of team types (payment processing, monitoring, data pipeline, auth, etc.). Wave 3 (week 15–20): remaining teams; by now, the system is proven and less hand-holding needed.\n\n- Daily retros with Wave 1 teams: what worked, what was confusing, what broke\n- Each wave produces a \"lessons learned\" document that informs the next\n- Track adoption metrics: how many incidents reported per team, alert quality, MTTD/MTTR\n- Address resistance: engineers who are skeptical of the system, on-call burden, tool friction; assign a \"change champion\" in each team", "dependencies": ["S18"]}, {"step_id": "S20", "title": "Run incident response drills and simulations", "description": "Practice incidents in a controlled setting so responders gain confidence and gaps are found before a real crisis.\n\nDrill 1 (week 8, pilot teams only): simulated database failover; test alert routing, IC assignment, on-call notification, status page update, postmortem. Drill 2 (week 12, all teams): full incident spanning multiple services; test escalation, communications, decision-making. Drill 3 (week 18): chaos drill with intentional confusion—what if IC is unreachable, what if comms lead is on PTO, what if an alert tool goes down? Quarterly thereafter: repeat drills with new on-call engineers.\n\n- Schedule drills at a time when key people are available (not Friday afternoon)\n- Run drill with live observers (reliability lead, manager) who provide feedback\n- Measure drill success: all steps completed on time, communications clear, postmortem written\n- Debrief immediately: what did we learn, what do we change?", "dependencies": ["S19"]}, {"step_id": "S21", "title": "Establish continuous improvement process", "description": "Build a feedback loop so the incident management system evolves based on what you learn.\n\nMonthly improvement cycle: in the monthly incident review (S15), leadership identifies one or two friction points from recent incidents (e.g., \"alert tool response is slow,\" \"postmortem meetings are boring,\" \"action items keep slipping\"). Assign an owner to investigate and propose a fix (change to playbook, tweak to tool config, additional training, process change). Test the fix with pilot team, gather feedback, roll out if successful. Annual review: can we reduce on-call burden, simplify communication, improve MTTD further? Do any teams need special handling (new services, difficult customers)?\n\n- Keep a public backlog of improvement ideas (Jira, shared doc) so any engineer can suggest\n- Quarterly all-hands: share wins (\"MTTD improved 50%,\" \"on-call satisfaction up,\" \"zero customer-detected incidents\") to build momentum\n- Link improvements back to metrics (S14): show that changes actually worked", "dependencies": ["S15", "S20"]}], "estimated_complexity": "high", "success_metrics": "- Median time to detect reduced from 22 minutes to <8 minutes within 6 months\n- Customer-detected incidents drop from 40% to <5% within 6 months\n- Median time to mitigate (MTTR) reduced from 3h 10min to <45 minutes for SEV-1 incidents within 6 months\n- Annual SLA credits decrease from $1.3M to <$100k within 12 months\n- Alert noise reduced from 3,400 per month (85% false positive) to <400 per month (>95% signal) within 3 months\n- Zero incidents with command-and-control ambiguity (>1 hour without clear IC) within 2 months\n- Postmortem action item completion rate reaches >80% (from 17%) within 4 months\n- On-call satisfaction score (survey) reaches >7/10 for on-call engineers within 3 months\n- All 28 teams integrated into incident management system with active on-call rotations by week 20\n- SOC 2 Type II audit passes incident response controls with no findings 8 months from start\n- Incident commander certification: 100% of active ICs trained and drilled within 2 months\n- Monthly incident review meeting established and attended by leadership; trends documented\n- New incident system integration complete: single alert tool, single dashboard, all 180 services feeding in, <5 min deployment"}Starts with a funded charter and a named process owner (S1), then a 12-month incident and alert baseline used as both problem statement and audit evidence (S2). Builds severity with a SEV0 for security/regulatory events, federated per-team on-call with a central IC roster, SLO- and synthetic-based detection, a 90-day noise sprint with a two-pages-per-shift budget, comms timing SLAs with pre-approved legal templates and a status page outside the failure domain, evidence-based action closure, then pilot, four gated waves and a SOC 2 dry run. Ends with a standing council and a separate resilience roadmap.
The plan produced
1. Programme charter, ownership and executive mandate
This step turns the CEO's email into a funded programme with a named owner and explicit authority. Without it, every downstream decision stalls in cross-team negotiation.
- Appoint a single accountable process owner (for example a Director of Incident Management) reporting to the CTO, with a dotted line to the COO for customer and SLA matters.
- Publish a one-page charter covering scope (all customer-impacting and money-moving incidents), decision rights, and the power to override team preferences during an active incident.
- Define the funding envelope: tooling licences, training time, exercise time and on-call compensation, with an indicative annual figure.
- Set the timeline against the SOC 2 date: a working process in four months, evidence accumulating from month two, audit-ready by month seven.
- Stand up a steering group with CTO, VP Engineering, Head of Support, Head of Compliance and one engineering manager per region.
- Agree that incident-process participation is a documented performance expectation for engineering managers, not an optional extra.
2. Baseline measurement and evidence pack (after 1)
You cannot demonstrate improvement without a defensible starting point, and you cannot win the argument about noise without numbers.
- Build a 12-month incident register: date, retro-assigned severity, detection source, time to detect, time to mitigate, customer impact, services involved and SLA credits paid.
- Quantify the alert estate per tool, per team and per service; compute page-to-action ratio, list the 50 noisiest rules and count off-hours interruptions per engineer.
- Survey on-call engineers and managers on burden, fairness and clarity of escalation, targeting a response rate above 70 percent.
- Interview Support, Account Management and Sales about how customers learn of incidents and what they complain about.
- Document exactly where the current process breaks: unclear command in the two known incidents, postmortem action closure at 11 of 64, and ad-hoc status-page authorship.
- Publish the pack internally as the problem statement and retain it as management-review evidence for the SOC 2 audit.
3. Severity taxonomy and trigger matrix (after 1, 2)
Severity is the keystone of the whole process. Every other rule, from paging to communications timing to postmortems, is keyed off it.
- Define four levels plus a special SEV0 for security or regulatory events: SEV1 for total or material loss of a payment path, SEV2 for degradation or single-region loss, SEV3 for limited impact with a workaround, SEV4 for internal-only issues and near-misses.
- Anchor each level in customer and money terms: customers affected, dollars at risk, regions lost, ledger integrity, settlement delay.
- Specify automatic triggers, for example loss of one AWS region, ledger write failures, a missed settlement cut-off, or payment success rate below threshold for five minutes.
- State who may declare each level (any engineer, Support or account manager may declare) and who may only recommend a downgrade (the Incident Commander alone).
- Map each level to SLA credit exposure and to the customer-visible status-page state.
- Include worked examples from the last 12 months so teams recognise their own incidents in the definitions.
- Add a review clause: the taxonomy is re-validated quarterly against real declarations.
4. Incident roles, command structure and decision rights (after 2, 3)
The two incidents where nobody knew who was in charge for over an hour are the direct brief for this step.
- Define roles with one-page role cards: Incident Commander, Deputy IC, Operations Lead, Internal Communications Lead, Customer Communications Lead, Scribe, Subject-Matter Responders and an Executive Sponsor for SEV1 only.
- State the core rule plainly: the IC owns the incident, not the fix, and does not debug.
- Give the IC explicit decision rights: declaring and escalating severity, freezing changes, halting deploys, approving customer messaging and calling additional responders.
- Define minimum viable role coverage per severity: SEV1 staffs every role, SEV3 staffs an IC and a scribe only.
- Define handover discipline: maximum four-hour IC shifts on SEV1, a written handover template, and a Deputy IC nominated within 15 minutes.
- Set behaviour standards for responders: one incident channel, one bridge, no side channels, and explicit asks with a named owner and a time.
- Publish role cards on the internal wiki and link them from every paging notification.
5. Escalation, paging and incident lifecycle policy (after 3, 4)
This step defines the mechanical path from an alert to a declared incident and back to normal service.
- Define lifecycle states with entry and exit criteria: Detected, Triaged, Declared, Mitigated, Resolved, Postmortem, Closed.
- Set acknowledgement targets: page acknowledged within five minutes, triage decision within fifteen, declaration within thirty for anything customer-facing.
- Define escalation ladders per layer (responder, service owner, team manager, IC on-call, VP Engineering) each with an automatic timer.
- Make escalation blameless and automatic: no responder is ever criticised for escalating, and timers fire whether or not a human asks.
- Define change freeze and rollback authority during SEV1 and SEV2, and the single condition that lifts the freeze.
- Enforce one incident, one record: the incident record is the sole source of truth for timeline, roles and communications.
- Require every SEV1 and SEV2 to produce an automatically captured timeline from channel and bridge, never one written from memory afterwards.
6. Detection strategy: SLOs, signals and customer-journey monitoring (after 2, 3)
Customers detected 40 percent of incidents first. That number is the reason this step exists.
- Define SLIs and SLOs for the top 20 customer journeys, including payment initiation, settlement, ledger read and write, API availability and webhook delivery, measured per region.
- Require symptom-based alerting on those SLOs rather than cause-based alerting on infrastructure metrics.
- Add synthetic transaction monitoring from outside AWS in both regions and a third location, on a one-minute cadence for money-moving paths.
- Add ledger and shared-PostgreSQL specific signals: replication lag, connection saturation, write latency, transaction ID exhaustion and checkpoint pressure.
- Open a customer-reported path so Support and account managers can raise an incident directly, and count that path as a detection source in reporting.
- Publish a detection contract per service: every one of the 180 services needs an owner, at least one symptom alert and a documented expected detect time.
- Fund a separate resilience track to reduce shared-cluster blast radius, because better detection will not save a single shared ledger during a corruption event.
7. Alert quality standard and noise-reduction programme (after 2, 3, 6)
3,400 alerts a month with 85 percent noise is the reason engineers resent the pager. Fixing it is the price of admission for everything else.
- Publish alert standards: every page must be symptom-based, actionable, owned, linked to a runbook and mapped to a severity. No runbook, no page.
- Separate paging from ticketing: only customer-impacting or imminent-impact signals may page; everything else becomes a ticket or a dashboard entry.
- Set a noise budget per team and per service, for example no service may exceed two pages per on-call shift, measured monthly.
- Define the default action for a noisy alert: fix, tune or delete within ten working days, with deletion treated as a legitimate and successful outcome.
- Introduce correlation and deduplication at the event pipeline so a single root cause produces one page instead of forty.
- Require expiry dates on every silencing rule and temporary threshold so suppression cannot become permanent blindness.
- Run a 90-day noise sprint with a visible burn-down of the top 100 noisiest rules, each owned by a named manager.
- Report page-to-action ratio per team in the monthly reliability review.
8. On-call architecture and 24x7 coverage model across 28 teams (after 3, 4)
This is the hardest political step. The answer to carrying a pager for another team's code is that every team carries its own, and the platform carries the shared risk.
- Adopt a federated model: every service has exactly one owning team, and that team's primary on-call carries its own pager. No team is paged for code it does not own.
- State the consequence honestly: 16 of 28 teams currently have no on-call. They must build one or formally transfer ownership of their services to a team that will.
- Merge small or low-traffic teams into shared rotations so no rotation has fewer than six engineers, below which coverage is not sustainable.
- Create a separate, centrally staffed Incident Commander on-call roster drawn from trained senior engineers across all teams, covering 24x7.
- Define primary and secondary per rotation, with the secondary engaged only on a no-acknowledge or an explicit request.
- Define coverage across the two AWS regions and New York business hours: one global IC rotation, service on-call aligned to their service's users.
- Define the unresponsive-team path: fifteen minutes without acknowledgement escalates to the team manager, thirty minutes to the IC, who may then direct any available engineer.
- Publish a coverage matrix of all 28 teams showing services, rotation size and gaps, reviewed monthly.
- Make on-call participation an explicit expectation in engineering job levels and hiring criteria.
9. On-call compensation, wellbeing and sustainability policy (after 8)
Unpaid on-call is the most cited reason for resistance. The policy must be settled before rollout, not negotiated during it.
- Move to paid on-call: a per-shift stipend or salary uplift agreed with HR and Finance and benchmarked to the New York market.
- Pay event-based compensation for incident callouts outside business hours, with a minimum call-out block.
- Provide compensatory rest: no engineer works a normal day after a night incident, and the rest day is documented, not granted as a favour.
- Cap intrusion by defining a maximum number of off-hours pages per shift, with a mandatory review triggered whenever it is exceeded.
- Define a voluntary opt-out path for engineers with genuine constraints, balanced by an explicit obligation that someone else is paid to take the shift.
- Include on-call expectation and compensation in offers and job descriptions so the commitment is set before hiring.
- Publish the policy with an effective date before any team is asked to join a new rotation.
- Review the policy every six months against actual page volumes, attrition and survey results.
10. Internal and customer communications policy with timing SLAs (after 3, 4)
Today the status page is written by whoever is around. This step replaces improvisation with a clock and a named owner.
- Set internal cadence: first update within 15 minutes of declaration, then every 30 minutes for SEV1 and 60 minutes for SEV2, regardless of whether there is progress.
- Set customer cadence: status page updated within 30 minutes of a SEV1 and 60 minutes of a SEV2, and a no-new-information update is still mandatory.
- Define the channel hierarchy: status page for everyone, direct email to affected customers on SEV1, named account-manager calls for the top 50 accounts.
- Prepare templates per severity in advance with legal and compliance pre-approval, covering detection, impact, workaround, mitigation and next-update time.
- Define regulatory obligations explicitly: money transmitter and banking regulator notification windows, security breach notification, and who signs off (Compliance, not Engineering).
- Prohibit speculation: customer communications never guess at cause or blame and never commit to a root cause before the postmortem.
- Define the closing communication: resolution notice, expected SLA credit handling and the committed date for a written report.
- Assign a named Customer Communications Lead per incident with a trained deputy on every SEV1.
11. Status page, notification tooling and account-manager playbook (after 10)
Policy without tooling collapses at three in the morning. This step makes publishing a five-minute action.
- Upgrade or replace the status page so components map to customer journeys rather than internal services, with subscriber control per component.
- Integrate the incident tool with the status page so the incident record drives the update and the public timeline.
- Provide one-click templates pre-filled with severity, impact language and next-update time.
- Give account managers a playbook: contact tree, what they may say, what they must not say, and how to escalate a customer question into the incident channel within minutes.
- Define the SLA credit process end to end, covering computation, approval, customer notification and finance treatment, so credits stop being a manual scramble.
- Host the status page outside the production failure domain so it survives a total platform outage.
- Test publishing during game days, including a simulated status-page outage and a simulated loss of the primary region.
12. Postmortem policy, template and blameless review process (after 3, 4)
Only 11 of 64 action items closed means the postmortem ritual is currently a writing exercise. This step rebuilds it around learning and tracking.
- Make postmortems mandatory for every SEV1 and SEV2, every SEV3 with customer impact or a repeat pattern, and any near-miss the IC flags.
- Set deadlines: draft within five business days, blameless review within ten, internal publication within fifteen.
- Adopt a single template: impact, timeline, detection, response, contributing factors, what went well, what went badly, action items.
- Train a pool of blameless facilitators and require a trained facilitator for every SEV1 review.
- Prohibit counterfactual and blame language in the template, and require contributing factors across tooling, process, organisation and human factors.
- Limit action items to a small number of concrete, verifiable, owned changes with dates.
- Publish all postmortems internally by default, with security review only for genuinely sensitive material.
- Produce a customer-facing root cause report variant for SEV1 incidents, especially those affecting regulated or top-tier accounts.
13. Corrective action tracking and reliability backlog governance (after 12)
A postmortem without durable action tracking is a complaint, not a control.
- Create a single reliability backlog in the engineering tracker with a mandatory label, owner, due date and link to the originating incident.
- Define closure criteria that require evidence: a merged change, a tested alert or a verified drill, never a self-reported status change.
- Protect capacity by reserving a fixed percentage of each team's sprint for reliability work, with unspent capacity visible to vice presidents.
- Run a weekly ageing review of open actions and escalate anything overdue by more than 30 days to the VP Engineering.
- Report closure rate and median age monthly, targeting more than 90 percent closed within 60 days.
- Require a repeat incident in the same area to trigger a design review rather than another action item.
14. Incident tooling consolidation and integration (after 3, 5, 7, 11)
Six alerting tools and no single incident record is a structural cause of the 22-minute detection and the three-hour mitigation.
- Select one incident management platform for paging, on-call schedules, escalation policies, incident records and postmortem workflow.
- Consolidate the six alerting sources into a single event pipeline feeding that platform, with deduplication and severity mapping applied at ingest.
- Integrate with platform and ledger observability so responders see dashboards and runbooks inside the incident record.
- Integrate chat and bridge: incident channel auto-created, timeline auto-captured, decisions logged as they happen.
- Define the data model and retention required for SOC 2 evidence: who did what, when, and under whose authority.
- Run a dual-run period alongside the old tools with a defined rollback, then switch off the legacy tools on a published date.
- Budget for licences, migration effort and a two-week hardening period after cutover.
15. Training, certification and exercise programme (after 4, 5, 10, 12)
A process that exists only on a wiki page fails on the first real page. Skills have to be built and tested before they are needed.
- Build a short practical curriculum: how to be on call, how to declare an incident, how to run an incident as IC, how to communicate and how to write a postmortem.
- Require certification before joining the IC on-call roster: a written assessment plus a live simulated incident.
- Train at least two certified ICs per team group so the central roster has depth across all 28 teams.
- Run monthly tabletops on realistic scenarios drawn from the last 12 months, including region loss and ledger corruption.
- Run quarterly game days with deliberately injected failure, including shared-PostgreSQL failover and status-page outage.
- Add a mandatory onboarding module for every engineer joining or transferring in, with completion records suitable for audit.
- Track training completion by team and publish it in the monthly reliability review.
16. Metrics, dashboards and review cadence (after 2, 3)
The programme needs a public scoreboard, or it will quietly rot after the audit.
- Define the outcome metrics: time to detect by source, time to mitigate, percentage of incidents detected by customers (target below ten), incidents by severity and SLA credits paid.
- Define the process metrics: declaration latency, page acknowledgement rate, IC roster coverage, first-update timeliness and update-cadence adherence.
- Define the health metrics: alert volume and noise ratio per team, off-hours pages per engineer, postmortem timeliness, action closure rate and action age.
- Publish live dashboards visible to every engineer, not only to managers, refreshed daily.
- Institute a weekly operational review of 30 minutes going incident by incident, a monthly reliability review of trends and actions, and a quarterly executive review with the CEO's office.
- Baseline every metric against the S2 evidence pack and set 90-day and 12-month targets.
- Require every review to end with decisions and owners, not just numbers.
17. Pilot with volunteer teams (after 5, 7, 9, 11, 12, 13, 14, 15, 16)
Do not roll out to 28 teams untested. Run the entire process end to end with a small cohort first.
- Recruit three to four volunteer teams covering a mix of criticality: one payment-path team, one ledger-adjacent team, one platform team and one low-traffic team.
- Run the complete process in the pilot: new severity scale, roles, escalation, communications, postmortems, action tracking and paid on-call.
- Instrument the pilot against the S16 metrics and compare results with the S2 baseline.
- Hold weekly retrospectives with pilot teams and iterate on the written policies, the tooling and the training.
- Fix the top issues found before any wider rollout and document what changed and why.
- Produce a pilot report with before-and-after numbers to carry into every rollout conversation.
- Set explicit pilot exit criteria: rotation coverage achieved, no unacknowledged pages over a defined period, postmortems delivered on time and actions tracked.
18. Phased rollout to all 28 teams (after 13, 16, 17)
Rollout is a staged migration with readiness gates, not an email announcement.
- Sequence the 28 teams into four waves of roughly seven teams, ordered by customer impact, with three weeks between waves.
- Define a per-team readiness checklist: services mapped and owned, alerts cleaned to standard, runbooks written, rotation staffed, training complete and manager briefed.
- Hold a gate review with the process owner before each team joins, and move unready teams to the next wave with a dated remediation plan.
- Give each wave a named champion and run an internal communications cadence that explains the why using pilot numbers.
- Handle resistance directly by publishing the own-your-code-own-your-pager rule and the paid on-call mechanics before each wave, not after.
- Retire legacy tools, escalation lists and the informal status-page process at the end of each wave on a published cutover date.
- Harvest feedback formally at each wave and push accepted changes back into the policy documents through change control.
19. SOC 2 incident-response control mapping and evidence framework (after 1, 3, 10, 12)
The audit will test incident response. Mapping controls now means evidence accumulates from the first real incident under the new process.
- Map the process to the relevant Trust Services Criteria for incident identification, response, evaluation of incidents and communication of security events.
- Write control statements in auditor language and name a single owner for each control.
- Specify the evidence artifact for each control: incident record, severity classification, paging log, communications log, postmortem, action tracker entry and training record.
- Set evidence retention and storage location so nothing depends on a laptop or on chat history that expires.
- Run an early walkthrough with an experienced compliance partner or the auditor's readiness team to test the design before the audit window.
- Flag controls that are blocked until rollout reaches the relevant teams, and keep a gap register with owners and dates reviewed fortnightly by the steering group.
20. SOC 2 dry run, gap remediation and audit support (after 18, 19)
This step converts a good process into a provable one, a few months before the auditors arrive.
- Run an internal dry run six weeks before the audit window, sampling real incidents from the pilot and early waves against each control's evidence requirements.
- Remediate every gap found, prioritising anything that risks a qualified opinion.
- Brief every engineer who may be interviewed so they can describe the process as they actually practise it.
- Prepare the auditor package: process documentation, sample incident records, training records, on-call schedules and the action tracking register.
- Designate a single audit liaison and a small evidence-request team so requests do not land on 28 teams at once.
- Rehearse the walkthrough with the incident commander and communications roles, since auditors probe realism under pressure.
- Schedule the dry run and the audit alongside rollout waves so no team is hit with both in the same week.
21. Standing governance, process ownership and continuous improvement (after 18, 20)
The classic post-audit failure is that the process freezes and then decays. This step keeps it alive.
- Establish a standing Incident Management Council chaired by the process owner, meeting monthly with engineering, support, compliance and product representation.
- Give the process owner a documented mandate to change standards, and require formal change control for any change to severity, roles, communications timings or compensation.
- Re-validate the severity taxonomy quarterly against real declarations and re-baseline the metric set annually.
- Fund a resilience roadmap separate from incident response: reducing shared-database blast radius, multi-region failover readiness and deploy safety.
- Report quarterly to the executive team on the metric set and on the top five systemic causes of incidents.
- Feed incident themes into architecture review, release engineering and hiring so the programme improves the system, not just the response.
- Refresh the training and tabletop programme annually and after any SEV1.
- Median time to detect reduced from 22 minutes to 5 minutes or less by month 6.
- Customer-detected incidents reduced from 40 percent to 10 percent or less of all customer-impacting incidents by month 9.
- Median time to mitigate reduced from 3 h 10 min to 60 minutes or less for SEV1 and SEV2 by month 9.
- Customer-impacting incidents reduced by at least 40 percent year over year from the 31 incident baseline.
- SLA credits paid reduced by at least 50 percent year over year from the $1.3M baseline.
- 100 percent of SEV1 incidents have a named Incident Commander within 5 minutes and a scribe within 15 minutes.
- Zero incidents in which command authority is unclear for more than 15 minutes.
- 28 of 28 teams have a documented on-call rotation, an owning service map and at least one trained on-call engineer by month 6.
- Central Incident Commander roster holds at least 40 certified ICs covering 24x7 with no single point of failure.
- 100 percent of on-call rotations are paid under a published policy by month 5.
- Monthly alert volume reduced from 3,400 to below 700, with a false-positive rate below 20 percent.
- No service exceeds 2 pages per on-call shift, measured monthly for three consecutive months.
- 100 percent of SEV1 and SEV2 postmortems published internally within 15 business days.
- At least 90 percent of postmortem action items closed within 60 days, up from 17 percent (11 of 64).
- Status-page first update published within 30 minutes on at least 95 percent of SEV1 incidents.
- Zero missed regulatory notification windows on any incident requiring notification.
- SOC 2 Type II audit passed with no findings related to incident response.
- Review cadence sustained: weekly operational review in at least 90 percent of weeks, 12 of 12 monthly reliability reviews, 4 of 4 quarterly executive reviews.
- On-call satisfaction at 70 percent or higher on the quarterly survey, with zero on-call-attributed voluntary attrition.
- 100 percent of new engineers complete the incident-response onboarding module within 30 days of joining.
[SYSTEM]
You are an expert assistant in complex project planning.
Your task is to generate a detailed and structured action plan to reach the main objective in the most professional and most detailed way, no matter how much work or steps will be needed to perform.
Use your internal reasoning processes to think deeply about the problem and create the most comprehensive plan possible.
Take as much time and space as you need to think through all aspects of the problem.
After your thorough analysis, answer with the plan in the requested structure.
Every text field you write will be read by a busy person, so write for them: short sentences, one idea each, in short paragraphs. Text fields accept Markdown: separate paragraphs with a blank line, use a bullet list when you enumerate things and plain paragraphs when you explain or argue, and bold at most one key phrase per paragraph. A text of more than three sentences must be split into paragraphs; never deliver a single unbroken block of text.
[HUMAN]
Main Objective: "A B2B payments platform in New York (260 engineers in 28 teams, 2,100 customers, $4B processed a month, a 99.95% SLA with service credits) runs about 180 services on Kubernetes across two AWS regions, with a shared PostgreSQL cluster for the ledger. In the last 12 months: 31 customer-impacting incidents, median time to detect 22 minutes (customers detected 40% of them first), median time to mitigate 3 h 10 min, $1.3M paid in SLA credits, two incidents in which nobody was sure who was in charge for over an hour, and a CEO email about "outages we hear about from clients". On-call exists in 12 of the 28 teams, unpaid, fed by alerts from six different tools — 3,400 alerts a month, 85% of them noise; there is no severity scale, status-page updates are written by whoever is around, and postmortems happen for some incidents, in various formats, with action items that are rarely tracked (11 of 64 closed). A SOC 2 Type II audit in eight months will test incident response. Engineers push back against "carrying a pager for other teams' code".
Define the incident management process: severity levels and what each one triggers; roles (incident commander, communications, scribe, subject-matter responders) and how they are staffed 24x7 across 28 teams; the on-call structure, rotations, compensation and the rules for alert quality; detection and escalation paths; internal and customer communications (status page, account managers, regulators when required) with their timings; postmortems (when mandatory, format, blameless review, ownership and tracking of actions); the metrics and reviews that show whether it works; and how it is introduced across the teams without waiting for the audit."
Generate an initial, detailed, and well-structured plan proposal.
The plan has these parts:
- "steps": the list of steps, each with "step_id" (a short label unique within this proposal, such as S1, S2, S3), "title" (a short, descriptive title), "description" (what the step does and how, in Markdown: a short opening paragraph and then, if there are several concrete things to say, a bullet list) and "dependencies" (the step_ids of the other steps of THIS proposal that must be completed first; empty if none).
- "estimated_complexity": "low", "medium" or "high".
- "success_metrics": clear and measurable success metrics, one per line as a bullet list.
Think deeply about the problem, considering all aspects and implications. Develop a comprehensive understanding of what needs to be done to achieve the objective, then produce the plan.
Answer with a single JSON object, and nothing else, that conforms to this JSON schema:
{"description": "The structure every proposing agent must answer with, in every round.", "properties": {"steps": {"items": {"description": "One step of a proposed plan. Dependencies are step_ids of the same proposal.", "properties": {"step_id": {"title": "Step Id", "type": "string"}, "title": {"title": "Title", "type": "string"}, "description": {"title": "Description", "type": "string"}, "dependencies": {"items": {"type": "string"}, "title": "Dependencies", "type": "array"}}, "required": ["step_id", "title", "description", "dependencies"], "title": "Step", "type": "object", "additionalProperties": false}, "title": "Steps", "type": "array"}, "estimated_complexity": {"title": "Estimated Complexity", "type": "string"}, "success_metrics": {"title": "Success Metrics", "type": "string"}}, "required": ["steps", "estimated_complexity", "success_metrics"], "title": "ProposalContent", "type": "object", "additionalProperties": false}
{"steps": [{"step_id": "S1", "title": "Programme charter, ownership and executive mandate", "description": "This step turns the CEO's email into a funded programme with a named owner and explicit authority. Without it, every downstream decision stalls in cross-team negotiation.\n\n- Appoint a single accountable process owner (for example a Director of Incident Management) reporting to the CTO, with a dotted line to the COO for customer and SLA matters.\n- Publish a one-page charter covering scope (all customer-impacting and money-moving incidents), decision rights, and the power to override team preferences during an active incident.\n- Define the funding envelope: tooling licences, training time, exercise time and on-call compensation, with an indicative annual figure.\n- Set the timeline against the SOC 2 date: a working process in four months, evidence accumulating from month two, audit-ready by month seven.\n- Stand up a steering group with CTO, VP Engineering, Head of Support, Head of Compliance and one engineering manager per region.\n- Agree that incident-process participation is a documented performance expectation for engineering managers, not an optional extra.", "dependencies": []}, {"step_id": "S2", "title": "Baseline measurement and evidence pack", "description": "You cannot demonstrate improvement without a defensible starting point, and you cannot win the argument about noise without numbers.\n\n- Build a 12-month incident register: date, retro-assigned severity, detection source, time to detect, time to mitigate, customer impact, services involved and SLA credits paid.\n- Quantify the alert estate per tool, per team and per service; compute page-to-action ratio, list the 50 noisiest rules and count off-hours interruptions per engineer.\n- Survey on-call engineers and managers on burden, fairness and clarity of escalation, targeting a response rate above 70 percent.\n- Interview Support, Account Management and Sales about how customers learn of incidents and what they complain about.\n- Document exactly where the current process breaks: unclear command in the two known incidents, postmortem action closure at 11 of 64, and ad-hoc status-page authorship.\n- Publish the pack internally as the problem statement and retain it as management-review evidence for the SOC 2 audit.", "dependencies": ["S1"]}, {"step_id": "S3", "title": "Severity taxonomy and trigger matrix", "description": "Severity is the keystone of the whole process. Every other rule, from paging to communications timing to postmortems, is keyed off it.\n\n- Define four levels plus a special SEV0 for security or regulatory events: SEV1 for total or material loss of a payment path, SEV2 for degradation or single-region loss, SEV3 for limited impact with a workaround, SEV4 for internal-only issues and near-misses.\n- Anchor each level in customer and money terms: customers affected, dollars at risk, regions lost, ledger integrity, settlement delay.\n- Specify automatic triggers, for example loss of one AWS region, ledger write failures, a missed settlement cut-off, or payment success rate below threshold for five minutes.\n- State who may declare each level (any engineer, Support or account manager may declare) and who may only recommend a downgrade (the Incident Commander alone).\n- Map each level to SLA credit exposure and to the customer-visible status-page state.\n- Include worked examples from the last 12 months so teams recognise their own incidents in the definitions.\n- Add a review clause: the taxonomy is re-validated quarterly against real declarations.", "dependencies": ["S1", "S2"]}, {"step_id": "S4", "title": "Incident roles, command structure and decision rights", "description": "The two incidents where nobody knew who was in charge for over an hour are the direct brief for this step.\n\n- Define roles with one-page role cards: Incident Commander, Deputy IC, Operations Lead, Internal Communications Lead, Customer Communications Lead, Scribe, Subject-Matter Responders and an Executive Sponsor for SEV1 only.\n- State the core rule plainly: the IC owns the incident, not the fix, and does not debug.\n- Give the IC explicit decision rights: declaring and escalating severity, freezing changes, halting deploys, approving customer messaging and calling additional responders.\n- Define minimum viable role coverage per severity: SEV1 staffs every role, SEV3 staffs an IC and a scribe only.\n- Define handover discipline: maximum four-hour IC shifts on SEV1, a written handover template, and a Deputy IC nominated within 15 minutes.\n- Set behaviour standards for responders: one incident channel, one bridge, no side channels, and explicit asks with a named owner and a time.\n- Publish role cards on the internal wiki and link them from every paging notification.", "dependencies": ["S2", "S3"]}, {"step_id": "S5", "title": "Escalation, paging and incident lifecycle policy", "description": "This step defines the mechanical path from an alert to a declared incident and back to normal service.\n\n- Define lifecycle states with entry and exit criteria: Detected, Triaged, Declared, Mitigated, Resolved, Postmortem, Closed.\n- Set acknowledgement targets: page acknowledged within five minutes, triage decision within fifteen, declaration within thirty for anything customer-facing.\n- Define escalation ladders per layer (responder, service owner, team manager, IC on-call, VP Engineering) each with an automatic timer.\n- Make escalation blameless and automatic: no responder is ever criticised for escalating, and timers fire whether or not a human asks.\n- Define change freeze and rollback authority during SEV1 and SEV2, and the single condition that lifts the freeze.\n- Enforce one incident, one record: the incident record is the sole source of truth for timeline, roles and communications.\n- Require every SEV1 and SEV2 to produce an automatically captured timeline from channel and bridge, never one written from memory afterwards.", "dependencies": ["S3", "S4"]}, {"step_id": "S6", "title": "Detection strategy: SLOs, signals and customer-journey monitoring", "description": "Customers detected 40 percent of incidents first. That number is the reason this step exists.\n\n- Define SLIs and SLOs for the top 20 customer journeys, including payment initiation, settlement, ledger read and write, API availability and webhook delivery, measured per region.\n- Require symptom-based alerting on those SLOs rather than cause-based alerting on infrastructure metrics.\n- Add synthetic transaction monitoring from outside AWS in both regions and a third location, on a one-minute cadence for money-moving paths.\n- Add ledger and shared-PostgreSQL specific signals: replication lag, connection saturation, write latency, transaction ID exhaustion and checkpoint pressure.\n- Open a customer-reported path so Support and account managers can raise an incident directly, and count that path as a detection source in reporting.\n- Publish a detection contract per service: every one of the 180 services needs an owner, at least one symptom alert and a documented expected detect time.\n- Fund a separate resilience track to reduce shared-cluster blast radius, because better detection will not save a single shared ledger during a corruption event.", "dependencies": ["S2", "S3"]}, {"step_id": "S7", "title": "Alert quality standard and noise-reduction programme", "description": "3,400 alerts a month with 85 percent noise is the reason engineers resent the pager. Fixing it is the price of admission for everything else.\n\n- Publish alert standards: every page must be symptom-based, actionable, owned, linked to a runbook and mapped to a severity. No runbook, no page.\n- Separate paging from ticketing: only customer-impacting or imminent-impact signals may page; everything else becomes a ticket or a dashboard entry.\n- Set a noise budget per team and per service, for example no service may exceed two pages per on-call shift, measured monthly.\n- Define the default action for a noisy alert: fix, tune or delete within ten working days, with deletion treated as a legitimate and successful outcome.\n- Introduce correlation and deduplication at the event pipeline so a single root cause produces one page instead of forty.\n- Require expiry dates on every silencing rule and temporary threshold so suppression cannot become permanent blindness.\n- Run a 90-day noise sprint with a visible burn-down of the top 100 noisiest rules, each owned by a named manager.\n- Report page-to-action ratio per team in the monthly reliability review.", "dependencies": ["S2", "S3", "S6"]}, {"step_id": "S8", "title": "On-call architecture and 24x7 coverage model across 28 teams", "description": "This is the hardest political step. The answer to carrying a pager for another team's code is that every team carries its own, and the platform carries the shared risk.\n\n- Adopt a federated model: every service has exactly one owning team, and that team's primary on-call carries its own pager. No team is paged for code it does not own.\n- State the consequence honestly: 16 of 28 teams currently have no on-call. They must build one or formally transfer ownership of their services to a team that will.\n- Merge small or low-traffic teams into shared rotations so no rotation has fewer than six engineers, below which coverage is not sustainable.\n- Create a separate, centrally staffed Incident Commander on-call roster drawn from trained senior engineers across all teams, covering 24x7.\n- Define primary and secondary per rotation, with the secondary engaged only on a no-acknowledge or an explicit request.\n- Define coverage across the two AWS regions and New York business hours: one global IC rotation, service on-call aligned to their service's users.\n- Define the unresponsive-team path: fifteen minutes without acknowledgement escalates to the team manager, thirty minutes to the IC, who may then direct any available engineer.\n- Publish a coverage matrix of all 28 teams showing services, rotation size and gaps, reviewed monthly.\n- Make on-call participation an explicit expectation in engineering job levels and hiring criteria.", "dependencies": ["S3", "S4"]}, {"step_id": "S9", "title": "On-call compensation, wellbeing and sustainability policy", "description": "Unpaid on-call is the most cited reason for resistance. The policy must be settled before rollout, not negotiated during it.\n\n- Move to paid on-call: a per-shift stipend or salary uplift agreed with HR and Finance and benchmarked to the New York market.\n- Pay event-based compensation for incident callouts outside business hours, with a minimum call-out block.\n- Provide compensatory rest: no engineer works a normal day after a night incident, and the rest day is documented, not granted as a favour.\n- Cap intrusion by defining a maximum number of off-hours pages per shift, with a mandatory review triggered whenever it is exceeded.\n- Define a voluntary opt-out path for engineers with genuine constraints, balanced by an explicit obligation that someone else is paid to take the shift.\n- Include on-call expectation and compensation in offers and job descriptions so the commitment is set before hiring.\n- Publish the policy with an effective date before any team is asked to join a new rotation.\n- Review the policy every six months against actual page volumes, attrition and survey results.", "dependencies": ["S8"]}, {"step_id": "S10", "title": "Internal and customer communications policy with timing SLAs", "description": "Today the status page is written by whoever is around. This step replaces improvisation with a clock and a named owner.\n\n- Set internal cadence: first update within 15 minutes of declaration, then every 30 minutes for SEV1 and 60 minutes for SEV2, regardless of whether there is progress.\n- Set customer cadence: status page updated within 30 minutes of a SEV1 and 60 minutes of a SEV2, and a no-new-information update is still mandatory.\n- Define the channel hierarchy: status page for everyone, direct email to affected customers on SEV1, named account-manager calls for the top 50 accounts.\n- Prepare templates per severity in advance with legal and compliance pre-approval, covering detection, impact, workaround, mitigation and next-update time.\n- Define regulatory obligations explicitly: money transmitter and banking regulator notification windows, security breach notification, and who signs off (Compliance, not Engineering).\n- Prohibit speculation: customer communications never guess at cause or blame and never commit to a root cause before the postmortem.\n- Define the closing communication: resolution notice, expected SLA credit handling and the committed date for a written report.\n- Assign a named Customer Communications Lead per incident with a trained deputy on every SEV1.", "dependencies": ["S3", "S4"]}, {"step_id": "S11", "title": "Status page, notification tooling and account-manager playbook", "description": "Policy without tooling collapses at three in the morning. This step makes publishing a five-minute action.\n\n- Upgrade or replace the status page so components map to customer journeys rather than internal services, with subscriber control per component.\n- Integrate the incident tool with the status page so the incident record drives the update and the public timeline.\n- Provide one-click templates pre-filled with severity, impact language and next-update time.\n- Give account managers a playbook: contact tree, what they may say, what they must not say, and how to escalate a customer question into the incident channel within minutes.\n- Define the SLA credit process end to end, covering computation, approval, customer notification and finance treatment, so credits stop being a manual scramble.\n- Host the status page outside the production failure domain so it survives a total platform outage.\n- Test publishing during game days, including a simulated status-page outage and a simulated loss of the primary region.", "dependencies": ["S10"]}, {"step_id": "S12", "title": "Postmortem policy, template and blameless review process", "description": "Only 11 of 64 action items closed means the postmortem ritual is currently a writing exercise. This step rebuilds it around learning and tracking.\n\n- Make postmortems mandatory for every SEV1 and SEV2, every SEV3 with customer impact or a repeat pattern, and any near-miss the IC flags.\n- Set deadlines: draft within five business days, blameless review within ten, internal publication within fifteen.\n- Adopt a single template: impact, timeline, detection, response, contributing factors, what went well, what went badly, action items.\n- Train a pool of blameless facilitators and require a trained facilitator for every SEV1 review.\n- Prohibit counterfactual and blame language in the template, and require contributing factors across tooling, process, organisation and human factors.\n- Limit action items to a small number of concrete, verifiable, owned changes with dates.\n- Publish all postmortems internally by default, with security review only for genuinely sensitive material.\n- Produce a customer-facing root cause report variant for SEV1 incidents, especially those affecting regulated or top-tier accounts.", "dependencies": ["S3", "S4"]}, {"step_id": "S13", "title": "Corrective action tracking and reliability backlog governance", "description": "A postmortem without durable action tracking is a complaint, not a control.\n\n- Create a single reliability backlog in the engineering tracker with a mandatory label, owner, due date and link to the originating incident.\n- Define closure criteria that require evidence: a merged change, a tested alert or a verified drill, never a self-reported status change.\n- Protect capacity by reserving a fixed percentage of each team's sprint for reliability work, with unspent capacity visible to vice presidents.\n- Run a weekly ageing review of open actions and escalate anything overdue by more than 30 days to the VP Engineering.\n- Report closure rate and median age monthly, targeting more than 90 percent closed within 60 days.\n- Require a repeat incident in the same area to trigger a design review rather than another action item.", "dependencies": ["S12"]}, {"step_id": "S14", "title": "Incident tooling consolidation and integration", "description": "Six alerting tools and no single incident record is a structural cause of the 22-minute detection and the three-hour mitigation.\n\n- Select one incident management platform for paging, on-call schedules, escalation policies, incident records and postmortem workflow.\n- Consolidate the six alerting sources into a single event pipeline feeding that platform, with deduplication and severity mapping applied at ingest.\n- Integrate with platform and ledger observability so responders see dashboards and runbooks inside the incident record.\n- Integrate chat and bridge: incident channel auto-created, timeline auto-captured, decisions logged as they happen.\n- Define the data model and retention required for SOC 2 evidence: who did what, when, and under whose authority.\n- Run a dual-run period alongside the old tools with a defined rollback, then switch off the legacy tools on a published date.\n- Budget for licences, migration effort and a two-week hardening period after cutover.", "dependencies": ["S3", "S5", "S7", "S11"]}, {"step_id": "S15", "title": "Training, certification and exercise programme", "description": "A process that exists only on a wiki page fails on the first real page. Skills have to be built and tested before they are needed.\n\n- Build a short practical curriculum: how to be on call, how to declare an incident, how to run an incident as IC, how to communicate and how to write a postmortem.\n- Require certification before joining the IC on-call roster: a written assessment plus a live simulated incident.\n- Train at least two certified ICs per team group so the central roster has depth across all 28 teams.\n- Run monthly tabletops on realistic scenarios drawn from the last 12 months, including region loss and ledger corruption.\n- Run quarterly game days with deliberately injected failure, including shared-PostgreSQL failover and status-page outage.\n- Add a mandatory onboarding module for every engineer joining or transferring in, with completion records suitable for audit.\n- Track training completion by team and publish it in the monthly reliability review.", "dependencies": ["S4", "S5", "S10", "S12"]}, {"step_id": "S16", "title": "Metrics, dashboards and review cadence", "description": "The programme needs a public scoreboard, or it will quietly rot after the audit.\n\n- Define the outcome metrics: time to detect by source, time to mitigate, percentage of incidents detected by customers (target below ten), incidents by severity and SLA credits paid.\n- Define the process metrics: declaration latency, page acknowledgement rate, IC roster coverage, first-update timeliness and update-cadence adherence.\n- Define the health metrics: alert volume and noise ratio per team, off-hours pages per engineer, postmortem timeliness, action closure rate and action age.\n- Publish live dashboards visible to every engineer, not only to managers, refreshed daily.\n- Institute a weekly operational review of 30 minutes going incident by incident, a monthly reliability review of trends and actions, and a quarterly executive review with the CEO's office.\n- Baseline every metric against the S2 evidence pack and set 90-day and 12-month targets.\n- Require every review to end with decisions and owners, not just numbers.", "dependencies": ["S2", "S3"]}, {"step_id": "S17", "title": "Pilot with volunteer teams", "description": "Do not roll out to 28 teams untested. Run the entire process end to end with a small cohort first.\n\n- Recruit three to four volunteer teams covering a mix of criticality: one payment-path team, one ledger-adjacent team, one platform team and one low-traffic team.\n- Run the complete process in the pilot: new severity scale, roles, escalation, communications, postmortems, action tracking and paid on-call.\n- Instrument the pilot against the S16 metrics and compare results with the S2 baseline.\n- Hold weekly retrospectives with pilot teams and iterate on the written policies, the tooling and the training.\n- Fix the top issues found before any wider rollout and document what changed and why.\n- Produce a pilot report with before-and-after numbers to carry into every rollout conversation.\n- Set explicit pilot exit criteria: rotation coverage achieved, no unacknowledged pages over a defined period, postmortems delivered on time and actions tracked.", "dependencies": ["S5", "S7", "S9", "S11", "S12", "S13", "S14", "S15", "S16"]}, {"step_id": "S18", "title": "Phased rollout to all 28 teams", "description": "Rollout is a staged migration with readiness gates, not an email announcement.\n\n- Sequence the 28 teams into four waves of roughly seven teams, ordered by customer impact, with three weeks between waves.\n- Define a per-team readiness checklist: services mapped and owned, alerts cleaned to standard, runbooks written, rotation staffed, training complete and manager briefed.\n- Hold a gate review with the process owner before each team joins, and move unready teams to the next wave with a dated remediation plan.\n- Give each wave a named champion and run an internal communications cadence that explains the why using pilot numbers.\n- Handle resistance directly by publishing the own-your-code-own-your-pager rule and the paid on-call mechanics before each wave, not after.\n- Retire legacy tools, escalation lists and the informal status-page process at the end of each wave on a published cutover date.\n- Harvest feedback formally at each wave and push accepted changes back into the policy documents through change control.", "dependencies": ["S13", "S16", "S17"]}, {"step_id": "S19", "title": "SOC 2 incident-response control mapping and evidence framework", "description": "The audit will test incident response. Mapping controls now means evidence accumulates from the first real incident under the new process.\n\n- Map the process to the relevant Trust Services Criteria for incident identification, response, evaluation of incidents and communication of security events.\n- Write control statements in auditor language and name a single owner for each control.\n- Specify the evidence artifact for each control: incident record, severity classification, paging log, communications log, postmortem, action tracker entry and training record.\n- Set evidence retention and storage location so nothing depends on a laptop or on chat history that expires.\n- Run an early walkthrough with an experienced compliance partner or the auditor's readiness team to test the design before the audit window.\n- Flag controls that are blocked until rollout reaches the relevant teams, and keep a gap register with owners and dates reviewed fortnightly by the steering group.", "dependencies": ["S1", "S3", "S10", "S12"]}, {"step_id": "S20", "title": "SOC 2 dry run, gap remediation and audit support", "description": "This step converts a good process into a provable one, a few months before the auditors arrive.\n\n- Run an internal dry run six weeks before the audit window, sampling real incidents from the pilot and early waves against each control's evidence requirements.\n- Remediate every gap found, prioritising anything that risks a qualified opinion.\n- Brief every engineer who may be interviewed so they can describe the process as they actually practise it.\n- Prepare the auditor package: process documentation, sample incident records, training records, on-call schedules and the action tracking register.\n- Designate a single audit liaison and a small evidence-request team so requests do not land on 28 teams at once.\n- Rehearse the walkthrough with the incident commander and communications roles, since auditors probe realism under pressure.\n- Schedule the dry run and the audit alongside rollout waves so no team is hit with both in the same week.", "dependencies": ["S18", "S19"]}, {"step_id": "S21", "title": "Standing governance, process ownership and continuous improvement", "description": "The classic post-audit failure is that the process freezes and then decays. This step keeps it alive.\n\n- Establish a standing Incident Management Council chaired by the process owner, meeting monthly with engineering, support, compliance and product representation.\n- Give the process owner a documented mandate to change standards, and require formal change control for any change to severity, roles, communications timings or compensation.\n- Re-validate the severity taxonomy quarterly against real declarations and re-baseline the metric set annually.\n- Fund a resilience roadmap separate from incident response: reducing shared-database blast radius, multi-region failover readiness and deploy safety.\n- Report quarterly to the executive team on the metric set and on the top five systemic causes of incidents.\n- Feed incident themes into architecture review, release engineering and hiring so the programme improves the system, not just the response.\n- Refresh the training and tabletop programme annually and after any SEV1.", "dependencies": ["S18", "S20"]}], "estimated_complexity": "high", "success_metrics": "- Median time to detect reduced from 22 minutes to 5 minutes or less by month 6.\n- Customer-detected incidents reduced from 40 percent to 10 percent or less of all customer-impacting incidents by month 9.\n- Median time to mitigate reduced from 3 h 10 min to 60 minutes or less for SEV1 and SEV2 by month 9.\n- Customer-impacting incidents reduced by at least 40 percent year over year from the 31 incident baseline.\n- SLA credits paid reduced by at least 50 percent year over year from the $1.3M baseline.\n- 100 percent of SEV1 incidents have a named Incident Commander within 5 minutes and a scribe within 15 minutes.\n- Zero incidents in which command authority is unclear for more than 15 minutes.\n- 28 of 28 teams have a documented on-call rotation, an owning service map and at least one trained on-call engineer by month 6.\n- Central Incident Commander roster holds at least 40 certified ICs covering 24x7 with no single point of failure.\n- 100 percent of on-call rotations are paid under a published policy by month 5.\n- Monthly alert volume reduced from 3,400 to below 700, with a false-positive rate below 20 percent.\n- No service exceeds 2 pages per on-call shift, measured monthly for three consecutive months.\n- 100 percent of SEV1 and SEV2 postmortems published internally within 15 business days.\n- At least 90 percent of postmortem action items closed within 60 days, up from 17 percent (11 of 64).\n- Status-page first update published within 30 minutes on at least 95 percent of SEV1 incidents.\n- Zero missed regulatory notification windows on any incident requiring notification.\n- SOC 2 Type II audit passed with no findings related to incident response.\n- Review cadence sustained: weekly operational review in at least 90 percent of weeks, 12 of 12 monthly reliability reviews, 4 of 4 quarterly executive reviews.\n- On-call satisfaction at 70 percent or higher on the quarterly survey, with zero on-call-attributed voluntary attrition.\n- 100 percent of new engineers complete the incident-response onboarding module within 30 days of joining."}Twelve steps: executive charter, severity matrix tied to transaction failure rates, a single mandatory paid on-call pool for all 28 teams, RACI roles plus a Rapid Response Team for the shared PostgreSQL cluster, tool consolidation with runbook-or-no-page hygiene, synthetic transactions and auto-escalation, tight status-page timings (SEV1 first update in 5 minutes), mandatory 5-day postmortems, Jira-integrated action tracking, game days, and a month-by-month rollout to month 7. Metrics target MTTD <5 min and 95% internal detection.
The plan produced
1. Establish Governance and Executive Sponsorship
Secure explicit CEO and CTO backing to mandate the program, framing it as a revenue protection and compliance necessity rather than an engineering burden. Form a cross-functional Incident Management Task Force.
- Identify a single executive owner (e.g., VP of Engineering or Head of Platform) with authority over budget and staffing.
- Create a charter defining the scope: severity definitions, on-call compensation, and tool consolidation.
- Communicate the strategic priority to all 28 team leads to preempt pushback regarding pager responsibility.
2. Define Severity Matrix and Response SLAs (after 1)
Create a rigorous, binary severity scale to eliminate ambiguity in escalation and communication speed. Align internal response times with the 99.95% SLA.
- Sev1 (Critical): System-wide outage, data corruption, or >5% transaction failure rate. Immediate executive page.
- Sev2 (Major): Single customer segment impacted or >1% transaction failure rate. High-priority team response.
- Sev3 (Minor): Degraded performance or isolated bug. Standard business-hours follow-up.
- Sev4 (Cosmetic): No functional impact. Backlog item.
- Define maximum detection and mitigation times for each level.
3. Restructure On-Call Organization and Compensation (after 1)
Transform the current unpaid, fragmented on-call system into a professional, compensated, and centralized rotation to ensure 24x7 coverage across all 28 teams. Address the 'pushback' by making it mandatory and paid.
- Consolidate the 12 current on-call teams into a unified, rotating pool covering all 28 teams.
- Implement a fair compensation model: base stipend + hourly rate for active hours + incident bonus for Sev1/2.
- Define a strict 'Buddy System' or 'Follow-the-Sun' rotation if geographically feasible, otherwise ensure adequate team size to rotate fairly.
- Mandate that every team must participate in the rotation; no exemptions.
4. Define Incident Roles and RACI Model (after 2)
Eliminate the 'who is in charge' ambiguity by pre-defining roles and establishing a clear chain of command for every incident. Roles are assigned per incident, not permanently held.
- Incident Commander (IC): Single point of authority. Owns the timeline, delegates tasks, and approves communications. Must be senior.
- Scribe: Maintains the real-time timeline log of actions, timestamps, and decisions.
- Communications Lead: Owns external/status page updates and internal Slack broadcasts. Shields the IC from interruptions.
- Technical Leads (SMEs): Deployed by IC to troubleshoot specific services. They do not coordinate with each other; they report to IC.
- Create a 'Rapid Response Team' (RRT) for high-severity incidents involving the shared PostgreSQL cluster.
5. Consolidate Monitoring and Enforce Alert Hygiene (after 3)
Reduce the 3,400 monthly alerts (85% noise) to a signal-to-noise ratio that allows engineers to trust the pager. Standardize tooling.
- Consolidate the six disparate alerting tools into a single platform (e.g., PagerDuty, Opsgenie, or Datadog).
- Implement 'Actionable Alerting': If an alert pages, there must be a documented, immediate runbook action. No 'informational' pages.
- Set a hard cap on alert volume per engineer/night. Breaching the cap triggers a mandatory alert quality review.
- Deprecate legacy tools that cannot map to the severity matrix defined in S2.
6. Build Detection and Escalation Automation (after 5)
Reduce the 22-minute mean time to detect (MTTD) by automating the path from metric breach to page. Ensure customers never find an issue before the internal pager does.
- Implement synthetic transactions for core payment flows to detect failures before real users do.
- Configure escalation policies: If not acknowledged in 5 minutes, auto-escalate to the team lead, then to the Incident Commander pool.
- Map specific Kubernetes and PostgreSQL metrics to the Severity Matrix (S2) to trigger automated pages.
- Create a 'Major Incident Channel' in Slack that is automatically opened and pinned when a Sev1/2 is triggered.
7. Standardize Communication Templates and Timing (after 4, 6)
Codify the rules for external and internal communications to stop the 'whoever is around' approach. Link these timings to the Severity Matrix.
- Status Page Updates:
- Sev1: First update within 5 minutes of detection. Updates every 15 minutes.
- Sev2: First update within 30 minutes. Updates every 60 minutes.
- Customer Direct Mail: For Sev1 impacting specific accounts, Account Managers send pre-approved templates within 1 hour of mitigation.
- Regulator Notification: Define legal triggers for NY financial regulations.
- Create a 'Comms War Room' Slack channel restricted to the IC and Comms Lead.
- Status Page Updates:
8. Formalize the Postmortem Process (after 4, 7)
Transform postmortems from optional/various formats into a mandatory, blameless, and tracked process. This directly addresses the SOC 2 audit requirement for continuous improvement.
- Mandate a postmortem for every Sev1 and Sev2 incident within 5 business days.
- Enforce a 'Blameless' charter: Focus on systemic/process failures, not individual errors.
- Adopt a standardized template: Timeline, Root Cause Analysis (5 Whys or Fishbone), Impact Analysis, and Action Items.
- Action Items must have: Owner, Due Date, and Priority. They are tracked in Jira/Linear, not the doc.
9. Implement Action Item Tracking and Governance (after 8)
Solve the '11 of 64 closed' problem by integrating postmortem actions into the core engineering workflow and executive reviews. Ensure fixes are shipped.
- Integrate Jira with the Postmortem tool (e.g., OpsLevel, Rootly, or Notion API). Auto-create tickets for every AI.
- Block 'Sev1 Closure' until all high-priority Action Items are closed or explicitly deferred with executive approval.
- Weekly Engineering Leadership review of open postmortem actions older than 30 days.
- Tie action item completion rates to team performance reviews.
10. Deploy Training and Runbook Library (after 6, 7, 8)
Address the skill gap and the 'pushback' by equipping engineers with the tools and knowledge to succeed on-call. Reduce anxiety through clarity.
- Audit all 180 services. Require every service to have a 'Page-Worthy' runbook.
- Run 'Game Day' (Chaos Engineering) exercises quarterly to practice role assignments (IC, Scribe, Comms) in a simulated Sev1.
- Conduct mandatory training for the new on-call rotation and compensation model.
- Provide a 'Playbook' for the Incident Commander role, including scripts for communication.
11. Phased Rollout and Change Management (after 1, 2, 3, 4, 5, 6, 7, 8, 9, 10)
Introduce the new process in phases to manage cultural resistance and technical debt, ensuring adoption before the SOC 2 audit. Do not 'big bang' the change.
- Month 1-2: Launch Severity Matrix and Consolidated Tooling (S2, S5). Pilot on 5 high-traffic teams.
- Month 3-4: Roll out new On-Call Compensation and Roles (S3, S4) to all 28 teams. Enforce new Communication Templates.
- Month 5-6: Enforce strict Postmortem and Action Item Tracking (S8, S9).
- Month 7: Final audit readiness check. Simulate SOC 2 auditor questions on incident response.
12. Define Metrics for Success and Continuous Review (after 2, 6, 8, 9)
Establish a dashboard to prove the new system works, linking directly to the $1.3M savings in SLA credits. Monitor the health of the incident process itself.
- Track MTTD (Target: <5 min), MTTR (Target: <30 min for Sev2), and Detection Ratio (Target: >95% internal detection).
- Monitor 'Alert Noise' (target: <5% of alerts result in a non-actionable page).
- Track 'Action Item Completion Rate' (Target: 100% on-time closure for Sev1/2).
- Calculate 'SLA Credit Avoidance' vs. On-Call Program Cost to prove ROI to leadership.
- MTTD reduced to under 10 minutes with >90% detection occurring before customer reporting.
- MTTR for Sev1/Sev2 incidents reduced to under 60 minutes.
- SLA credits paid reduced by 50% within 12 months.
- On-call alert volume reduced by 80% (from 3,400 to <600 actionable alerts/month).
- 100% of Sev1/Sev2 incidents have a blameless postmortem completed within 5 days.
-
90% completion rate of postmortem action items within 30 days of assignment.
- Zero ambiguity in incident command: 100% of major incidents have a designated IC logged in the timeline.
- SOC 2 Type II audit passes incident response controls with zero critical findings.
[SYSTEM] You are an expert assistant in complex project planning. Your task is to generate a detailed and structured action plan to reach the main objective in the most professional and most detailed way, no matter how much work or steps will be needed to perform. Use your internal reasoning processes to think deeply about the problem and create the most comprehensive plan possible. Take as much time and space as you need to think through all aspects of the problem. After your thorough analysis, answer with the plan in the requested structure. Every text field you write will be read by a busy person, so write for them: short sentences, one idea each, in short paragraphs. Text fields accept Markdown: separate paragraphs with a blank line, use a bullet list when you enumerate things and plain paragraphs when you explain or argue, and bold at most one key phrase per paragraph. A text of more than three sentences must be split into paragraphs; never deliver a single unbroken block of text. [HUMAN] Main Objective: "A B2B payments platform in New York (260 engineers in 28 teams, 2,100 customers, $4B processed a month, a 99.95% SLA with service credits) runs about 180 services on Kubernetes across two AWS regions, with a shared PostgreSQL cluster for the ledger. In the last 12 months: 31 customer-impacting incidents, median time to detect 22 minutes (customers detected 40% of them first), median time to mitigate 3 h 10 min, $1.3M paid in SLA credits, two incidents in which nobody was sure who was in charge for over an hour, and a CEO email about "outages we hear about from clients". On-call exists in 12 of the 28 teams, unpaid, fed by alerts from six different tools — 3,400 alerts a month, 85% of them noise; there is no severity scale, status-page updates are written by whoever is around, and postmortems happen for some incidents, in various formats, with action items that are rarely tracked (11 of 64 closed). A SOC 2 Type II audit in eight months will test incident response. Engineers push back against "carrying a pager for other teams' code". Define the incident management process: severity levels and what each one triggers; roles (incident commander, communications, scribe, subject-matter responders) and how they are staffed 24x7 across 28 teams; the on-call structure, rotations, compensation and the rules for alert quality; detection and escalation paths; internal and customer communications (status page, account managers, regulators when required) with their timings; postmortems (when mandatory, format, blameless review, ownership and tracking of actions); the metrics and reviews that show whether it works; and how it is introduced across the teams without waiting for the audit." Generate an initial, detailed, and well-structured plan proposal. The plan has these parts: - "steps": the list of steps, each with "step_id" (a short label unique within this proposal, such as S1, S2, S3), "title" (a short, descriptive title), "description" (what the step does and how, in Markdown: a short opening paragraph and then, if there are several concrete things to say, a bullet list) and "dependencies" (the step_ids of the other steps of THIS proposal that must be completed first; empty if none). - "estimated_complexity": "low", "medium" or "high". - "success_metrics": clear and measurable success metrics, one per line as a bullet list. Think deeply about the problem, considering all aspects and implications. Develop a comprehensive understanding of what needs to be done to achieve the objective, then produce the plan.
{"steps": [{"step_id": "S1", "title": "Establish Governance and Executive Sponsorship", "description": "Secure explicit CEO and CTO backing to mandate the program, framing it as a revenue protection and compliance necessity rather than an engineering burden. Form a cross-functional Incident Management Task Force.\n\n- Identify a single executive owner (e.g., VP of Engineering or Head of Platform) with authority over budget and staffing.\n- Create a charter defining the scope: severity definitions, on-call compensation, and tool consolidation.\n- Communicate the strategic priority to all 28 team leads to preempt pushback regarding pager responsibility.", "dependencies": []}, {"step_id": "S2", "title": "Define Severity Matrix and Response SLAs", "description": "Create a rigorous, binary severity scale to eliminate ambiguity in escalation and communication speed. Align internal response times with the 99.95% SLA.\n\n- **Sev1 (Critical)**: System-wide outage, data corruption, or >5% transaction failure rate. Immediate executive page.\n- **Sev2 (Major)**: Single customer segment impacted or >1% transaction failure rate. High-priority team response.\n- **Sev3 (Minor)**: Degraded performance or isolated bug. Standard business-hours follow-up.\n- **Sev4 (Cosmetic)**: No functional impact. Backlog item.\n- Define maximum detection and mitigation times for each level.", "dependencies": ["S1"]}, {"step_id": "S3", "title": "Restructure On-Call Organization and Compensation", "description": "Transform the current unpaid, fragmented on-call system into a professional, compensated, and centralized rotation to ensure 24x7 coverage across all 28 teams. Address the 'pushback' by making it mandatory and paid.\n\n- Consolidate the 12 current on-call teams into a unified, rotating pool covering all 28 teams.\n- Implement a fair compensation model: base stipend + hourly rate for active hours + incident bonus for Sev1/2.\n- Define a strict 'Buddy System' or 'Follow-the-Sun' rotation if geographically feasible, otherwise ensure adequate team size to rotate fairly.\n- Mandate that every team must participate in the rotation; no exemptions.", "dependencies": ["S1"]}, {"step_id": "S4", "title": "Define Incident Roles and RACI Model", "description": "Eliminate the 'who is in charge' ambiguity by pre-defining roles and establishing a clear chain of command for every incident. Roles are assigned per incident, not permanently held.\n\n- **Incident Commander (IC)**: Single point of authority. Owns the timeline, delegates tasks, and approves communications. Must be senior.\n- **Scribe**: Maintains the real-time timeline log of actions, timestamps, and decisions.\n- **Communications Lead**: Owns external/status page updates and internal Slack broadcasts. Shields the IC from interruptions.\n- **Technical Leads (SMEs)**: Deployed by IC to troubleshoot specific services. They do *not* coordinate with each other; they report to IC.\n- Create a 'Rapid Response Team' (RRT) for high-severity incidents involving the shared PostgreSQL cluster.", "dependencies": ["S2"]}, {"step_id": "S5", "title": "Consolidate Monitoring and Enforce Alert Hygiene", "description": "Reduce the 3,400 monthly alerts (85% noise) to a signal-to-noise ratio that allows engineers to trust the pager. Standardize tooling.\n\n- Consolidate the six disparate alerting tools into a single platform (e.g., PagerDuty, Opsgenie, or Datadog).\n- Implement 'Actionable Alerting': If an alert pages, there must be a documented, immediate runbook action. No 'informational' pages.\n- Set a hard cap on alert volume per engineer/night. Breaching the cap triggers a mandatory alert quality review.\n- Deprecate legacy tools that cannot map to the severity matrix defined in S2.", "dependencies": ["S3"]}, {"step_id": "S6", "title": "Build Detection and Escalation Automation", "description": "Reduce the 22-minute mean time to detect (MTTD) by automating the path from metric breach to page. Ensure customers never find an issue before the internal pager does.\n\n- Implement synthetic transactions for core payment flows to detect failures before real users do.\n- Configure escalation policies: If not acknowledged in 5 minutes, auto-escalate to the team lead, then to the Incident Commander pool.\n- Map specific Kubernetes and PostgreSQL metrics to the Severity Matrix (S2) to trigger automated pages.\n- Create a 'Major Incident Channel' in Slack that is automatically opened and pinned when a Sev1/2 is triggered.", "dependencies": ["S5"]}, {"step_id": "S7", "title": "Standardize Communication Templates and Timing", "description": "Codify the rules for external and internal communications to stop the 'whoever is around' approach. Link these timings to the Severity Matrix.\n\n- **Status Page Updates**:\n - Sev1: First update within 5 minutes of detection. Updates every 15 minutes.\n - Sev2: First update within 30 minutes. Updates every 60 minutes.\n- **Customer Direct Mail**: For Sev1 impacting specific accounts, Account Managers send pre-approved templates within 1 hour of mitigation.\n- **Regulator Notification**: Define legal triggers for NY financial regulations.\n- Create a 'Comms War Room' Slack channel restricted to the IC and Comms Lead.", "dependencies": ["S4", "S6"]}, {"step_id": "S8", "title": "Formalize the Postmortem Process", "description": "Transform postmortems from optional/various formats into a mandatory, blameless, and tracked process. This directly addresses the SOC 2 audit requirement for continuous improvement.\n\n- Mandate a postmortem for every Sev1 and Sev2 incident within 5 business days.\n- Enforce a 'Blameless' charter: Focus on systemic/process failures, not individual errors.\n- Adopt a standardized template: Timeline, Root Cause Analysis (5 Whys or Fishbone), Impact Analysis, and Action Items.\n- Action Items must have: Owner, Due Date, and Priority. They are tracked in Jira/Linear, not the doc.", "dependencies": ["S4", "S7"]}, {"step_id": "S9", "title": "Implement Action Item Tracking and Governance", "description": "Solve the '11 of 64 closed' problem by integrating postmortem actions into the core engineering workflow and executive reviews. Ensure fixes are shipped.\n\n- Integrate Jira with the Postmortem tool (e.g., OpsLevel, Rootly, or Notion API). Auto-create tickets for every AI.\n- Block 'Sev1 Closure' until all high-priority Action Items are closed or explicitly deferred with executive approval.\n- Weekly Engineering Leadership review of open postmortem actions older than 30 days.\n- Tie action item completion rates to team performance reviews.", "dependencies": ["S8"]}, {"step_id": "S10", "title": "Deploy Training and Runbook Library", "description": "Address the skill gap and the 'pushback' by equipping engineers with the tools and knowledge to succeed on-call. Reduce anxiety through clarity.\n\n- Audit all 180 services. Require every service to have a 'Page-Worthy' runbook.\n- Run 'Game Day' (Chaos Engineering) exercises quarterly to practice role assignments (IC, Scribe, Comms) in a simulated Sev1.\n- Conduct mandatory training for the new on-call rotation and compensation model.\n- Provide a 'Playbook' for the Incident Commander role, including scripts for communication.", "dependencies": ["S6", "S7", "S8"]}, {"step_id": "S11", "title": "Phased Rollout and Change Management", "description": "Introduce the new process in phases to manage cultural resistance and technical debt, ensuring adoption before the SOC 2 audit. Do not 'big bang' the change.\n\n- **Month 1-2**: Launch Severity Matrix and Consolidated Tooling (S2, S5). Pilot on 5 high-traffic teams.\n- **Month 3-4**: Roll out new On-Call Compensation and Roles (S3, S4) to all 28 teams. Enforce new Communication Templates.\n- **Month 5-6**: Enforce strict Postmortem and Action Item Tracking (S8, S9).\n- **Month 7**: Final audit readiness check. Simulate SOC 2 auditor questions on incident response.", "dependencies": ["S1", "S2", "S3", "S4", "S5", "S6", "S7", "S8", "S9", "S10"]}, {"step_id": "S12", "title": "Define Metrics for Success and Continuous Review", "description": "Establish a dashboard to prove the new system works, linking directly to the $1.3M savings in SLA credits. Monitor the health of the incident process itself.\n\n- Track MTTD (Target: <5 min), MTTR (Target: <30 min for Sev2), and Detection Ratio (Target: >95% internal detection).\n- Monitor 'Alert Noise' (target: <5% of alerts result in a non-actionable page).\n- Track 'Action Item Completion Rate' (Target: 100% on-time closure for Sev1/2).\n- Calculate 'SLA Credit Avoidance' vs. On-Call Program Cost to prove ROI to leadership.", "dependencies": ["S2", "S6", "S8", "S9"]}], "estimated_complexity": "high", "success_metrics": "- MTTD reduced to under 10 minutes with >90% detection occurring before customer reporting.\n- MTTR for Sev1/Sev2 incidents reduced to under 60 minutes.\n- SLA credits paid reduced by 50% within 12 months.\n- On-call alert volume reduced by 80% (from 3,400 to <600 actionable alerts/month).\n- 100% of Sev1/Sev2 incidents have a blameless postmortem completed within 5 days.\n- >90% completion rate of postmortem action items within 30 days of assignment.\n- Zero ambiguity in incident command: 100% of major incidents have a designated IC logged in the timeline.\n- SOC 2 Type II audit passes incident response controls with zero critical findings."}Round 1 — refinement 1 of 2
P2's round-0 architecture became the de facto template: P1 and P3 both rebuilt their plans on it step-for-step, while P2 itself deepened its version with genuinely new mechanics (Triage Owner, severity x class, priced opt-out, three-action cap, regulator clock matrix, SOC 2 evidence clock). The round converged strongly on structure; what still separates the plans is depth of incident mechanics, realism of comms timings and where audit work sits in the sequence.
What still separates them
- Command mechanics between alert and declaration. P2 alone closes the gap that caused the two hour-long ownership failures: Triage Owner rule (step 5), Watch state with a 30-minute timer, "declaring is free", the ambiguity rule and the two-simultaneous-SEV-1 rule (step 6). P1 has an escalation ladder (step 9) but no owner-from-first-ack rule; P3 has no escalation or lifecycle step at all and dropped its round-0 5-minute auto-escalation.
- On-call shape. P2 builds three rotations including a paid Platform Duty for the shared PostgreSQL/Kubernetes estate (step 8) and prices opting out against a paid pool (step 9). P1 (steps 10–11) and P3 (step 5) stop at federated team rotations plus a central IC roster, leaving shared infrastructure ownership implicit.
- Customer communication timings and money. P1 demands a status-page update within 3 minutes and SEV-1 updates every 5 minutes (step 13) — contradicting its own success metric of 30 minutes. P2 uses 30/60 minutes (step 12) plus a regulator clock matrix naming NYDFS Part 500 and a customer-impact ledger driving SLA credit automation (step 13). P3 uses 15/30 minutes (step 9) with no credit process at all.
- Where audit work sits. P2 maps controls and defines the "golden incident file" in month one (step 3) on the argument that Type II evidence cannot be backfilled. P1 places control mapping at step 21, P3 at step 12 after postmortems. P3 also schedules game days (step 16) and the metrics dashboard (step 17) only after full rollout, so nothing is drilled or measured during the pilot.
Who took what from whom
- P1 and P3 rebuilt on P2's round-0 skeleton almost wholesale: charter (P2 s1), baseline evidence pack (s2), severity trigger matrix (s3), role cards (s4), SLO/synthetic detection (s6), "no runbook, no page" and the 90-day noise sprint (s7), federated on-call with a six-engineer floor (s8), paid on-call with compensatory rest (s9), pilot then four waves with readiness gates (s17, s18), control mapping and dry run (s19, s20).
- P1 kept only one structural idea of its own: severity playbooks with decision trees (its round-0 s9, now step 12); everything else in its 23 steps mirrors P2's ordering.
- P2 took P3's ROI framing (P3 s12, credit avoidance vs programme cost) into its step 13, and P1's mobile-pager requirement (P1 s5/s8) into step 11 ("run a SEV-1 from a phone at 3am").
- Nobody adopted P1's round-0 per-incident bonus (s4) — P2 explicitly bans pay attached to incident counts (step 9) — and nobody adopted P3's tying of action-item completion to performance reviews (s9), which P2 contradicts with a published amnesty.
- P1 carried over P2's round-0 metric of 40+ certified ICs, which P2 itself cut to 12–16 this round as more realistic for 260 engineers.
The calls of this round
Influences: who took what from whom
| Round 1 ↓ · round 0 → | Proposal 1 | Proposal 2 | Proposal 3 | New steps |
|---|---|---|---|---|
| Proposal 1 |
kept1 | same titles19 analyst sees+4 / −1 | same titles1 analyst sees+1 / −1 | new2 |
| Proposal 2 |
same titles0 analyst sees+2 / −1 | kept6 | same titles0 analyst sees+1 / −2 | new15 |
| Proposal 3 |
same titles1 analyst sees+1 / −1 | same titles8 analyst sees+4 / −2 | kept2 | new9 |
P1 abandoned its generic round-0 framework and adopted P2's structure nearly step-for-step, gaining a charter, baseline evidence pack, SLO-based detection, a paging contract, a federated on-call model and a pilot-then-waves rollout. It kept its own useful severity playbooks step. Residual weaknesses are internal inconsistencies in timings and severity definitions.
- Added step 1 (executive mandate, named program owner, $400–600K budget) and step 2 (12-month incident register, alert census, survey) — the round-0 plan started at severity definitions with no baseline.
- Added step 7 detection strategy: SLOs for 20 customer journeys, external synthetics in three locations, PostgreSQL-specific signals (replication lag, txid exhaustion), detection contract per service — directly attacks the 40% customer-detected figure that round-0 only listed as a metric.
- Step 8 replaces vague noise rules with a paging contract ("no runbook, no page"), page budgets, expiry dates on silences and a 90-day burn-down of the top 100 rules.
- Step 16 now requires artifact-based closure of action items and a repeat-incident design review, instead of round-0's self-reported Jira tracking.
- Metrics gained dates and process/health tiers (first-update timeliness, action median age, IC roster coverage).
- Step 3 announces "a special SEV0 for security/regulatory events" then never defines it — a dangling category borrowed from P2 without its content.
- Comms timings contradict the success metrics: step 13 promises a status-page update within 3 minutes on SEV-1, the metric list says 30 minutes on ≥95%.
- Step 8's "no service may exceed two pages per on-call shift per month" garbles P2's budget into an unmeasurable unit.
- Target of 40+ certified ICs by month 4 is implausible for a 260-engineer org and inflates the training load in step 18.
- Step 19 (pilot) carries 12 dependencies, so almost nothing can be validated early; SOC 2 control mapping lands at step 21, losing P2's point that Type II evidence must accumulate from month one.
- Proposal 2 : Programme charter with a single accountable owner, steering group, and a 12-month baseline evidence pack.
- Proposal 2 : Symptom-based SLO alerting on customer journeys plus the alert-quality standard and noise sprint.
- Proposal 2 : Federated ownership ("no team paged for code it does not own"), six-engineer rotation floor, central IC roster, paid on-call with documented rest and opt-out.
- Proposal 2 : Status page hosted outside the production failure domain, evidence-based action closure, SOC 2 control mapping and dry run.
- Proposal 3 : A hard cap on alert volume whose breach triggers a mandatory alert-quality review.
- Proposal 2 : Status-page first update within 30 minutes of a SEV-1.
- Proposal 3 : Blocking SEV-1 closure until high-priority action items are done and tying completion to performance reviews.
+ Executive mandate and governance structure+ Alert consolidation and event pipeline+ Playbooks and communication templates by severityProgramme charter, ownership and executive mandate
The plan produced
1. Executive mandate and governance structure new
Turn the CEO email into a funded, authorized program with clear ownership and decision rights. Without executive backing, every downstream decision stalls in negotiation.
- Appoint a single program owner (e.g., Director of Incident Management) reporting to the CTO and COO.
- Publish a one-page charter covering scope (all customer-impacting incidents), authority to override team preferences during incidents, and funding for tooling, training, and on-call compensation.
- Establish a standing Incident Management Steering Group with CTO, VP Engineering, VP Support, Head of Compliance, and one engineering manager per region meeting monthly.
- Secure budget envelope: tool licenses, training time, incident-response infrastructure, and on-call compensation (estimated $400–600K annually).
- Make incident-response participation a documented performance expectation for all engineering managers and team leads.
2. Baseline measurement and evidence pack (after 1) from P2 step 2
You cannot prove improvement without defensible baseline numbers, and you cannot win arguments about noise or impact without data.
- Build a 12-month incident register: date, detection source, impact scope, time to detect, time to mitigate, SLA credits paid, and services involved.
- Audit the current alert estate: count alerts per tool, per team, per service; compute page-to-action ratio; identify top 50 noisiest rules and off-hours interruptions per engineer.
- Survey on-call engineers and managers on burden, fairness, and escalation clarity (target >70% response rate).
- Interview Support and Account Management: how do customers discover incidents, what do they complain about?
- Document the two command-ambiguity incidents: exactly when unclear who was in charge and why, how long it lasted.
- Publish this pack internally as the problem statement and retain all artifacts for SOC 2 audit evidence.
3. Severity taxonomy and trigger matrix (after 2) from P2 step 3
Severity is the keystone. Every other rule—paging, communications, postmortems, compensation—keys off it. Define four levels plus a special SEV0 for security/regulatory events.
- SEV1 (Critical): Complete service outage, data corruption, or >5% payment-path failure rate for >5 min. Every minute costs money. IC required; 99.99% uptime threatened.
- SEV2 (Major): Significant degradation, single region loss, or 1–5% transaction failure. IC typically required; service credit exposure.
- SEV3 (Minor): Limited customer impact with workaround available, or internal issues affecting operations. On-call SME + escalation if SLA at risk.
- SEV4 (Cosmetic): Observations, non-impacting bugs, alerts. Alert-driven, no escalation unless pattern emerges.
- Specify automatic triggers: region loss, ledger write failures, missed settlement window, payment success rate thresholds.
- Define who may declare (any engineer, Support, account manager) and who may downgrade (IC only).
- Include worked examples from the last 12 months so teams recognize their incidents in the definitions.
- Review and re-validate quarterly against actual declarations.
4. Incident roles, command structure, and decision rights (after 3) from P2 step 4
The two incidents with >1 hour of command ambiguity prove this step is non-negotiable. Define clear roles with explicit decision authority.
- Incident Commander: Owns the incident timeline, not the fix. Declares severity, decides escalation, approves communications, freezes changes, calls responders. Non-technical ICs are acceptable.
- Deputy IC: Shadows IC; takes over if IC unavailable. Nominated within 5 minutes of incident declaration.
- Communications Lead: Owns internal Slack updates and status-page messaging. Shields IC from interruptions.
- Scribe: Records real-time timeline: decisions, who did what, key timestamps. Not responsible for fixing.
- Subject-Matter Responders: Engineers with service context. Take IC direction without debate. Report only to IC.
- Operations Lead (SEV1 only): Coordinates across multiple responders, manages incident bridge.
- Write one-page role cards with mission, decision authority, and escalation upward. Publish on wiki and link from every paging notification.
- Define minimum viable coverage per severity: SEV1 staffs all roles; SEV2 staffs IC, Comms, Scribe; SEV3 staffs IC + Scribe.
- Establish handover discipline: maximum 4-hour IC shifts on SEV1, written handover template required.
5. Incident tooling consolidation and integration (after 1, 3) from P2 step 14
Six alert tools and ad-hoc incident records are structural causes of the 22-minute detection and 3+ hour mitigation. Consolidate to a single incident platform that is the source of truth.
- Select an incident management platform (PagerDuty, Opsgenie, Incident.io, etc.) that supports paging, schedules, escalation, incident records, and postmortem workflow.
- Requirement: the platform must integrate with observability tools, auto-create and pin incident channels in Slack, auto-capture timeline from chat, and support API-driven playbook automation.
- Plan a dual-run period alongside legacy tools with a published cutover date; define rollback criteria.
- Integrate incident record with the 180 services' monitoring and dashboards so responders see everything in one place.
- Define data retention and audit trail to satisfy SOC 2 evidence requirements: who did what, when, under whose authority.
- Budget for licenses, migration effort, and two-week hardening period post-cutover.
6. Alert consolidation and event pipeline (after 5) new
Replace six alert sources with a single ingestion point. Deduplicate and route alerts with minimal manual judgment, removing a major source of detection delay.
- Consolidate alert endpoints from six tools into a single event pipeline; this often sits in front of the incident platform (S5).
- Implement deduplication and correlation so a single outage triggering alerts from five monitoring tools produces one page, not five.
- Map every alert to a severity level from S3 (SEV1, SEV2, SEV3, SEV4) at ingestion.
- Log every alert for postmortem analysis and trending.
- Ensure the platform's mobile app works reliably; on-call responders need to engage from any device.
7. Detection strategy: SLOs, signals, and customer-journey monitoring (after 3, 6) from P2 step 6
Customers detected 40% of incidents first—a detection gap that must be closed. Build symptom-based alerting that detects outages before customers do.
- Define SLIs and SLOs for the top 20 customer journeys: payment initiation, settlement, ledger read/write, API availability, webhook delivery, measured per region.
- Require symptom-based alerting on SLOs, not cause-based infrastructure metrics (e.g., alert on "payment success rate <99%" not "database CPU >80%").
- Add synthetic transaction monitoring from outside AWS in both regions and a third location, one-minute cadence, for all money-moving paths.
- Add ledger-critical signals: PostgreSQL replication lag, connection saturation, write latency, transaction ID exhaustion, checkpoint pressure.
- Create a detection contract per service: owner identified, at least one symptom alert defined, expected detect time documented.
- Open a customer-reported incident path: Support and account managers can declare an incident directly, counted as a detection source in metrics.
- Fund a separate resilience roadmap to reduce shared-database blast radius, because detection improvements do not protect against ledger corruption.
8. Alert quality standards and noise-reduction program (after 3, 6, 7) from P2 step 7
3,400 monthly alerts with 85% noise is the reason engineers resent the pager. Cutting noise is the price of admission for on-call buy-in.
- Publish alert standards: every page must be symptom-based, have an immediate runbook action, be owned by a team, and map to a severity level. No runbook, no page.
- Separate paging from ticketing: only customer-impacting or imminent-impact signals trigger pages. Everything else becomes a dashboard entry or a ticket.
- Set a noise budget per team and service: no service may exceed two pages per on-call shift per month. Breaching triggers a mandatory alert quality review.
- Define the default action for a noisy alert: fix the root cause, tune the threshold, or delete it within ten working days. Deletion is a legitimate successful outcome.
- Implement automatic suppression rules: silence alerts if service auto-recovered within 30 seconds; suppress known maintenance windows; group flapping alerts (>5 in 2 min) into one page; rate-limit noisy services (max 1 alert per 5 min until condition clears).
- Require expiry dates on all silencing rules so suppression cannot become permanent blindness.
- Run a 90-day noise sprint with a visible burn-down of the top 100 noisiest rules, each owned by a named manager.
9. Escalation policies and incident lifecycle (after 3, 4, 5, 6) from P2 step 5
Define the mechanical path from alert to incident declaration to resolution. Escalation must be automatic and blameless.
- Define incident lifecycle states with clear entry/exit criteria: Detected → Triaged → Declared → Mitigated → Resolved → Postmortem → Closed.
- Set acknowledgement targets: page acknowledged in 5 minutes; triage decision (is this real?) in 15 minutes; severity declaration (is this customer-impacting?) in 30 minutes.
- Build escalation ladders: if responder does not acknowledge in 5 min, page escalates to service owner, then team manager, then IC on-call. Escalation is automatic, not manual.
- Implement escalation for SEV1: IC paged via phone call + SMS + Slack + mobile; if not acknowledged in 2 min, Deputy IC paged simultaneously; Communications Lead pinged at same time.
- Define change freeze during SEV1 and SEV2: no deployments except to fix the incident. Freeze lifts only when mitigation is confirmed.
- Enforce one incident, one record: the incident record is the sole source of truth. Auto-capture timeline from Slack and bridge; never write timeline from memory later.
- Test all escalation paths weekly via synthetic page to on-call; adjust timings based on first month of operations.
10. On-call architecture and 24x7 coverage model (after 3, 4, 9) from P2 step 8
The answer to "carrying a pager for another team's code" is that every team carries its own, and the platform carries shared risk. Design a sustainable model.
- Adopt a federated model: every service has one owning team; that team's on-call carries its service's pager. No team is paged for code it does not own.
- State the consequence clearly: 16 of 28 teams currently have no on-call. They must either build one or formally transfer service ownership to a team that will.
- Create a separate, centrally staffed Incident Commander on-call roster drawn from trained senior engineers across all teams, covering 24x7 with no single point of failure.
- Merge small or low-traffic teams into shared rotations so no rotation has fewer than six engineers (below which coverage is unsustainable).
- Define primary and secondary per slot: secondary engages only on no-acknowledge or explicit request from IC.
- Align coverage across two AWS regions and New York business hours: one global IC rotation; service on-call aligned to service users' time zones.
- Define unresponsive-team escalation: 15 min without acknowledgement escalates to team manager; 30 min escalates to IC, who may direct any available engineer.
- Publish a coverage matrix of all 28 teams showing services, rotation size, and gaps, reviewed monthly.
11. On-call compensation, wellbeing, and sustainability policy (after 10) from P2 step 9
Unpaid on-call is the most cited reason for resistance. Settle compensation before rollout, not during it. Make it sustainable.
- Introduce paid on-call: a weekly stipend ($500–1,000) while on-call, regardless of incident volume, benchmarked to New York market rates.
- Pay event-based compensation: 1.5× hourly rate for time spent mitigating incidents during off-hours (minimum one-hour block per callout).
- Provide compensatory rest: no engineer works a normal business day after a night incident requiring >2 hours mitigation. Rest day is documented, not granted as a favor.
- Cap intrusion: define maximum off-hours pages per week (e.g., no more than three per shift). Mandatory review and escalation if exceeded.
- Offer a voluntary opt-out path for engineers with genuine constraints, balanced by explicit obligation that someone else is paid to cover.
- Include on-call expectation and compensation in job descriptions and hiring conversations so commitment is understood before joining.
- Publish the policy with an effective date before any team is asked to join a rotation.
- Review the policy every six months against actual page volumes, attrition rates, and survey feedback.
12. Playbooks and communication templates by severity (after 3, 4) from P3 step 7
Playbooks remove ambiguity and decision fatigue during incidents. Templates ensure consistent, compliant messaging.
- Create a one-page (or one-screen) playbook for each severity level: who gets paged (roles, order); first questions (is it real, how big, who knows); what IC declares first (status page text, account manager notification, regulatory trigger); escalation timeline.
- SEV1 playbook: Immediate IC + Comms + CTO notification; customer status every 5 minutes; sample message templates.
- SEV2 playbook: IC + Comms + tech lead notification; status every 15 minutes; decision tree for escalation to executive team.
- SEV3 playbook: On-call SME + Comms if customer-visible; status every 30 minutes or as resolved.
- SEV4 playbook: On-call SME only; update customers only if promised SLA is at risk.
- Include decision trees: is this SEV1 or SEV2? Is it our code or dependency? Escalate or containment?
- Prepare customer-communication templates pre-approved by legal and compliance: sample language for detection, impact, workaround, mitigation phases.
- Publish playbooks on wiki and embed links in incident management platform.
13. Internal and customer communications workflows (after 4, 9, 12)
Specify who informs whom, in what order, via what channel. Prevent gaps like "nobody knew who was in charge for an hour."
- Internal cadence: First update to #incidents Slack channel within 3 minutes of declaration (even if "investigating"). Updates every 5 minutes (SEV1), 15 minutes (SEV2), 30 minutes (SEV3) or when material change occurs.
- IC calls CTO/VP Eng and incident channel lead within 1 minute of declaration (SEV1); incident declared in Slack with severity, IC name, and service affected.
- SME on-call for the failing service joins incident bridge automatically; escalation call includes them within 5 minutes.
- Designate a single Customer Communications Lead per incident (pre-identified on-call roster) who owns external messaging exclusively. Shields IC from customer contact.
- Customer communication: Status page updated within 3 minutes (SEV1) or 10 minutes (SEV2) even if root cause unknown; post "Investigating" with next-update ETA. Account managers of affected top-tier customers called within 5 minutes (SEV1) with templated language.
- Create a phone tree or escalation list accessible to responders; set expectation: "If you don't hear from IC in 2 minutes, call them."
- Use a single incident Slack channel per incident (auto-created by incident tool); log all communications for postmortem review.
- Define regulatory notification path: compliance must approve before sending, but do not wait for root cause; flag incidents triggering payment-processing regulations to legal immediately.
14. Status page, customer notifications, and account-manager playbook (after 5, 12, 13) from P2 step 11
Policy without tooling collapses at 3 AM. Make publishing a five-minute action.
- Upgrade or replace status page so components map to customer journeys ("payments", "settlements", "ledger") not internal services. Allow customers to subscribe per component.
- Integrate incident tool (S5) with status page so incident record drives updates and public timeline auto-populates.
- Provide one-click templates pre-filled with severity, impact language, and next-update time; reduce typing and errors.
- Create account-manager playbook: contact tree for top 50 customers, what they may say (facts only), what they must not say (speculation, blame, false ETAs), escalation path if customer escalates.
- Define SLA credit process end to end: impact detection → credit calculation (based on duration × severity) → approval → customer notification → finance treatment. Automate where possible.
- Host status page outside production failure domain on separate infrastructure so it survives total platform outage.
- Test status-page reliability during game days (S17), including simulated status-page outage and total region loss.
15. Postmortem policy, blameless process, and facilitation (after 3, 4) from P2 step 12
Only 11 of 64 action items closed means postmortems are currently a writing exercise. Rebuild around learning and tracking.
- Make postmortems mandatory: all SEV1 and SEV2, all SEV3 with customer impact or repeat pattern, any near-miss the IC flags.
- Set deadlines: draft within 5 business days, blameless review within 10 days, internal publication within 15 days.
- Adopt a single standardized template: impact and duration, timeline (detection through resolution), root cause (not "human error" but system condition that enabled error), contributing factors, what went well, what went badly, action items.
- Train a pool of blameless facilitators; require a trained facilitator (not the IC) for every SEV1 and SEV2 review.
- Prohibit counterfactual and blame language in postmortems; require contributing factors addressing tooling, process, organization, and human factors.
- Limit action items to small number of concrete, verifiable, owned changes with dates.
- Publish all postmortems internally by default, with security review only for genuinely sensitive material (e.g., unpatched vulnerability details).
- For SEV1 incidents affecting regulated customers, produce a variant customer-facing root cause report.
16. Action item tracking, reliability backlog, and completion governance (after 15) from P2 step 13
A postmortem without durable action tracking is a complaint. Solve the 11-of-64 problem.
- Create a single reliability backlog in the engineering tracker (Jira, Linear, etc.) with mandatory label, owner, due date, and link to originating incident.
- Define closure criteria: evidence required (merged code change, tested alert, verified drill) not self-reported status.
- Protect capacity: reserve a fixed percentage of each team's sprint (e.g., 10%) for reliability work; track unspent capacity and report to VP Engineering.
- Run a weekly ageing review of open actions; escalate anything overdue by >30 days to team lead and VP Engineering.
- Require IC or postmortem facilitator to sign off on action completion.
- Report completion rate and median action age in monthly incident review (target: >90% closed within 60 days).
- If the same service repeats an incident in the same area, trigger a design review rather than another action item; break the cycle.
17. Metrics, dashboards, and review cadence (after 2, 3) from P2 step 16
Measure to prove the system works. Publish dashboards so everyone sees the scoreboard.
- Define outcome metrics: time to detect (by source, target <5 min internally detected); time to mitigate SEV1/SEV2 (target <60 min); customer-detected incidents per month (target <2); SLA credits paid (target <$100K/year by month 12).
- Define process metrics: declaration latency, page acknowledgement rate, IC roster coverage (no single point of failure), first-update timeliness (% within SLA), update-cadence adherence.
- Define health metrics: alert volume and noise ratio per team, off-hours pages per engineer per month, postmortem timeliness, action closure rate and median action age.
- Build live dashboards visible to every engineer (not just managers); auto-populate from incident tool, update daily.
- Institute review cadence: weekly operational review (30 min, prior week incidents), monthly reliability review (trends, top causes, action status), quarterly executive review (CEO's office, SLA cost, systemic changes).
- Baseline every metric against S2 evidence pack; set 90-day and 12-month targets.
- End every review with decisions and owners, not just numbers.
18. Training, certification, and exercise program (after 4, 12, 13, 14, 15) from P2 step 15
A process that exists only on a wiki fails on the first real page. Build skills before deployment.
- Build curriculum: how to be on-call, how to declare an incident, how to run incidents as IC, how to communicate, how to write blameless postmortems.
- Create role-specific tracks: (1) All engineers: 30-min video on severity levels, communication expectations, postmortem format, action items (mandatory); (2) On-call responders: 2-hour workshop on playbooks (S12), alert tool (S5), escalation paths (S9), when to call manager, case studies, mobile app walkthrough; (3) Incident commanders: 4-hour IC bootcamp on leadership under pressure, decision-making, communicating with executives, status-page discipline, postmortem facilitation, practiced drills; (4) Communications leads: 2-hour training on templates, update timings, how to talk to customers, regulatory rules.
- Require certification before joining IC on-call roster: written assessment plus live simulated incident (pass/fail).
- Record videos so async teams can learn on their schedule; create runbooks and quick-reference cards (print + digital); pair new on-call engineers with experienced responder for first week.
- Run monthly tabletop exercises on realistic scenarios from the prior 12 months: region loss, ledger corruption, cascading failures.
- Run quarterly game days with intentional failure injection (database failover, status-page outage, alert tool downtime); include all on-call roles.
19. Pilot with volunteer teams (after 5, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 18) from P2 step 17
Do not roll out untested to 28 teams. Run the entire process end to end with a small cohort first.
- Recruit 3–4 volunteer teams covering a mix of criticality: one payment-path team, one ledger-adjacent team, one platform team, one low-traffic team.
- Run complete process in pilot: new severity scale (S3), consolidated tooling (S5, S6), roles (S4), escalation (S9), communications (S13, S14), postmortems (S15), action tracking (S16), paid on-call (S11), training (S18), metrics (S17).
- Instrument pilot against S17 metrics; compare results with S2 baseline.
- Hold weekly retrospectives with pilot teams; iterate on written policies, tooling, training based on feedback.
- Fix top issues found before wider rollout; document what changed and why.
- Produce pilot report with before/after numbers (MTTD, MTTR, alert noise, action completion rate) to carry into rollout conversations.
- Set explicit pilot exit criteria: rotation coverage achieved, zero unacknowledged pages over 2 weeks, all postmortems delivered on time, >80% of action items tracked.
20. Phased rollout to all 28 teams (after 16, 17, 19) from P2 step 18
Rollout is a staged migration with readiness gates, not an email announcement.
- Sequence the 28 teams into four waves of ~7 teams each, ordered by customer-impact criticality and readiness; space waves three weeks apart.
- Define per-team readiness checklist: services mapped and owned, alerts cleaned to standard (S8), runbooks written, rotation staffed, training complete, manager briefed.
- Hold gate review with program owner before each team joins; move unready teams to next wave with dated remediation plan.
- Assign named champion per wave; run internal communications cadence explaining why using pilot numbers from S19.
- Handle resistance directly: publish the "own-your-code, own-your-pager" rule and paid on-call mechanics before each wave, not after.
- Retire legacy tools, informal escalation lists, and ad-hoc status-page process at end of each wave on published cutover date.
- Harvest feedback formally at each wave and push accepted process changes through change control.
21. SOC 2 control mapping and evidence framework (after 1, 3, 13, 15) from P2 step 19
The audit will test incident response. Mapping controls now means evidence accumulates from the first real incident under the new process.
- Map the process to Trust Services Criteria for incident identification, response, evaluation, and communication of security incidents.
- Write control statements in auditor language; name a single owner per control.
- Specify the evidence artifact for each control: incident record, severity classification, paging log, communications log, postmortem, action tracker entry, training record.
- Define evidence retention and storage location (not on laptops, not on chat history that expires); plan for audit access.
- Identify which controls are blocked until certain rollout waves complete; keep a gap register with owners and review fortnightly with steering group.
- Run early walkthrough with an experienced compliance partner or pre-audit readiness team to test control design before the audit window.
22. SOC 2 dry run, gap remediation, and audit support (after 20, 21) from P2 step 20
Convert a good process into a provable one, a few months before auditors arrive. Prove the system works at scale.
- Schedule a dry run six weeks before audit window, sampling real incidents from completed waves against each control's evidence requirements.
- Remediate every gap found; prioritize anything risking a qualified opinion.
- Brief every engineer who may be interviewed so they can describe the process as practiced.
- Prepare auditor package: process documentation, sample incident records, training records, on-call schedules, action tracking register, alert quality metrics.
- Designate a single audit liaison and small evidence-request team so requests do not scatter across teams.
- Rehearse IC and Communications Lead roles under interview conditions; auditors probe realism under pressure.
- Ensure all postmortems, incidents, and evidence are retained, searchable, and accessible to auditors for the required audit period.
23. Standing governance, process ownership, and continuous improvement (after 20, 22) from P2 step 21
The classic post-audit failure is the process freezing and decaying. Lock in continuous improvement as a permanent structure.
- Establish a standing Incident Management Council chaired by the program owner, meeting monthly with engineering, support, compliance, and product representation.
- Give program owner documented mandate to change standards; require formal change control for any change to severity, roles, communications timings, or compensation.
- Re-validate severity taxonomy quarterly against real declarations; re-baseline metrics annually.
- Feed incident themes into architecture review and hiring so the program improves the system, not just the response.
- Report quarterly to executive team on metric set (MTTD, MTTR, SLA credits, customer-detected %) and top five systemic causes of incidents.
- Fund a resilience roadmap separate from incident response: reducing shared-database blast radius, multi-region failover, deploy safety.
- Refresh training and tabletop program annually and after any SEV1.
- Keep public backlog of improvement ideas (Jira, shared doc) so any engineer can suggest; publish monthly "incident newsletter" to all engineers with wins and learnings.
- Median time to detect reduced from 22 minutes to under 5 minutes by month 6, with >90% internal detection (vs. 40% customer-detected now).
- Median time to mitigate reduced from 3h 10min to under 60 minutes for SEV1 and SEV2 by month 9.
- Customer-impacting incidents detected by customers drop from 40% to <5% of all customer-impacting incidents.
- Alert volume reduced from 3,400 per month to <600 per month; signal-to-noise ratio improves from 15:85 to >95:5.
- SLA credits paid reduced from $1.3M to <$100K annually by month 12.
- Zero incidents with command-and-control ambiguity lasting >15 minutes; all SEV1/2 incidents have named IC logged in timeline within 5 minutes.
- Postmortem action item completion rate reaches >80% (from 11 of 64, or 17%) by month 4.
- 100% of SEV1 and SEV2 postmortems published within 15 business days by month 5.
- All 28 teams integrated into incident management system with active on-call rotations by week 20; no team unresponsive to pages for >30 minutes.
- On-call satisfaction score reaches >7/10 on survey; zero on-call-attributed voluntary attrition by month 6.
- Incident commander roster: 40+ certified ICs covering 24x7 with no single point of failure by month 4.
- Status-page first update published within 30 minutes on ≥95% of SEV1 incidents by month 3.
- SOC 2 Type II audit passes incident response controls with zero findings by month 8.
- Weekly incident review cadence sustained in ≥90% of weeks; monthly reliability reviews 12 of 12; quarterly executive reviews 4 of 4.
- 100% of new engineers complete incident-response onboarding within 30 days of joining.
[SYSTEM]
You are an expert assistant in complex project planning.
Your task is to generate a detailed and structured action plan to reach the main objective in the most professional and most detailed way, no matter how much work or steps will be needed to perform.
Use your internal reasoning processes to think deeply about the problem and create the most comprehensive plan possible.
Take as much time and space as you need to think through all aspects of the problem.
After your thorough analysis, answer with the plan in the requested structure.
Every text field you write will be read by a busy person, so write for them: short sentences, one idea each, in short paragraphs. Text fields accept Markdown: separate paragraphs with a blank line, use a bullet list when you enumerate things and plain paragraphs when you explain or argue, and bold at most one key phrase per paragraph. A text of more than three sentences must be split into paragraphs; never deliver a single unbroken block of text.
[HUMAN]
Main Objective: "A B2B payments platform in New York (260 engineers in 28 teams, 2,100 customers, $4B processed a month, a 99.95% SLA with service credits) runs about 180 services on Kubernetes across two AWS regions, with a shared PostgreSQL cluster for the ledger. In the last 12 months: 31 customer-impacting incidents, median time to detect 22 minutes (customers detected 40% of them first), median time to mitigate 3 h 10 min, $1.3M paid in SLA credits, two incidents in which nobody was sure who was in charge for over an hour, and a CEO email about "outages we hear about from clients". On-call exists in 12 of the 28 teams, unpaid, fed by alerts from six different tools — 3,400 alerts a month, 85% of them noise; there is no severity scale, status-page updates are written by whoever is around, and postmortems happen for some incidents, in various formats, with action items that are rarely tracked (11 of 64 closed). A SOC 2 Type II audit in eight months will test incident response. Engineers push back against "carrying a pager for other teams' code".
Define the incident management process: severity levels and what each one triggers; roles (incident commander, communications, scribe, subject-matter responders) and how they are staffed 24x7 across 28 teams; the on-call structure, rotations, compensation and the rules for alert quality; detection and escalation paths; internal and customer communications (status page, account managers, regulators when required) with their timings; postmortems (when mandatory, format, blameless review, ownership and tracking of actions); the metrics and reviews that show whether it works; and how it is introduced across the teams without waiting for the audit."
For your consideration and refinement, here are proposals from the previous round:
Previous Proposal 1 (ID: 895aa761-b883-4277-9609-6e76ca669198, Agent: claudeHaiku4.5_initial_1, LLM: anthropic/claude-haiku-4-5):
Estimated Complexity: high
Success Metrics: - Median time to detect reduced from 22 minutes to <8 minutes within 6 months
- Customer-detected incidents drop from 40% to <5% within 6 months
- Median time to mitigate (MTTR) reduced from 3h 10min to <45 minutes for SEV-1 incidents within 6 months
- Annual SLA credits decrease from $1.3M to <$100k within 12 months
- Alert noise reduced from 3,400 per month (85% false positive) to <400 per month (>95% signal) within 3 months
- Zero incidents with command-and-control ambiguity (>1 hour without clear IC) within 2 months
- Postmortem action item completion rate reaches >80% (from 17%) within 4 months
- On-call satisfaction score (survey) reaches >7/10 for on-call engineers within 3 months
- All 28 teams integrated into incident management system with active on-call rotations by week 20
- SOC 2 Type II audit passes incident response controls with no findings 8 months from start
- Incident commander certification: 100% of active ICs trained and drilled within 2 months
- Monthly incident review meeting established and attended by leadership; trends documented
- New incident system integration complete: single alert tool, single dashboard, all 180 services feeding in, <5 min deployment
Steps (21):
1. Define severity levels and decision criteria
Create a four-tier severity framework (SEV-1 through SEV-4) that guides all downstream decisions about response, escalation, and communications.
Each level must specify: customer impact (revenue at risk, customers affected, data loss risk); financial threshold triggering service credits; whether an incident commander is required; response time SLA (e.g., SEV-1 < 5 min notification, SEV-4 < 2 hours); and the go-live decision tree (when to declare and when to resolve).
- SEV-1: Complete service down or critical path broken for >5% of customers; every minute costs money; IC required; 99.99% uptime threatened
- SEV-2: Significant degradation, features unavailable, affecting 1–5% of customers; IC typically required
- SEV-3: Minor impact, limited customer footprint or workaround exists; escalation path but not automatic IC
- SEV-4: Observations or minor issues; alert-driven, no escalation unless pattern emerges
2. Define incident roles and responsibilities (depends on: 1)
Create the organizational roles that operate during an incident: who is in charge, who talks to customers, who writes down what happened, who fixes the system, and how decisions are made under pressure.
Each role must have a single-sentence mission, decision authority, and escalation upward.
- Incident Commander: owns decision-making and timeline; declares severity; resolves conflicts; may or may not be technical
- Deputy IC: shadow to IC and takes over if IC becomes unavailable
- Communications Lead: writes status page, notifies account teams, manages customer perception
- Scribe: records decisions, who did what, key timestamps; not responsible for fixing
- SME Responders: engineers with context on the failing service(s); take IC's direction without debate
3. Design 24x7 on-call rotation structure (depends on: 2)
Build a rotation model that covers all 28 teams with primary and backup on-call engineers every hour across weekdays, evenings, weekends, and holidays; addresses the pager-carrying resistance.
Key design decisions: Is coverage per-team (each team owns its services) or pooled (shared responder pool handles anything)? How many people per rotation? How long are shifts (one week, two weeks)? When can engineers opt out without leaving the team exposed? Which roles are on-call (IC, communications, SME)?
- Recommend: dedicated IC pool (4–6 people in fast rotation) + per-team SME on-call for each team's own services
- Recommend: two-week rotation blocks to reduce handoff friction
- Recommend: one primary, one secondary per slot; secondary handles during primary's escalation
- Provide swaps, blackout dates, and a rule that no engineer is on-call more than 2 weeks per quarter
4. Define on-call compensation and incentives (depends on: 3)
Create a pay model that makes on-call acceptable and rewards engineers who carry the pager; ties compensation to real business risk.
- Base on-call stipend: e.g., $500–1,000 per week while on-call (regardless of incidents)
- Callback pay: 1.5× hourly rate for time spent mitigating incidents during off-hours
- Incident bonus: $50–100 extra per SEV-1 or SEV-2 incident mitigated (recognition)
- Comp time: full business day off after an incident that required >2 hours mitigation during night/weekend
- Annual bonus tie-in: 10–20% bonus multiplier for flawless on-call reviews
- Communicate: position as investment in reliability, not punishment for being online
5. Consolidate alert routing infrastructure
Replace six alert tools with a single ingestion and routing system; stop engineers from being woken by duplicate alerts, and make escalation automated instead of manual.
Evaluate existing tools (likely candidates: PagerDuty, Opsgenie, or Incident.io) or build a lightweight wrapper. The system must: accept alerts from all 180 services; deduplicate and correlate (same outage, different monitoring source); route to correct on-call engineer; expose an API for playbook automation; log every alert for postmortem analysis.
- Choose tool by week 2 of S5
- Migrate alerting endpoints from 6 sources to 1 by week 4
- Set up audit trail and retention
- Ensure mobile app works (on-call engineers need to engage from phone)
6. Build alert quality rules to cut noise (depends on: 1, 5)
Implement rules that automatically suppress the 85% of alerts that are noise (flapping, transient errors, auto-recovered conditions). Target: <400 actionable alerts per month.
Rules to implement: suppress alerts if service auto-recovered within 30 seconds; deduplicate same alert from multiple monitoring sources; suppress alerts for known maintenance windows; group flapping alerts (same service, >5 occurrences in 2 minutes) into a single page to on-call; rate-limit alerts from noisy services (e.g., max 1 alert per 5 minutes per service until silence clears).
- Audit existing 3,400 alerts per month: which are true signals, which are noise
- Tag each alert source with severity level (S1, S2, S3, S4 from S1)
- Create exceptions list: services known to be noisy, require different rules
- Weekly review: alert teams that trigger >50 alerts per week for reduction strategies
7. Implement automated detection and escalation paths (depends on: 1, 5, 6)
Wire the alert system to automatically escalate based on time or severity; removes the need for manual judgment calls during chaos.
Logic: SEV-1 alert arrives → IC notified instantly via phone call + SMS + Slack + mobile; if IC does not acknowledge within 2 minutes, page deputy IC; Communications Lead pinged simultaneously. SEV-2: on-call SME for that service + IC notify via Slack and mobile, escalate to IC's manager if not acknowledged in 10 min. SEV-3/4: on-call SME only, escalate after 30 min.
- Implement in alert routing system (S5)
- Test all paths weekly via synthetic page to on-call
- Track escalation metrics: how many pages reach secondary, how many hit manager
- Adjust timing based on first month of operations
8. Build incident dashboard and status tracking (depends on: 5)
Create a single source of truth during an incident that every responder sees in real time: who is on-call, incident timeline, who said what, current status, next steps.
Dashboard displays: active incidents and their severity; who is the IC and communications lead; timeline of all events (alert fired, IC assigned, customer notified, mitigation started, resolved); Slack channel and mobile notification status; on-call rosters (who is on-call right now for each team); postmortem link as soon as incident closes.
- Integrate with alert tool (S5) to auto-populate incident creation and initial severity
- Push updates to status page and customer account managers automatically
- Log all timeline entries for audit and postmortem completeness
- Mobile-optimized so IC can work from any device
9. Write incident playbooks for each severity (depends on: 1, 2, 3)
Create a one-page (or one-screen) reference for the IC and SMEs during an incident; sequences the steps and removes ambiguity.
Each severity level gets its own playbook: who gets paged (roles, order); first questions to ask (is it real, how big, who knows); what the IC should declare in first message (status page text, account manager notification, regulatory trigger); how long before escalating to executive team; decision rules for going dark vs. continuing to update customers.
- SEV-1 playbook: immediate IC + comms + CTO notification; customer status every 5 minutes
- SEV-2 playbook: IC + comms + tech lead notification; status every 15 minutes
- SEV-3 playbook: on-call SME + comms if customer-visible; status every 30 min or as resolved
- SEV-4 playbook: on-call SME only; update customers only if promised SLA is at risk
- Include decision trees: is this SEV-1 or SEV-2? Is it our code or dependency? Escalate or containment?
10. Define internal communication workflows (depends on: 2, 3)
Specify who informs whom, in what order, via what channel (call, Slack, email) during an incident; prevents gaps like "nobody knew who was in charge for an hour."
Workflow for SEV-1: IC assigned → IC calls CTO/VP Eng and incident channel lead within 1 minute; incident declared in #incidents Slack channel with severity, IC name, service affected; SME on-call for that service joins call automatically; IC pushes updates to #incidents every 5 minutes or when material change occurs. For SEV-2: IC notifies team leads via Slack, updates #incidents every 15 min. Define escalation: if IC is unreachable, deputy IC takes over and announces it.
- Create a phone tree or on-call list accessible to responders
- Set expectations: "If you don't hear from IC in 2 minutes, call them"
- Use a single incident Slack channel per incident (auto-created by incident tool)
- Log all comms in the incident dashboard for postmortem review
11. Design customer communication and status page process (depends on: 1, 2)
Plan when and how to inform customers, account managers, and regulators; ensure 2,100 customers are not learning about outages from Twitter before you tell them.
Rules by severity: SEV-1 detected → status page updated within 3 minutes (even if root cause unknown; post "Investigating"); account managers of affected customers called within 5 minutes; regulatory notification (if payment processing down) queued for approval; customer email within 10 minutes with ETA for next update. SEV-2: status page within 10 min, account managers called within 15 min, email if affecting >10 customers. SEV-3/4: no customer communication unless SLA at risk.
- Empower Communications Lead to update status page without IC approval if delay >3 min
- Prepare templated messages for common scenarios (database failover, data pipeline stuck, service crashed)
- Route regulatory notifications through legal/compliance; don't wait for perfect root cause
- Track customer impact in real time: how many customers affected by severity
12. Establish blameless postmortem process and format (depends on: 1, 2)
Build a systematic way to learn from incidents so the same failure does not happen twice; counter the fear that admitting a mistake leads to being blamed.
Mandatory postmortems: all SEV-1 and SEV-2 incidents, within 48 hours of resolution. Optional but encouraged: SEV-3 if interesting or if >3 of same type in 30 days. Format: what was the user-visible impact and for how long; what was the root cause (not "human error" but the system condition that made error possible); timeline of discovery and response; action items with owner and deadline; blameless tone (focus on process and system design, not individual mistakes).
- Assign a facilitator (not the on-call IC) to run postmortem
- Attendees: IC, comms lead, SMEs involved, team lead, customer success if customer-facing
- Write postmortem in shared doc; make it findable (searchable, linked from incident)
- No discussion of "who screwed up"; only "why did the system allow this to happen"
13. Build action item tracking and accountability (depends on: 12)
Create a system that tracks postmortem action items so they are not forgotten; currently 11 of 64 (17%) are being tracked, leaving 53 unfinished improvements.
System: each postmortem generates action items (e.g., "add monitoring for X," "update runbook for Y," "write test for Z"). Each item gets: clear description, owner (engineer's name), due date (1–4 weeks based on priority), severity (critical = must do before similar incident happens again; important = improve next month; nice-to-have = backlog). Action items live in a dedicated Jira project visible to all teams; owners are accountable (their manager reviews quarterly). Weekly: incident commander reviews open items due that week. Monthly: each team's postmortem items reviewed in their standup.
- Export action items from postmortem document to tracking system automatically
- Require IC to sign off that an action is complete before closing
- Report on completion rate as a metric (target: >80% by month 3)
14. Define incident metrics and KPIs
Establish what "good" looks like; measure so you can improve. Target metrics for 12 months out: mean time to detect 8 minutes (vs. 22 now), customers detect first <5% of incidents (vs. 40%), MTTR 45 minutes for SEV-1 (vs. 190), SLA credits <$100k/year.
Metrics to track: (1) MTTD = time from incident start to first alert/report; disaggregate: external report vs. internal detection. (2) MTTR = time from first report to full mitigation; track by severity and by service. (3) Customer-reported incidents per month (should drop to <2 per month). (4) Alert signal-to-noise ratio (goal: <5% false positive after S6 rules). (5) On-call satisfaction (survey: would you do this again?). (6) Postmortem action completion rate. (7) Incident commander and responder utilization (hours per week per person).
- Dashboard: auto-populated from incident tool, updated daily
- Disaggregate by team and service: which teams have bad MTTR? Which service is most incident-prone?
15. Create review cadence and governance process (depends on: 14)
Establish regular rhythm to inspect the metrics, spot trends, and adjust the process itself; prevent the system from calcifying.
Weekly: incident commander and on-call lead review prior week—number of incidents, any escalations, any communication gaps. Monthly: director-level incident review—trends by service, top causes of incidents, action item status, whether severity classification is working. Quarterly: full leadership review—MTTD, MTTR, customer impact, SLA credit spend, on-call satisfaction score, any systemic changes needed. Annually: audit the entire process for SOC 2 compliance.
- Assign meeting owners: weekly = on-call lead; monthly = director of reliability; quarterly = VP Eng + CFO (SLA cost) + customer success
- Use same data dashboard (S14) for all reviews
- Publish a monthly "incident newsletter" to all engineers: what happened, what we learned, what's improving
16. Prepare SOC 2 Type II audit checklist (depends on: 1, 2, 9, 12, 13, 14, 15)
Document that the incident management system meets the control requirements for a SOC 2 audit; audit happens in 8 months, so this work builds confidence in coverage.
Audit will test: (1) Is there a defined incident response process? (2) Are roles and responsibilities clear? (3) Are incidents logged and tracked? (4) Is root cause analysis performed? (5) Are action items tracked and completed? (6) Is on-call staffing adequate? (7) Are communications timely? (8) Are postmortems documented and blameless? Create a control mapping document that links each SOC 2 requirement to your process (S1–S15). Collect evidence: incident logs, postmortem documents, action item tickets, metrics reports, training records.
- Designate a compliance owner (often a reliability lead or security engineer)
- Run a mock audit at month 6 to identify gaps
- Ensure all postmortems and incidents are retained and searchable for auditor review
17. Develop implementation and rollout plan (depends on: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16)
Create a phased timeline to roll out the incident management system across all 28 teams; avoids big-bang failure and builds credibility.
Recommended structure: Phase 1 (weeks 1–4): build and test infrastructure (S5, S8, alerting); deploy severity levels and roles (S1, S2); pick pilot teams (2–3 high-traffic teams). Phase 2 (weeks 5–12): train pilot teams, run incident drills, refine playbooks based on learning; expand to half of remaining teams. Phase 3 (weeks 13–20): full rollout to all 28 teams; continue drills; track metrics. Phase 4 (weeks 21–28): stabilize, iterate on metrics, prepare for audit.
- Assign a release manager to coordinate across teams
- Create a detailed Gantt chart with swim lanes (infra, process, training, rollout)
- Identify risks: competing priorities, engineers worried about pager burden, tool adoption friction
- Plan stakeholder engagement: weekly updates to eng leadership, monthly town halls for all engineers
18. Build training and documentation (depends on: 17)
Create role-specific education so engineers understand the new system and are confident executing during an incident.
Training tracks: (1) All engineers: 30-min video on severity levels, communication expectations, postmortem format, action items (mandatory); (2) On-call responders: 2-hour workshop on playbooks, alert tool, escalation paths, when to call manager, case studies of real incidents, mobile app walkthrough; (3) Incident commanders: 4-hour IC bootcamp covering leadership under pressure, decision-making, communicating with executives, status page discipline, postmortem facilitation, practiced drills; (4) Communications leads: templates, when to update, how to talk to customers, regulatory notification rules.
- Record videos so async teams can learn on their schedule
- Create runbooks and quick-reference cards for each role (print + digital)
- Pair new on-call engineers with experienced responder for first week
- Require IC certification before anyone joins IC rotation (pass a practical drill)
19. Execute staged rollout across teams (depends on: 18)
Progressively activate the incident management system with feedback loops at each stage; reduces risk of system-wide failure.
Wave 1 (week 6–8): 3–4 pilot teams begin on-call rotations and incident response using new system; capture feedback daily. Wave 2 (week 10–14): 8–10 additional teams, incorporating lessons from Wave 1; ensure diversity of team types (payment processing, monitoring, data pipeline, auth, etc.). Wave 3 (week 15–20): remaining teams; by now, the system is proven and less hand-holding needed.
- Daily retros with Wave 1 teams: what worked, what was confusing, what broke
- Each wave produces a "lessons learned" document that informs the next
- Track adoption metrics: how many incidents reported per team, alert quality, MTTD/MTTR
- Address resistance: engineers who are skeptical of the system, on-call burden, tool friction; assign a "change champion" in each team
20. Run incident response drills and simulations (depends on: 19)
Practice incidents in a controlled setting so responders gain confidence and gaps are found before a real crisis.
Drill 1 (week 8, pilot teams only): simulated database failover; test alert routing, IC assignment, on-call notification, status page update, postmortem. Drill 2 (week 12, all teams): full incident spanning multiple services; test escalation, communications, decision-making. Drill 3 (week 18): chaos drill with intentional confusion—what if IC is unreachable, what if comms lead is on PTO, what if an alert tool goes down? Quarterly thereafter: repeat drills with new on-call engineers.
- Schedule drills at a time when key people are available (not Friday afternoon)
- Run drill with live observers (reliability lead, manager) who provide feedback
- Measure drill success: all steps completed on time, communications clear, postmortem written
- Debrief immediately: what did we learn, what do we change?
21. Establish continuous improvement process (depends on: 15, 20)
Build a feedback loop so the incident management system evolves based on what you learn.
Monthly improvement cycle: in the monthly incident review (S15), leadership identifies one or two friction points from recent incidents (e.g., "alert tool response is slow," "postmortem meetings are boring," "action items keep slipping"). Assign an owner to investigate and propose a fix (change to playbook, tweak to tool config, additional training, process change). Test the fix with pilot team, gather feedback, roll out if successful. Annual review: can we reduce on-call burden, simplify communication, improve MTTD further? Do any teams need special handling (new services, difficult customers)?
- Keep a public backlog of improvement ideas (Jira, shared doc) so any engineer can suggest
- Quarterly all-hands: share wins ("MTTD improved 50%," "on-call satisfaction up," "zero customer-detected incidents") to build momentum
- Link improvements back to metrics (S14): show that changes actually worked
Previous Proposal 2 (ID: bdeac435-cbbd-4c22-bd5c-08a0ab80adec, Agent: deepseek-flash_initial_2, LLM: deepseek/deepseek-flash):
Estimated Complexity: high
Success Metrics: - Median time to detect reduced from 22 minutes to 5 minutes or less by month 6.
- Customer-detected incidents reduced from 40 percent to 10 percent or less of all customer-impacting incidents by month 9.
- Median time to mitigate reduced from 3 h 10 min to 60 minutes or less for SEV1 and SEV2 by month 9.
- Customer-impacting incidents reduced by at least 40 percent year over year from the 31 incident baseline.
- SLA credits paid reduced by at least 50 percent year over year from the $1.3M baseline.
- 100 percent of SEV1 incidents have a named Incident Commander within 5 minutes and a scribe within 15 minutes.
- Zero incidents in which command authority is unclear for more than 15 minutes.
- 28 of 28 teams have a documented on-call rotation, an owning service map and at least one trained on-call engineer by month 6.
- Central Incident Commander roster holds at least 40 certified ICs covering 24x7 with no single point of failure.
- 100 percent of on-call rotations are paid under a published policy by month 5.
- Monthly alert volume reduced from 3,400 to below 700, with a false-positive rate below 20 percent.
- No service exceeds 2 pages per on-call shift, measured monthly for three consecutive months.
- 100 percent of SEV1 and SEV2 postmortems published internally within 15 business days.
- At least 90 percent of postmortem action items closed within 60 days, up from 17 percent (11 of 64).
- Status-page first update published within 30 minutes on at least 95 percent of SEV1 incidents.
- Zero missed regulatory notification windows on any incident requiring notification.
- SOC 2 Type II audit passed with no findings related to incident response.
- Review cadence sustained: weekly operational review in at least 90 percent of weeks, 12 of 12 monthly reliability reviews, 4 of 4 quarterly executive reviews.
- On-call satisfaction at 70 percent or higher on the quarterly survey, with zero on-call-attributed voluntary attrition.
- 100 percent of new engineers complete the incident-response onboarding module within 30 days of joining.
Steps (21):
1. Programme charter, ownership and executive mandate
This step turns the CEO's email into a funded programme with a named owner and explicit authority. Without it, every downstream decision stalls in cross-team negotiation.
- Appoint a single accountable process owner (for example a Director of Incident Management) reporting to the CTO, with a dotted line to the COO for customer and SLA matters.
- Publish a one-page charter covering scope (all customer-impacting and money-moving incidents), decision rights, and the power to override team preferences during an active incident.
- Define the funding envelope: tooling licences, training time, exercise time and on-call compensation, with an indicative annual figure.
- Set the timeline against the SOC 2 date: a working process in four months, evidence accumulating from month two, audit-ready by month seven.
- Stand up a steering group with CTO, VP Engineering, Head of Support, Head of Compliance and one engineering manager per region.
- Agree that incident-process participation is a documented performance expectation for engineering managers, not an optional extra.
2. Baseline measurement and evidence pack (depends on: 1)
You cannot demonstrate improvement without a defensible starting point, and you cannot win the argument about noise without numbers.
- Build a 12-month incident register: date, retro-assigned severity, detection source, time to detect, time to mitigate, customer impact, services involved and SLA credits paid.
- Quantify the alert estate per tool, per team and per service; compute page-to-action ratio, list the 50 noisiest rules and count off-hours interruptions per engineer.
- Survey on-call engineers and managers on burden, fairness and clarity of escalation, targeting a response rate above 70 percent.
- Interview Support, Account Management and Sales about how customers learn of incidents and what they complain about.
- Document exactly where the current process breaks: unclear command in the two known incidents, postmortem action closure at 11 of 64, and ad-hoc status-page authorship.
- Publish the pack internally as the problem statement and retain it as management-review evidence for the SOC 2 audit.
3. Severity taxonomy and trigger matrix (depends on: 1, 2)
Severity is the keystone of the whole process. Every other rule, from paging to communications timing to postmortems, is keyed off it.
- Define four levels plus a special SEV0 for security or regulatory events: SEV1 for total or material loss of a payment path, SEV2 for degradation or single-region loss, SEV3 for limited impact with a workaround, SEV4 for internal-only issues and near-misses.
- Anchor each level in customer and money terms: customers affected, dollars at risk, regions lost, ledger integrity, settlement delay.
- Specify automatic triggers, for example loss of one AWS region, ledger write failures, a missed settlement cut-off, or payment success rate below threshold for five minutes.
- State who may declare each level (any engineer, Support or account manager may declare) and who may only recommend a downgrade (the Incident Commander alone).
- Map each level to SLA credit exposure and to the customer-visible status-page state.
- Include worked examples from the last 12 months so teams recognise their own incidents in the definitions.
- Add a review clause: the taxonomy is re-validated quarterly against real declarations.
4. Incident roles, command structure and decision rights (depends on: 2, 3)
The two incidents where nobody knew who was in charge for over an hour are the direct brief for this step.
- Define roles with one-page role cards: Incident Commander, Deputy IC, Operations Lead, Internal Communications Lead, Customer Communications Lead, Scribe, Subject-Matter Responders and an Executive Sponsor for SEV1 only.
- State the core rule plainly: the IC owns the incident, not the fix, and does not debug.
- Give the IC explicit decision rights: declaring and escalating severity, freezing changes, halting deploys, approving customer messaging and calling additional responders.
- Define minimum viable role coverage per severity: SEV1 staffs every role, SEV3 staffs an IC and a scribe only.
- Define handover discipline: maximum four-hour IC shifts on SEV1, a written handover template, and a Deputy IC nominated within 15 minutes.
- Set behaviour standards for responders: one incident channel, one bridge, no side channels, and explicit asks with a named owner and a time.
- Publish role cards on the internal wiki and link them from every paging notification.
5. Escalation, paging and incident lifecycle policy (depends on: 3, 4)
This step defines the mechanical path from an alert to a declared incident and back to normal service.
- Define lifecycle states with entry and exit criteria: Detected, Triaged, Declared, Mitigated, Resolved, Postmortem, Closed.
- Set acknowledgement targets: page acknowledged within five minutes, triage decision within fifteen, declaration within thirty for anything customer-facing.
- Define escalation ladders per layer (responder, service owner, team manager, IC on-call, VP Engineering) each with an automatic timer.
- Make escalation blameless and automatic: no responder is ever criticised for escalating, and timers fire whether or not a human asks.
- Define change freeze and rollback authority during SEV1 and SEV2, and the single condition that lifts the freeze.
- Enforce one incident, one record: the incident record is the sole source of truth for timeline, roles and communications.
- Require every SEV1 and SEV2 to produce an automatically captured timeline from channel and bridge, never one written from memory afterwards.
6. Detection strategy: SLOs, signals and customer-journey monitoring (depends on: 2, 3)
Customers detected 40 percent of incidents first. That number is the reason this step exists.
- Define SLIs and SLOs for the top 20 customer journeys, including payment initiation, settlement, ledger read and write, API availability and webhook delivery, measured per region.
- Require symptom-based alerting on those SLOs rather than cause-based alerting on infrastructure metrics.
- Add synthetic transaction monitoring from outside AWS in both regions and a third location, on a one-minute cadence for money-moving paths.
- Add ledger and shared-PostgreSQL specific signals: replication lag, connection saturation, write latency, transaction ID exhaustion and checkpoint pressure.
- Open a customer-reported path so Support and account managers can raise an incident directly, and count that path as a detection source in reporting.
- Publish a detection contract per service: every one of the 180 services needs an owner, at least one symptom alert and a documented expected detect time.
- Fund a separate resilience track to reduce shared-cluster blast radius, because better detection will not save a single shared ledger during a corruption event.
7. Alert quality standard and noise-reduction programme (depends on: 2, 3, 6)
3,400 alerts a month with 85 percent noise is the reason engineers resent the pager. Fixing it is the price of admission for everything else.
- Publish alert standards: every page must be symptom-based, actionable, owned, linked to a runbook and mapped to a severity. No runbook, no page.
- Separate paging from ticketing: only customer-impacting or imminent-impact signals may page; everything else becomes a ticket or a dashboard entry.
- Set a noise budget per team and per service, for example no service may exceed two pages per on-call shift, measured monthly.
- Define the default action for a noisy alert: fix, tune or delete within ten working days, with deletion treated as a legitimate and successful outcome.
- Introduce correlation and deduplication at the event pipeline so a single root cause produces one page instead of forty.
- Require expiry dates on every silencing rule and temporary threshold so suppression cannot become permanent blindness.
- Run a 90-day noise sprint with a visible burn-down of the top 100 noisiest rules, each owned by a named manager.
- Report page-to-action ratio per team in the monthly reliability review.
8. On-call architecture and 24x7 coverage model across 28 teams (depends on: 3, 4)
This is the hardest political step. The answer to carrying a pager for another team's code is that every team carries its own, and the platform carries the shared risk.
- Adopt a federated model: every service has exactly one owning team, and that team's primary on-call carries its own pager. No team is paged for code it does not own.
- State the consequence honestly: 16 of 28 teams currently have no on-call. They must build one or formally transfer ownership of their services to a team that will.
- Merge small or low-traffic teams into shared rotations so no rotation has fewer than six engineers, below which coverage is not sustainable.
- Create a separate, centrally staffed Incident Commander on-call roster drawn from trained senior engineers across all teams, covering 24x7.
- Define primary and secondary per rotation, with the secondary engaged only on a no-acknowledge or an explicit request.
- Define coverage across the two AWS regions and New York business hours: one global IC rotation, service on-call aligned to their service's users.
- Define the unresponsive-team path: fifteen minutes without acknowledgement escalates to the team manager, thirty minutes to the IC, who may then direct any available engineer.
- Publish a coverage matrix of all 28 teams showing services, rotation size and gaps, reviewed monthly.
- Make on-call participation an explicit expectation in engineering job levels and hiring criteria.
9. On-call compensation, wellbeing and sustainability policy (depends on: 8)
Unpaid on-call is the most cited reason for resistance. The policy must be settled before rollout, not negotiated during it.
- Move to paid on-call: a per-shift stipend or salary uplift agreed with HR and Finance and benchmarked to the New York market.
- Pay event-based compensation for incident callouts outside business hours, with a minimum call-out block.
- Provide compensatory rest: no engineer works a normal day after a night incident, and the rest day is documented, not granted as a favour.
- Cap intrusion by defining a maximum number of off-hours pages per shift, with a mandatory review triggered whenever it is exceeded.
- Define a voluntary opt-out path for engineers with genuine constraints, balanced by an explicit obligation that someone else is paid to take the shift.
- Include on-call expectation and compensation in offers and job descriptions so the commitment is set before hiring.
- Publish the policy with an effective date before any team is asked to join a new rotation.
- Review the policy every six months against actual page volumes, attrition and survey results.
10. Internal and customer communications policy with timing SLAs (depends on: 3, 4)
Today the status page is written by whoever is around. This step replaces improvisation with a clock and a named owner.
- Set internal cadence: first update within 15 minutes of declaration, then every 30 minutes for SEV1 and 60 minutes for SEV2, regardless of whether there is progress.
- Set customer cadence: status page updated within 30 minutes of a SEV1 and 60 minutes of a SEV2, and a no-new-information update is still mandatory.
- Define the channel hierarchy: status page for everyone, direct email to affected customers on SEV1, named account-manager calls for the top 50 accounts.
- Prepare templates per severity in advance with legal and compliance pre-approval, covering detection, impact, workaround, mitigation and next-update time.
- Define regulatory obligations explicitly: money transmitter and banking regulator notification windows, security breach notification, and who signs off (Compliance, not Engineering).
- Prohibit speculation: customer communications never guess at cause or blame and never commit to a root cause before the postmortem.
- Define the closing communication: resolution notice, expected SLA credit handling and the committed date for a written report.
- Assign a named Customer Communications Lead per incident with a trained deputy on every SEV1.
11. Status page, notification tooling and account-manager playbook (depends on: 10)
Policy without tooling collapses at three in the morning. This step makes publishing a five-minute action.
- Upgrade or replace the status page so components map to customer journeys rather than internal services, with subscriber control per component.
- Integrate the incident tool with the status page so the incident record drives the update and the public timeline.
- Provide one-click templates pre-filled with severity, impact language and next-update time.
- Give account managers a playbook: contact tree, what they may say, what they must not say, and how to escalate a customer question into the incident channel within minutes.
- Define the SLA credit process end to end, covering computation, approval, customer notification and finance treatment, so credits stop being a manual scramble.
- Host the status page outside the production failure domain so it survives a total platform outage.
- Test publishing during game days, including a simulated status-page outage and a simulated loss of the primary region.
12. Postmortem policy, template and blameless review process (depends on: 3, 4)
Only 11 of 64 action items closed means the postmortem ritual is currently a writing exercise. This step rebuilds it around learning and tracking.
- Make postmortems mandatory for every SEV1 and SEV2, every SEV3 with customer impact or a repeat pattern, and any near-miss the IC flags.
- Set deadlines: draft within five business days, blameless review within ten, internal publication within fifteen.
- Adopt a single template: impact, timeline, detection, response, contributing factors, what went well, what went badly, action items.
- Train a pool of blameless facilitators and require a trained facilitator for every SEV1 review.
- Prohibit counterfactual and blame language in the template, and require contributing factors across tooling, process, organisation and human factors.
- Limit action items to a small number of concrete, verifiable, owned changes with dates.
- Publish all postmortems internally by default, with security review only for genuinely sensitive material.
- Produce a customer-facing root cause report variant for SEV1 incidents, especially those affecting regulated or top-tier accounts.
13. Corrective action tracking and reliability backlog governance (depends on: 12)
A postmortem without durable action tracking is a complaint, not a control.
- Create a single reliability backlog in the engineering tracker with a mandatory label, owner, due date and link to the originating incident.
- Define closure criteria that require evidence: a merged change, a tested alert or a verified drill, never a self-reported status change.
- Protect capacity by reserving a fixed percentage of each team's sprint for reliability work, with unspent capacity visible to vice presidents.
- Run a weekly ageing review of open actions and escalate anything overdue by more than 30 days to the VP Engineering.
- Report closure rate and median age monthly, targeting more than 90 percent closed within 60 days.
- Require a repeat incident in the same area to trigger a design review rather than another action item.
14. Incident tooling consolidation and integration (depends on: 3, 5, 7, 11)
Six alerting tools and no single incident record is a structural cause of the 22-minute detection and the three-hour mitigation.
- Select one incident management platform for paging, on-call schedules, escalation policies, incident records and postmortem workflow.
- Consolidate the six alerting sources into a single event pipeline feeding that platform, with deduplication and severity mapping applied at ingest.
- Integrate with platform and ledger observability so responders see dashboards and runbooks inside the incident record.
- Integrate chat and bridge: incident channel auto-created, timeline auto-captured, decisions logged as they happen.
- Define the data model and retention required for SOC 2 evidence: who did what, when, and under whose authority.
- Run a dual-run period alongside the old tools with a defined rollback, then switch off the legacy tools on a published date.
- Budget for licences, migration effort and a two-week hardening period after cutover.
15. Training, certification and exercise programme (depends on: 4, 5, 10, 12)
A process that exists only on a wiki page fails on the first real page. Skills have to be built and tested before they are needed.
- Build a short practical curriculum: how to be on call, how to declare an incident, how to run an incident as IC, how to communicate and how to write a postmortem.
- Require certification before joining the IC on-call roster: a written assessment plus a live simulated incident.
- Train at least two certified ICs per team group so the central roster has depth across all 28 teams.
- Run monthly tabletops on realistic scenarios drawn from the last 12 months, including region loss and ledger corruption.
- Run quarterly game days with deliberately injected failure, including shared-PostgreSQL failover and status-page outage.
- Add a mandatory onboarding module for every engineer joining or transferring in, with completion records suitable for audit.
- Track training completion by team and publish it in the monthly reliability review.
16. Metrics, dashboards and review cadence (depends on: 2, 3)
The programme needs a public scoreboard, or it will quietly rot after the audit.
- Define the outcome metrics: time to detect by source, time to mitigate, percentage of incidents detected by customers (target below ten), incidents by severity and SLA credits paid.
- Define the process metrics: declaration latency, page acknowledgement rate, IC roster coverage, first-update timeliness and update-cadence adherence.
- Define the health metrics: alert volume and noise ratio per team, off-hours pages per engineer, postmortem timeliness, action closure rate and action age.
- Publish live dashboards visible to every engineer, not only to managers, refreshed daily.
- Institute a weekly operational review of 30 minutes going incident by incident, a monthly reliability review of trends and actions, and a quarterly executive review with the CEO's office.
- Baseline every metric against the S2 evidence pack and set 90-day and 12-month targets.
- Require every review to end with decisions and owners, not just numbers.
17. Pilot with volunteer teams (depends on: 5, 7, 9, 11, 12, 13, 14, 15, 16)
Do not roll out to 28 teams untested. Run the entire process end to end with a small cohort first.
- Recruit three to four volunteer teams covering a mix of criticality: one payment-path team, one ledger-adjacent team, one platform team and one low-traffic team.
- Run the complete process in the pilot: new severity scale, roles, escalation, communications, postmortems, action tracking and paid on-call.
- Instrument the pilot against the S16 metrics and compare results with the S2 baseline.
- Hold weekly retrospectives with pilot teams and iterate on the written policies, the tooling and the training.
- Fix the top issues found before any wider rollout and document what changed and why.
- Produce a pilot report with before-and-after numbers to carry into every rollout conversation.
- Set explicit pilot exit criteria: rotation coverage achieved, no unacknowledged pages over a defined period, postmortems delivered on time and actions tracked.
18. Phased rollout to all 28 teams (depends on: 13, 16, 17)
Rollout is a staged migration with readiness gates, not an email announcement.
- Sequence the 28 teams into four waves of roughly seven teams, ordered by customer impact, with three weeks between waves.
- Define a per-team readiness checklist: services mapped and owned, alerts cleaned to standard, runbooks written, rotation staffed, training complete and manager briefed.
- Hold a gate review with the process owner before each team joins, and move unready teams to the next wave with a dated remediation plan.
- Give each wave a named champion and run an internal communications cadence that explains the why using pilot numbers.
- Handle resistance directly by publishing the own-your-code-own-your-pager rule and the paid on-call mechanics before each wave, not after.
- Retire legacy tools, escalation lists and the informal status-page process at the end of each wave on a published cutover date.
- Harvest feedback formally at each wave and push accepted changes back into the policy documents through change control.
19. SOC 2 incident-response control mapping and evidence framework (depends on: 1, 3, 10, 12)
The audit will test incident response. Mapping controls now means evidence accumulates from the first real incident under the new process.
- Map the process to the relevant Trust Services Criteria for incident identification, response, evaluation of incidents and communication of security events.
- Write control statements in auditor language and name a single owner for each control.
- Specify the evidence artifact for each control: incident record, severity classification, paging log, communications log, postmortem, action tracker entry and training record.
- Set evidence retention and storage location so nothing depends on a laptop or on chat history that expires.
- Run an early walkthrough with an experienced compliance partner or the auditor's readiness team to test the design before the audit window.
- Flag controls that are blocked until rollout reaches the relevant teams, and keep a gap register with owners and dates reviewed fortnightly by the steering group.
20. SOC 2 dry run, gap remediation and audit support (depends on: 18, 19)
This step converts a good process into a provable one, a few months before the auditors arrive.
- Run an internal dry run six weeks before the audit window, sampling real incidents from the pilot and early waves against each control's evidence requirements.
- Remediate every gap found, prioritising anything that risks a qualified opinion.
- Brief every engineer who may be interviewed so they can describe the process as they actually practise it.
- Prepare the auditor package: process documentation, sample incident records, training records, on-call schedules and the action tracking register.
- Designate a single audit liaison and a small evidence-request team so requests do not land on 28 teams at once.
- Rehearse the walkthrough with the incident commander and communications roles, since auditors probe realism under pressure.
- Schedule the dry run and the audit alongside rollout waves so no team is hit with both in the same week.
21. Standing governance, process ownership and continuous improvement (depends on: 18, 20)
The classic post-audit failure is that the process freezes and then decays. This step keeps it alive.
- Establish a standing Incident Management Council chaired by the process owner, meeting monthly with engineering, support, compliance and product representation.
- Give the process owner a documented mandate to change standards, and require formal change control for any change to severity, roles, communications timings or compensation.
- Re-validate the severity taxonomy quarterly against real declarations and re-baseline the metric set annually.
- Fund a resilience roadmap separate from incident response: reducing shared-database blast radius, multi-region failover readiness and deploy safety.
- Report quarterly to the executive team on the metric set and on the top five systemic causes of incidents.
- Feed incident themes into architecture review, release engineering and hiring so the programme improves the system, not just the response.
- Refresh the training and tabletop programme annually and after any SEV1.
Previous Proposal 3 (ID: bba03d5c-7ae7-4435-9aef-990402088c21, Agent: qwen3.8-flash_initial_3, LLM: alibaba/qwen3.8-flash):
Estimated Complexity: high
Success Metrics: - MTTD reduced to under 10 minutes with >90% detection occurring before customer reporting.
- MTTR for Sev1/Sev2 incidents reduced to under 60 minutes.
- SLA credits paid reduced by 50% within 12 months.
- On-call alert volume reduced by 80% (from 3,400 to <600 actionable alerts/month).
- 100% of Sev1/Sev2 incidents have a blameless postmortem completed within 5 days.
- >90% completion rate of postmortem action items within 30 days of assignment.
- Zero ambiguity in incident command: 100% of major incidents have a designated IC logged in the timeline.
- SOC 2 Type II audit passes incident response controls with zero critical findings.
Steps (12):
1. Establish Governance and Executive Sponsorship
Secure explicit CEO and CTO backing to mandate the program, framing it as a revenue protection and compliance necessity rather than an engineering burden. Form a cross-functional Incident Management Task Force.
- Identify a single executive owner (e.g., VP of Engineering or Head of Platform) with authority over budget and staffing.
- Create a charter defining the scope: severity definitions, on-call compensation, and tool consolidation.
- Communicate the strategic priority to all 28 team leads to preempt pushback regarding pager responsibility.
2. Define Severity Matrix and Response SLAs (depends on: 1)
Create a rigorous, binary severity scale to eliminate ambiguity in escalation and communication speed. Align internal response times with the 99.95% SLA.
- **Sev1 (Critical)**: System-wide outage, data corruption, or >5% transaction failure rate. Immediate executive page.
- **Sev2 (Major)**: Single customer segment impacted or >1% transaction failure rate. High-priority team response.
- **Sev3 (Minor)**: Degraded performance or isolated bug. Standard business-hours follow-up.
- **Sev4 (Cosmetic)**: No functional impact. Backlog item.
- Define maximum detection and mitigation times for each level.
3. Restructure On-Call Organization and Compensation (depends on: 1)
Transform the current unpaid, fragmented on-call system into a professional, compensated, and centralized rotation to ensure 24x7 coverage across all 28 teams. Address the 'pushback' by making it mandatory and paid.
- Consolidate the 12 current on-call teams into a unified, rotating pool covering all 28 teams.
- Implement a fair compensation model: base stipend + hourly rate for active hours + incident bonus for Sev1/2.
- Define a strict 'Buddy System' or 'Follow-the-Sun' rotation if geographically feasible, otherwise ensure adequate team size to rotate fairly.
- Mandate that every team must participate in the rotation; no exemptions.
4. Define Incident Roles and RACI Model (depends on: 2)
Eliminate the 'who is in charge' ambiguity by pre-defining roles and establishing a clear chain of command for every incident. Roles are assigned per incident, not permanently held.
- **Incident Commander (IC)**: Single point of authority. Owns the timeline, delegates tasks, and approves communications. Must be senior.
- **Scribe**: Maintains the real-time timeline log of actions, timestamps, and decisions.
- **Communications Lead**: Owns external/status page updates and internal Slack broadcasts. Shields the IC from interruptions.
- **Technical Leads (SMEs)**: Deployed by IC to troubleshoot specific services. They do *not* coordinate with each other; they report to IC.
- Create a 'Rapid Response Team' (RRT) for high-severity incidents involving the shared PostgreSQL cluster.
5. Consolidate Monitoring and Enforce Alert Hygiene (depends on: 3)
Reduce the 3,400 monthly alerts (85% noise) to a signal-to-noise ratio that allows engineers to trust the pager. Standardize tooling.
- Consolidate the six disparate alerting tools into a single platform (e.g., PagerDuty, Opsgenie, or Datadog).
- Implement 'Actionable Alerting': If an alert pages, there must be a documented, immediate runbook action. No 'informational' pages.
- Set a hard cap on alert volume per engineer/night. Breaching the cap triggers a mandatory alert quality review.
- Deprecate legacy tools that cannot map to the severity matrix defined in S2.
6. Build Detection and Escalation Automation (depends on: 5)
Reduce the 22-minute mean time to detect (MTTD) by automating the path from metric breach to page. Ensure customers never find an issue before the internal pager does.
- Implement synthetic transactions for core payment flows to detect failures before real users do.
- Configure escalation policies: If not acknowledged in 5 minutes, auto-escalate to the team lead, then to the Incident Commander pool.
- Map specific Kubernetes and PostgreSQL metrics to the Severity Matrix (S2) to trigger automated pages.
- Create a 'Major Incident Channel' in Slack that is automatically opened and pinned when a Sev1/2 is triggered.
7. Standardize Communication Templates and Timing (depends on: 4, 6)
Codify the rules for external and internal communications to stop the 'whoever is around' approach. Link these timings to the Severity Matrix.
- **Status Page Updates**:
- Sev1: First update within 5 minutes of detection. Updates every 15 minutes.
- Sev2: First update within 30 minutes. Updates every 60 minutes.
- **Customer Direct Mail**: For Sev1 impacting specific accounts, Account Managers send pre-approved templates within 1 hour of mitigation.
- **Regulator Notification**: Define legal triggers for NY financial regulations.
- Create a 'Comms War Room' Slack channel restricted to the IC and Comms Lead.
8. Formalize the Postmortem Process (depends on: 4, 7)
Transform postmortems from optional/various formats into a mandatory, blameless, and tracked process. This directly addresses the SOC 2 audit requirement for continuous improvement.
- Mandate a postmortem for every Sev1 and Sev2 incident within 5 business days.
- Enforce a 'Blameless' charter: Focus on systemic/process failures, not individual errors.
- Adopt a standardized template: Timeline, Root Cause Analysis (5 Whys or Fishbone), Impact Analysis, and Action Items.
- Action Items must have: Owner, Due Date, and Priority. They are tracked in Jira/Linear, not the doc.
9. Implement Action Item Tracking and Governance (depends on: 8)
Solve the '11 of 64 closed' problem by integrating postmortem actions into the core engineering workflow and executive reviews. Ensure fixes are shipped.
- Integrate Jira with the Postmortem tool (e.g., OpsLevel, Rootly, or Notion API). Auto-create tickets for every AI.
- Block 'Sev1 Closure' until all high-priority Action Items are closed or explicitly deferred with executive approval.
- Weekly Engineering Leadership review of open postmortem actions older than 30 days.
- Tie action item completion rates to team performance reviews.
10. Deploy Training and Runbook Library (depends on: 6, 7, 8)
Address the skill gap and the 'pushback' by equipping engineers with the tools and knowledge to succeed on-call. Reduce anxiety through clarity.
- Audit all 180 services. Require every service to have a 'Page-Worthy' runbook.
- Run 'Game Day' (Chaos Engineering) exercises quarterly to practice role assignments (IC, Scribe, Comms) in a simulated Sev1.
- Conduct mandatory training for the new on-call rotation and compensation model.
- Provide a 'Playbook' for the Incident Commander role, including scripts for communication.
11. Phased Rollout and Change Management (depends on: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10)
Introduce the new process in phases to manage cultural resistance and technical debt, ensuring adoption before the SOC 2 audit. Do not 'big bang' the change.
- **Month 1-2**: Launch Severity Matrix and Consolidated Tooling (S2, S5). Pilot on 5 high-traffic teams.
- **Month 3-4**: Roll out new On-Call Compensation and Roles (S3, S4) to all 28 teams. Enforce new Communication Templates.
- **Month 5-6**: Enforce strict Postmortem and Action Item Tracking (S8, S9).
- **Month 7**: Final audit readiness check. Simulate SOC 2 auditor questions on incident response.
12. Define Metrics for Success and Continuous Review (depends on: 2, 6, 8, 9)
Establish a dashboard to prove the new system works, linking directly to the $1.3M savings in SLA credits. Monitor the health of the incident process itself.
- Track MTTD (Target: <5 min), MTTR (Target: <30 min for Sev2), and Detection Ratio (Target: >95% internal detection).
- Monitor 'Alert Noise' (target: <5% of alerts result in a non-actionable page).
- Track 'Action Item Completion Rate' (Target: 100% on-time closure for Sev1/2).
- Calculate 'SLA Credit Avoidance' vs. On-Call Program Cost to prove ROI to leadership.
Please, considering the previous proposals as ideas that could be considered, focus on the main objective and generate an IMPROVED proposal or a completely DIFFERENT perspective if you deem it appropriate. Only if you consider any of them is amazing and impossible to improve, answer with the same proposal.
The plan has these parts:
- "steps": the list of steps, each with "step_id" (a short label unique within this proposal, such as S1, S2, S3), "title" (a short, descriptive title), "description" (what the step does and how, in Markdown: a short opening paragraph and then, if there are several concrete things to say, a bullet list) and "dependencies" (the step_ids of the other steps of THIS proposal that must be completed first; empty if none).
- "estimated_complexity": "low", "medium" or "high".
- "success_metrics": clear and measurable success metrics, one per line as a bullet list.
Think deeply about the problem, considering all aspects and implications. Develop a comprehensive understanding of what needs to be done to achieve the objective, then produce the plan.
{"steps": [{"step_id": "S1", "title": "Executive mandate and governance structure", "description": "Turn the CEO email into a funded, authorized program with clear ownership and decision rights. Without executive backing, every downstream decision stalls in negotiation.\n\n- Appoint a single program owner (e.g., Director of Incident Management) reporting to the CTO and COO.\n- Publish a one-page charter covering scope (all customer-impacting incidents), authority to override team preferences during incidents, and funding for tooling, training, and on-call compensation.\n- Establish a standing Incident Management Steering Group with CTO, VP Engineering, VP Support, Head of Compliance, and one engineering manager per region meeting monthly.\n- Secure budget envelope: tool licenses, training time, incident-response infrastructure, and on-call compensation (estimated $400–600K annually).\n- Make incident-response participation a documented performance expectation for all engineering managers and team leads.", "dependencies": []}, {"step_id": "S2", "title": "Baseline measurement and evidence pack", "description": "You cannot prove improvement without defensible baseline numbers, and you cannot win arguments about noise or impact without data.\n\n- Build a 12-month incident register: date, detection source, impact scope, time to detect, time to mitigate, SLA credits paid, and services involved.\n- Audit the current alert estate: count alerts per tool, per team, per service; compute page-to-action ratio; identify top 50 noisiest rules and off-hours interruptions per engineer.\n- Survey on-call engineers and managers on burden, fairness, and escalation clarity (target >70% response rate).\n- Interview Support and Account Management: how do customers discover incidents, what do they complain about?\n- Document the two command-ambiguity incidents: exactly when unclear who was in charge and why, how long it lasted.\n- Publish this pack internally as the problem statement and retain all artifacts for SOC 2 audit evidence.", "dependencies": ["S1"]}, {"step_id": "S3", "title": "Severity taxonomy and trigger matrix", "description": "Severity is the keystone. Every other rule—paging, communications, postmortems, compensation—keys off it. Define four levels plus a special SEV0 for security/regulatory events.\n\n- **SEV1 (Critical)**: Complete service outage, data corruption, or >5% payment-path failure rate for >5 min. Every minute costs money. IC required; 99.99% uptime threatened.\n- **SEV2 (Major)**: Significant degradation, single region loss, or 1–5% transaction failure. IC typically required; service credit exposure.\n- **SEV3 (Minor)**: Limited customer impact with workaround available, or internal issues affecting operations. On-call SME + escalation if SLA at risk.\n- **SEV4 (Cosmetic)**: Observations, non-impacting bugs, alerts. Alert-driven, no escalation unless pattern emerges.\n- Specify automatic triggers: region loss, ledger write failures, missed settlement window, payment success rate thresholds.\n- Define who may declare (any engineer, Support, account manager) and who may downgrade (IC only).\n- Include worked examples from the last 12 months so teams recognize their incidents in the definitions.\n- Review and re-validate quarterly against actual declarations.", "dependencies": ["S2"]}, {"step_id": "S4", "title": "Incident roles, command structure, and decision rights", "description": "The two incidents with >1 hour of command ambiguity prove this step is non-negotiable. Define clear roles with explicit decision authority.\n\n- **Incident Commander**: Owns the incident timeline, not the fix. Declares severity, decides escalation, approves communications, freezes changes, calls responders. Non-technical ICs are acceptable.\n- **Deputy IC**: Shadows IC; takes over if IC unavailable. Nominated within 5 minutes of incident declaration.\n- **Communications Lead**: Owns internal Slack updates and status-page messaging. Shields IC from interruptions.\n- **Scribe**: Records real-time timeline: decisions, who did what, key timestamps. Not responsible for fixing.\n- **Subject-Matter Responders**: Engineers with service context. Take IC direction without debate. Report only to IC.\n- **Operations Lead** (SEV1 only): Coordinates across multiple responders, manages incident bridge.\n- Write one-page role cards with mission, decision authority, and escalation upward. Publish on wiki and link from every paging notification.\n- Define minimum viable coverage per severity: SEV1 staffs all roles; SEV2 staffs IC, Comms, Scribe; SEV3 staffs IC + Scribe.\n- Establish handover discipline: maximum 4-hour IC shifts on SEV1, written handover template required.", "dependencies": ["S3"]}, {"step_id": "S5", "title": "Incident tooling consolidation and integration", "description": "Six alert tools and ad-hoc incident records are structural causes of the 22-minute detection and 3+ hour mitigation. Consolidate to a single incident platform that is the source of truth.\n\n- Select an incident management platform (PagerDuty, Opsgenie, Incident.io, etc.) that supports paging, schedules, escalation, incident records, and postmortem workflow.\n- Requirement: the platform must integrate with observability tools, auto-create and pin incident channels in Slack, auto-capture timeline from chat, and support API-driven playbook automation.\n- Plan a dual-run period alongside legacy tools with a published cutover date; define rollback criteria.\n- Integrate incident record with the 180 services' monitoring and dashboards so responders see everything in one place.\n- Define data retention and audit trail to satisfy SOC 2 evidence requirements: who did what, when, under whose authority.\n- Budget for licenses, migration effort, and two-week hardening period post-cutover.", "dependencies": ["S1", "S3"]}, {"step_id": "S6", "title": "Alert consolidation and event pipeline", "description": "Replace six alert sources with a single ingestion point. Deduplicate and route alerts with minimal manual judgment, removing a major source of detection delay.\n\n- Consolidate alert endpoints from six tools into a single event pipeline; this often sits in front of the incident platform (S5).\n- Implement deduplication and correlation so a single outage triggering alerts from five monitoring tools produces one page, not five.\n- Map every alert to a severity level from S3 (SEV1, SEV2, SEV3, SEV4) at ingestion.\n- Log every alert for postmortem analysis and trending.\n- Ensure the platform's mobile app works reliably; on-call responders need to engage from any device.", "dependencies": ["S5"]}, {"step_id": "S7", "title": "Detection strategy: SLOs, signals, and customer-journey monitoring", "description": "Customers detected 40% of incidents first—a detection gap that must be closed. Build symptom-based alerting that detects outages before customers do.\n\n- Define SLIs and SLOs for the top 20 customer journeys: payment initiation, settlement, ledger read/write, API availability, webhook delivery, measured per region.\n- Require symptom-based alerting on SLOs, not cause-based infrastructure metrics (e.g., alert on \"payment success rate <99%\" not \"database CPU >80%\").\n- Add synthetic transaction monitoring from outside AWS in both regions and a third location, one-minute cadence, for all money-moving paths.\n- Add ledger-critical signals: PostgreSQL replication lag, connection saturation, write latency, transaction ID exhaustion, checkpoint pressure.\n- Create a detection contract per service: owner identified, at least one symptom alert defined, expected detect time documented.\n- Open a customer-reported incident path: Support and account managers can declare an incident directly, counted as a detection source in metrics.\n- Fund a separate resilience roadmap to reduce shared-database blast radius, because detection improvements do not protect against ledger corruption.", "dependencies": ["S3", "S6"]}, {"step_id": "S8", "title": "Alert quality standards and noise-reduction program", "description": "3,400 monthly alerts with 85% noise is the reason engineers resent the pager. Cutting noise is the price of admission for on-call buy-in.\n\n- Publish alert standards: every page must be symptom-based, have an immediate runbook action, be owned by a team, and map to a severity level. No runbook, no page.\n- Separate paging from ticketing: only customer-impacting or imminent-impact signals trigger pages. Everything else becomes a dashboard entry or a ticket.\n- Set a noise budget per team and service: no service may exceed two pages per on-call shift per month. Breaching triggers a mandatory alert quality review.\n- Define the default action for a noisy alert: fix the root cause, tune the threshold, or delete it within ten working days. Deletion is a legitimate successful outcome.\n- Implement automatic suppression rules: silence alerts if service auto-recovered within 30 seconds; suppress known maintenance windows; group flapping alerts (>5 in 2 min) into one page; rate-limit noisy services (max 1 alert per 5 min until condition clears).\n- Require expiry dates on all silencing rules so suppression cannot become permanent blindness.\n- Run a 90-day noise sprint with a visible burn-down of the top 100 noisiest rules, each owned by a named manager.", "dependencies": ["S3", "S6", "S7"]}, {"step_id": "S9", "title": "Escalation policies and incident lifecycle", "description": "Define the mechanical path from alert to incident declaration to resolution. Escalation must be automatic and blameless.\n\n- Define incident lifecycle states with clear entry/exit criteria: Detected → Triaged → Declared → Mitigated → Resolved → Postmortem → Closed.\n- Set acknowledgement targets: page acknowledged in 5 minutes; triage decision (is this real?) in 15 minutes; severity declaration (is this customer-impacting?) in 30 minutes.\n- Build escalation ladders: if responder does not acknowledge in 5 min, page escalates to service owner, then team manager, then IC on-call. Escalation is automatic, not manual.\n- Implement escalation for SEV1: IC paged via phone call + SMS + Slack + mobile; if not acknowledged in 2 min, Deputy IC paged simultaneously; Communications Lead pinged at same time.\n- Define change freeze during SEV1 and SEV2: no deployments except to fix the incident. Freeze lifts only when mitigation is confirmed.\n- Enforce one incident, one record: the incident record is the sole source of truth. Auto-capture timeline from Slack and bridge; never write timeline from memory later.\n- Test all escalation paths weekly via synthetic page to on-call; adjust timings based on first month of operations.", "dependencies": ["S3", "S4", "S5", "S6"]}, {"step_id": "S10", "title": "On-call architecture and 24x7 coverage model", "description": "The answer to \"carrying a pager for another team's code\" is that every team carries its own, and the platform carries shared risk. Design a sustainable model.\n\n- Adopt a federated model: every service has one owning team; that team's on-call carries its service's pager. No team is paged for code it does not own.\n- State the consequence clearly: 16 of 28 teams currently have no on-call. They must either build one or formally transfer service ownership to a team that will.\n- Create a separate, centrally staffed Incident Commander on-call roster drawn from trained senior engineers across all teams, covering 24x7 with no single point of failure.\n- Merge small or low-traffic teams into shared rotations so no rotation has fewer than six engineers (below which coverage is unsustainable).\n- Define primary and secondary per slot: secondary engages only on no-acknowledge or explicit request from IC.\n- Align coverage across two AWS regions and New York business hours: one global IC rotation; service on-call aligned to service users' time zones.\n- Define unresponsive-team escalation: 15 min without acknowledgement escalates to team manager; 30 min escalates to IC, who may direct any available engineer.\n- Publish a coverage matrix of all 28 teams showing services, rotation size, and gaps, reviewed monthly.", "dependencies": ["S3", "S4", "S9"]}, {"step_id": "S11", "title": "On-call compensation, wellbeing, and sustainability policy", "description": "Unpaid on-call is the most cited reason for resistance. Settle compensation before rollout, not during it. Make it sustainable.\n\n- Introduce paid on-call: a weekly stipend ($500–1,000) while on-call, regardless of incident volume, benchmarked to New York market rates.\n- Pay event-based compensation: 1.5× hourly rate for time spent mitigating incidents during off-hours (minimum one-hour block per callout).\n- Provide compensatory rest: no engineer works a normal business day after a night incident requiring >2 hours mitigation. Rest day is documented, not granted as a favor.\n- Cap intrusion: define maximum off-hours pages per week (e.g., no more than three per shift). Mandatory review and escalation if exceeded.\n- Offer a voluntary opt-out path for engineers with genuine constraints, balanced by explicit obligation that someone else is paid to cover.\n- Include on-call expectation and compensation in job descriptions and hiring conversations so commitment is understood before joining.\n- Publish the policy with an effective date before any team is asked to join a rotation.\n- Review the policy every six months against actual page volumes, attrition rates, and survey feedback.", "dependencies": ["S10"]}, {"step_id": "S12", "title": "Playbooks and communication templates by severity", "description": "Playbooks remove ambiguity and decision fatigue during incidents. Templates ensure consistent, compliant messaging.\n\n- Create a one-page (or one-screen) playbook for each severity level: who gets paged (roles, order); first questions (is it real, how big, who knows); what IC declares first (status page text, account manager notification, regulatory trigger); escalation timeline.\n- **SEV1 playbook**: Immediate IC + Comms + CTO notification; customer status every 5 minutes; sample message templates.\n- **SEV2 playbook**: IC + Comms + tech lead notification; status every 15 minutes; decision tree for escalation to executive team.\n- **SEV3 playbook**: On-call SME + Comms if customer-visible; status every 30 minutes or as resolved.\n- **SEV4 playbook**: On-call SME only; update customers only if promised SLA is at risk.\n- Include decision trees: is this SEV1 or SEV2? Is it our code or dependency? Escalate or containment?\n- Prepare customer-communication templates pre-approved by legal and compliance: sample language for detection, impact, workaround, mitigation phases.\n- Publish playbooks on wiki and embed links in incident management platform.", "dependencies": ["S3", "S4"]}, {"step_id": "S13", "title": "Internal and customer communications workflows", "description": "Specify who informs whom, in what order, via what channel. Prevent gaps like \"nobody knew who was in charge for an hour.\"\n\n- **Internal cadence**: First update to #incidents Slack channel within 3 minutes of declaration (even if \"investigating\"). Updates every 5 minutes (SEV1), 15 minutes (SEV2), 30 minutes (SEV3) or when material change occurs.\n- IC calls CTO/VP Eng and incident channel lead within 1 minute of declaration (SEV1); incident declared in Slack with severity, IC name, and service affected.\n- SME on-call for the failing service joins incident bridge automatically; escalation call includes them within 5 minutes.\n- Designate a single Customer Communications Lead per incident (pre-identified on-call roster) who owns external messaging exclusively. Shields IC from customer contact.\n- **Customer communication**: Status page updated within 3 minutes (SEV1) or 10 minutes (SEV2) even if root cause unknown; post \"Investigating\" with next-update ETA. Account managers of affected top-tier customers called within 5 minutes (SEV1) with templated language.\n- Create a phone tree or escalation list accessible to responders; set expectation: \"If you don't hear from IC in 2 minutes, call them.\"\n- Use a single incident Slack channel per incident (auto-created by incident tool); log all communications for postmortem review.\n- Define regulatory notification path: compliance must approve before sending, but do not wait for root cause; flag incidents triggering payment-processing regulations to legal immediately.", "dependencies": ["S4", "S9", "S12"]}, {"step_id": "S14", "title": "Status page, customer notifications, and account-manager playbook", "description": "Policy without tooling collapses at 3 AM. Make publishing a five-minute action.\n\n- Upgrade or replace status page so components map to customer journeys (\"payments\", \"settlements\", \"ledger\") not internal services. Allow customers to subscribe per component.\n- Integrate incident tool (S5) with status page so incident record drives updates and public timeline auto-populates.\n- Provide one-click templates pre-filled with severity, impact language, and next-update time; reduce typing and errors.\n- Create account-manager playbook: contact tree for top 50 customers, what they may say (facts only), what they must not say (speculation, blame, false ETAs), escalation path if customer escalates.\n- Define SLA credit process end to end: impact detection → credit calculation (based on duration × severity) → approval → customer notification → finance treatment. Automate where possible.\n- Host status page outside production failure domain on separate infrastructure so it survives total platform outage.\n- Test status-page reliability during game days (S17), including simulated status-page outage and total region loss.", "dependencies": ["S5", "S12", "S13"]}, {"step_id": "S15", "title": "Postmortem policy, blameless process, and facilitation", "description": "Only 11 of 64 action items closed means postmortems are currently a writing exercise. Rebuild around learning and tracking.\n\n- Make postmortems mandatory: all SEV1 and SEV2, all SEV3 with customer impact or repeat pattern, any near-miss the IC flags.\n- Set deadlines: draft within 5 business days, blameless review within 10 days, internal publication within 15 days.\n- Adopt a single standardized template: impact and duration, timeline (detection through resolution), root cause (not \"human error\" but system condition that enabled error), contributing factors, what went well, what went badly, action items.\n- Train a pool of blameless facilitators; require a trained facilitator (not the IC) for every SEV1 and SEV2 review.\n- Prohibit counterfactual and blame language in postmortems; require contributing factors addressing tooling, process, organization, and human factors.\n- Limit action items to small number of concrete, verifiable, owned changes with dates.\n- Publish all postmortems internally by default, with security review only for genuinely sensitive material (e.g., unpatched vulnerability details).\n- For SEV1 incidents affecting regulated customers, produce a variant customer-facing root cause report.", "dependencies": ["S3", "S4"]}, {"step_id": "S16", "title": "Action item tracking, reliability backlog, and completion governance", "description": "A postmortem without durable action tracking is a complaint. Solve the 11-of-64 problem.\n\n- Create a single reliability backlog in the engineering tracker (Jira, Linear, etc.) with mandatory label, owner, due date, and link to originating incident.\n- Define closure criteria: evidence required (merged code change, tested alert, verified drill) not self-reported status.\n- Protect capacity: reserve a fixed percentage of each team's sprint (e.g., 10%) for reliability work; track unspent capacity and report to VP Engineering.\n- Run a weekly ageing review of open actions; escalate anything overdue by >30 days to team lead and VP Engineering.\n- Require IC or postmortem facilitator to sign off on action completion.\n- Report completion rate and median action age in monthly incident review (target: >90% closed within 60 days).\n- If the same service repeats an incident in the same area, trigger a design review rather than another action item; break the cycle.", "dependencies": ["S15"]}, {"step_id": "S17", "title": "Metrics, dashboards, and review cadence", "description": "Measure to prove the system works. Publish dashboards so everyone sees the scoreboard.\n\n- Define outcome metrics: time to detect (by source, target <5 min internally detected); time to mitigate SEV1/SEV2 (target <60 min); customer-detected incidents per month (target <2); SLA credits paid (target <$100K/year by month 12).\n- Define process metrics: declaration latency, page acknowledgement rate, IC roster coverage (no single point of failure), first-update timeliness (% within SLA), update-cadence adherence.\n- Define health metrics: alert volume and noise ratio per team, off-hours pages per engineer per month, postmortem timeliness, action closure rate and median action age.\n- Build live dashboards visible to every engineer (not just managers); auto-populate from incident tool, update daily.\n- Institute review cadence: weekly operational review (30 min, prior week incidents), monthly reliability review (trends, top causes, action status), quarterly executive review (CEO's office, SLA cost, systemic changes).\n- Baseline every metric against S2 evidence pack; set 90-day and 12-month targets.\n- End every review with decisions and owners, not just numbers.", "dependencies": ["S2", "S3"]}, {"step_id": "S18", "title": "Training, certification, and exercise program", "description": "A process that exists only on a wiki fails on the first real page. Build skills before deployment.\n\n- Build curriculum: how to be on-call, how to declare an incident, how to run incidents as IC, how to communicate, how to write blameless postmortems.\n- Create role-specific tracks: (1) All engineers: 30-min video on severity levels, communication expectations, postmortem format, action items (mandatory); (2) On-call responders: 2-hour workshop on playbooks (S12), alert tool (S5), escalation paths (S9), when to call manager, case studies, mobile app walkthrough; (3) Incident commanders: 4-hour IC bootcamp on leadership under pressure, decision-making, communicating with executives, status-page discipline, postmortem facilitation, practiced drills; (4) Communications leads: 2-hour training on templates, update timings, how to talk to customers, regulatory rules.\n- Require certification before joining IC on-call roster: written assessment plus live simulated incident (pass/fail).\n- Record videos so async teams can learn on their schedule; create runbooks and quick-reference cards (print + digital); pair new on-call engineers with experienced responder for first week.\n- Run monthly tabletop exercises on realistic scenarios from the prior 12 months: region loss, ledger corruption, cascading failures.\n- Run quarterly game days with intentional failure injection (database failover, status-page outage, alert tool downtime); include all on-call roles.", "dependencies": ["S4", "S12", "S13", "S14", "S15"]}, {"step_id": "S19", "title": "Pilot with volunteer teams", "description": "Do not roll out untested to 28 teams. Run the entire process end to end with a small cohort first.\n\n- Recruit 3–4 volunteer teams covering a mix of criticality: one payment-path team, one ledger-adjacent team, one platform team, one low-traffic team.\n- Run complete process in pilot: new severity scale (S3), consolidated tooling (S5, S6), roles (S4), escalation (S9), communications (S13, S14), postmortems (S15), action tracking (S16), paid on-call (S11), training (S18), metrics (S17).\n- Instrument pilot against S17 metrics; compare results with S2 baseline.\n- Hold weekly retrospectives with pilot teams; iterate on written policies, tooling, training based on feedback.\n- Fix top issues found before wider rollout; document what changed and why.\n- Produce pilot report with before/after numbers (MTTD, MTTR, alert noise, action completion rate) to carry into rollout conversations.\n- Set explicit pilot exit criteria: rotation coverage achieved, zero unacknowledged pages over 2 weeks, all postmortems delivered on time, >80% of action items tracked.", "dependencies": ["S5", "S7", "S8", "S9", "S10", "S11", "S12", "S13", "S14", "S15", "S16", "S18"]}, {"step_id": "S20", "title": "Phased rollout to all 28 teams", "description": "Rollout is a staged migration with readiness gates, not an email announcement.\n\n- Sequence the 28 teams into four waves of ~7 teams each, ordered by customer-impact criticality and readiness; space waves three weeks apart.\n- Define per-team readiness checklist: services mapped and owned, alerts cleaned to standard (S8), runbooks written, rotation staffed, training complete, manager briefed.\n- Hold gate review with program owner before each team joins; move unready teams to next wave with dated remediation plan.\n- Assign named champion per wave; run internal communications cadence explaining why using pilot numbers from S19.\n- Handle resistance directly: publish the \"own-your-code, own-your-pager\" rule and paid on-call mechanics before each wave, not after.\n- Retire legacy tools, informal escalation lists, and ad-hoc status-page process at end of each wave on published cutover date.\n- Harvest feedback formally at each wave and push accepted process changes through change control.", "dependencies": ["S16", "S17", "S19"]}, {"step_id": "S21", "title": "SOC 2 control mapping and evidence framework", "description": "The audit will test incident response. Mapping controls now means evidence accumulates from the first real incident under the new process.\n\n- Map the process to Trust Services Criteria for incident identification, response, evaluation, and communication of security incidents.\n- Write control statements in auditor language; name a single owner per control.\n- Specify the evidence artifact for each control: incident record, severity classification, paging log, communications log, postmortem, action tracker entry, training record.\n- Define evidence retention and storage location (not on laptops, not on chat history that expires); plan for audit access.\n- Identify which controls are blocked until certain rollout waves complete; keep a gap register with owners and review fortnightly with steering group.\n- Run early walkthrough with an experienced compliance partner or pre-audit readiness team to test control design before the audit window.", "dependencies": ["S1", "S3", "S13", "S15"]}, {"step_id": "S22", "title": "SOC 2 dry run, gap remediation, and audit support", "description": "Convert a good process into a provable one, a few months before auditors arrive. Prove the system works at scale.\n\n- Schedule a dry run six weeks before audit window, sampling real incidents from completed waves against each control's evidence requirements.\n- Remediate every gap found; prioritize anything risking a qualified opinion.\n- Brief every engineer who may be interviewed so they can describe the process as practiced.\n- Prepare auditor package: process documentation, sample incident records, training records, on-call schedules, action tracking register, alert quality metrics.\n- Designate a single audit liaison and small evidence-request team so requests do not scatter across teams.\n- Rehearse IC and Communications Lead roles under interview conditions; auditors probe realism under pressure.\n- Ensure all postmortems, incidents, and evidence are retained, searchable, and accessible to auditors for the required audit period.", "dependencies": ["S20", "S21"]}, {"step_id": "S23", "title": "Standing governance, process ownership, and continuous improvement", "description": "The classic post-audit failure is the process freezing and decaying. Lock in continuous improvement as a permanent structure.\n\n- Establish a standing Incident Management Council chaired by the program owner, meeting monthly with engineering, support, compliance, and product representation.\n- Give program owner documented mandate to change standards; require formal change control for any change to severity, roles, communications timings, or compensation.\n- Re-validate severity taxonomy quarterly against real declarations; re-baseline metrics annually.\n- Feed incident themes into architecture review and hiring so the program improves the system, not just the response.\n- Report quarterly to executive team on metric set (MTTD, MTTR, SLA credits, customer-detected %) and top five systemic causes of incidents.\n- Fund a resilience roadmap separate from incident response: reducing shared-database blast radius, multi-region failover, deploy safety.\n- Refresh training and tabletop program annually and after any SEV1.\n- Keep public backlog of improvement ideas (Jira, shared doc) so any engineer can suggest; publish monthly \"incident newsletter\" to all engineers with wins and learnings.", "dependencies": ["S20", "S22"]}], "estimated_complexity": "high", "success_metrics": "- Median time to detect reduced from 22 minutes to under 5 minutes by month 6, with >90% internal detection (vs. 40% customer-detected now).\n- Median time to mitigate reduced from 3h 10min to under 60 minutes for SEV1 and SEV2 by month 9.\n- Customer-impacting incidents detected by customers drop from 40% to <5% of all customer-impacting incidents.\n- Alert volume reduced from 3,400 per month to <600 per month; signal-to-noise ratio improves from 15:85 to >95:5.\n- SLA credits paid reduced from $1.3M to <$100K annually by month 12.\n- Zero incidents with command-and-control ambiguity lasting >15 minutes; all SEV1/2 incidents have named IC logged in timeline within 5 minutes.\n- Postmortem action item completion rate reaches >80% (from 11 of 64, or 17%) by month 4.\n- 100% of SEV1 and SEV2 postmortems published within 15 business days by month 5.\n- All 28 teams integrated into incident management system with active on-call rotations by week 20; no team unresponsive to pages for >30 minutes.\n- On-call satisfaction score reaches >7/10 on survey; zero on-call-attributed voluntary attrition by month 6.\n- Incident commander roster: 40+ certified ICs covering 24x7 with no single point of failure by month 4.\n- Status-page first update published within 30 minutes on ≥95% of SEV1 incidents by month 3.\n- SOC 2 Type II audit passes incident response controls with zero findings by month 8.\n- Weekly incident review cadence sustained in ≥90% of weeks; monthly reliability reviews 12 of 12; quarterly executive reviews 4 of 4.\n- 100% of new engineers complete incident-response onboarding within 30 days of joining."}P2 kept its 21-step shape but added several mechanisms that close real gaps rather than restating policy: the SOC 2 evidence clock, severity x class, the Triage Owner rule, three rotations, a priced opt-out, an action-item cap and a regulator clock matrix. It also made its own targets more realistic.
- Step 1's evidence clock: names the Type II operating-period problem and starts evidence collection in week one; step 3 moves control mapping and the auto-assembled "golden incident file" into month one.
- Step 5's Triage Owner rule — ownership from first acknowledgement until an IC takes over — targets the actual failure mode (the unowned gap before declaration), which round 0 only addressed at declaration.
- Step 4's severity x class matrix with "class can raise a response, never lower it" gives a data-integrity SEV-2 a SEV-1 posture, a real improvement for a ledger business.
- Step 6 adds a Watch state with a 30-minute timer, free false declarations tracked as a metric, the two-responder ambiguity rule and a no-IC-runs-two-incidents rule.
- Step 8 adds a paid Platform Duty rotation for the shared PostgreSQL/Kubernetes estate — the one place "own your code" breaks down — and caps load at two weeks per quarter by tool configuration.
- Step 9 prices the opt-out (the team buys coverage from a paid pool) and publishes an amnesty so incident records never reach performance reviews.
- Step 15 diagnoses 11-of-64 as an over-generation problem and caps postmortems at three actions with artifact-based definitions of done.
- Step 12 adds a regulator clock matrix naming NYDFS Part 500, money-transmitter, breach and card-network windows; step 13 adds a customer-impact ledger feeding comms, credits, regulatory reporting and ROI.
- Step 17's "never publish incident count as a team metric" plus reporting metrics (near-misses, detection gaps, false declarations) guards against hidden incidents.
- IC roster target cut from 40 to 12–16 certified ICs — more credible staffing.
- Step 18 (pilot) still depends on 12 prior steps, so the pilot cannot begin until nearly everything is built; no minimum viable process is defined for the interim weeks.
- Success-metric list has grown to 23 items, several of which (page volume, false-positive rate, no service above two pages) overlap and will be costly to report monthly.
- No costed figure for compensation or tooling despite step 1 promising a funding envelope — P1 at least names $400–600K.
- Alert probation and page budgets are asserted without a fallback if a team simply cannot meet the standard by its wave gate, beyond deferral.
- Proposal 3 : Calculate SLA credit avoidance against programme cost to prove ROI to leadership.
- Proposal 1 : The pager and incident tooling must work from a mobile device.
- Proposal 1 : Severity-keyed playbooks and pre-written messages.
- Proposal 1 : A $50–100 bonus per SEV-1/SEV-2 incident mitigated.
- Proposal 3 : Tie action-item completion rates to team performance reviews.
- Proposal 3 : Consolidate the 12 on-call teams into one unified pool covering all 28 teams, mandatory with no exemptions.
+ Charter, mandate and the evidence clock+ Baseline evidence and problem statement+ Severity times class taxonomy+ Roles, command and the never-without-an-owner rule+ Declaration, lifecycle and escalation policy+ Three on-call rotations across 28 teams+ Compensation, rest and the price of opting out+ Paging contract and alert quality+ Communications: internal, customer and regulator+ Customer-impact ledger and SLA credit automation+ Postmortem policy with three artifact levels+ Action items: capped, verifiable, with a repeat-incident rule+ Pilot with three to four teams, using real incidents+ Phased rollout sequenced by cost of failure+ Audit dry run and evidence reviewProgramme charter, ownership and executive mandateBaseline measurement and evidence packSeverity taxonomy and trigger matrixIncident roles, command structure and decision rightsEscalation, paging and incident lifecycle policyAlert quality standard and noise-reduction programmeOn-call architecture and 24x7 coverage model across 28 teamsOn-call compensation, wellbeing and sustainability policyInternal and customer communications policy with timing SLAsStatus page, notification tooling and account-manager playbookPostmortem policy, template and blameless review processCorrective action tracking and reliability backlog governancePilot with volunteer teamsPhased rollout to all 28 teamsSOC 2 dry run, gap remediation and audit support
The plan produced
1. Charter, mandate and the evidence clock new
This step turns the CEO's email into a funded programme with one accountable owner and explicit authority, and it starts the SOC 2 clock on day one.
- Appoint a single accountable process owner — a Director of Incident Management reporting to the CTO, with a dotted line to the COO for customer and SLA matters.
- Publish a one-page charter: scope (every customer-impacting, money-moving, security and data-integrity event across all 28 teams and both regions), decision rights during an active incident, and the power to freeze deploys and override team preferences.
- Fix the funding envelope up front: tooling licences, training and drill time, and on-call compensation, with an indicative annual figure and the expected return in avoided SLA credits.
- Start the evidence clock immediately. A SOC 2 Type II report covers a period of operating effectiveness, so evidence cannot be backfilled; an audit in eight months means the operating period begins in the first weeks of this programme.
- Stand up a steering group of CTO, VP Engineering, Head of Support, Head of Compliance and two engineering managers, meeting fortnightly.
- Make incident-process participation a documented performance expectation for every engineering manager, not an optional extra.
- Agree the timeline explicitly: a working process in four months, evidence accumulating from month two, audit-ready by month seven.
2. Baseline evidence and problem statement (after 1) new
You cannot demonstrate improvement without a defensible starting point, and you cannot win the argument about noise without numbers.
- Build a 12-month incident register and re-classify all 31 customer-impacting incidents: date, severity, class, detection source, time to detect, time to mitigate, customers affected, services involved and SLA credits paid.
- Run an alert census per tool, per team and per service: total volume, page-to-action ratio, off-hours interruptions per engineer, and the 50 noisiest rules with a named owner.
- Build a silent-failure register: incidents in which no internal alert fired at all. This is the number that explains the 40% customer-detected rate.
- Survey every on-call engineer and manager on burden, fairness, escalation clarity and pay expectations, targeting above 70% response.
- Interview Support, Account Management and Sales about how customers learn of incidents, what they complain about, and what they are promised.
- Reconstruct the two command-ambiguity incidents minute by minute, from alert to mitigation, to find exactly where ownership lapsed.
- Publish the pack internally as the problem statement and retain it as management-review evidence for the audit.
3. Control mapping and evidence architecture (after 1)
Most programmes leave compliance to the end. This one maps controls in the first month, because the mapping determines what the process must capture from day one.
- Map the process to the relevant Trust Services Criteria for incident identification, evaluation, response, containment, recovery and communication of events (CC7.1 to CC7.5, CC2.2, CC2.3, CC4.1, CC3.x).
- Write each control as a plain-language statement with one named owner and its evidence artifact: incident record, severity classification, paging log, communications log, postmortem, action tracker entry, training record.
- Define the golden incident file: one auto-assembled, single-click export per incident containing timeline, roles, severity, communications, postmortem, actions and closure evidence.
- Set retention, storage location and immutability so no control depends on a laptop, a private Slack channel or chat history that expires.
- Keep a gap register with owners and dates, reviewed fortnightly by the steering group.
- Run an early design walkthrough with the auditor's readiness team inside the first 90 days, to test the design before building on it.
4. Severity times class taxonomy (after 2) new
Severity alone is not enough. Severity decides how loud the response is; class decides which playbook runs and what must never be traded away.
- Define severity by impact in customer and money terms: customers affected, dollars at risk, regions lost, ledger integrity, settlement delay, data loss and SLA exposure.
- Use SEV-0 for security, privacy and regulatory events; SEV-1 for total or material loss of a payment path; SEV-2 for degradation or single-region loss; SEV-3 for limited impact with a workaround; SEV-4 for internal-only issues and near-misses.
- Define response classes: Availability, Performance, Data integrity and Ledger, Security and Privacy, Third-party dependency, Settlement and Financial, and Process failure.
- Class can raise a response, never lower it. A SEV-2 data-integrity incident gets SEV-1 posture, because integrity failures are not recoverable by moving faster.
- Specify automatic triggers: loss of one AWS region, ledger write failures, replication lag above threshold, a missed settlement cut-off, payment success rate below threshold for five minutes.
- State who may declare: any engineer, Support agent or account manager. State who may downgrade: the Incident Commander alone.
- Map every level to its SLA credit exposure and to the customer-visible status page state.
- Include worked examples from the last 12 months so each of the 28 teams recognises its own incidents, and re-validate the taxonomy quarterly against real declarations.
5. Roles, command and the never-without-an-owner rule (after 4) new
The two incidents where nobody was in charge for over an hour did not fail at declaration. They failed in the gap before it, when an alert had fired and no one owned it.
- Create one-page role cards: Triage Owner, Incident Commander, Deputy IC, Operations Lead, Internal Communications Lead, Customer Communications Lead, Scribe, Subject-Matter Responders and Executive Sponsor for SEV-1 only.
- Introduce the Triage Owner rule: from the moment a page is acknowledged, that person owns the incident until an IC takes over or the incident is stood down. There is never an unowned minute between first page and close.
- State the core rule plainly: the IC owns the incident, not the fix, and does not debug. An IC who starts troubleshooting has abandoned command.
- Give the IC explicit decision rights: declare and escalate severity, freeze changes, halt deploys, approve customer messaging, and pull in any engineer in the company.
- Define minimum viable staffing per severity: SEV-1 staffs every role; SEV-2 staffs IC, scribe, comms and responders; SEV-3 staffs an IC and a scribe only.
- Set handover discipline: four-hour maximum IC shifts on SEV-1, a written handover template, and a deputy named within 15 minutes of declaration.
- Set behaviour standards for responders: one incident channel, one bridge, no side channels, and asks phrased with a named owner and a time.
- Publish the role cards on the internal wiki and link them from every paging notification.
6. Declaration, lifecycle and escalation policy (after 5) new
This step defines the mechanical path from an alert to a declared incident and back to normal service, and it removes judgment calls from the worst moments.
- Define lifecycle states with entry and exit criteria: Detected, Triaged, Declared, Mitigated, Resolved, Postmortem, Closed — plus a Watch state with a hard 30-minute timer, after which the incident is either declared or stood down.
- Make declaring free. A declaration that turns out to be a false alarm is closed as a false declaration, with no blame and no follow-up, and it is tracked as a metric so the cost of caution stays visible.
- Apply the ambiguity rule: if two responders disagree about whether this is an incident, it is an incident.
- Set acknowledgement and declaration targets: page acknowledged within five minutes, triage decision within fifteen, declaration within thirty for anything customer-facing.
- Build automatic escalation ladders with timers at every layer: responder, service owner, team manager, IC on-call, VP Engineering. Escalation never requires a human decision and is never criticised.
- Define the unresponsive-team path: fifteen minutes without acknowledgement escalates to the team manager, thirty minutes to the IC, who may then direct any available engineer.
- Define change freeze and rollback authority during SEV-1 and SEV-2, and the single named condition that lifts the freeze.
- Enforce one incident, one record, with the timeline auto-captured from the channel and bridge rather than written from memory afterwards.
- Add the two-simultaneous-incidents rule: a second IC is designated, and no IC ever runs two incidents at once.
7. Detection strategy: journeys, synthetic signals and customer-report intake (after 4)
Customers detected 40% of incidents first. That number is the reason this step exists, and it is fixed by measuring customer journeys rather than infrastructure.
- Define SLIs and SLOs for the top 20 customer journeys — payment initiation, authorisation, settlement, ledger read and write, API availability, webhook delivery and payout — measured per region.
- Require symptom-based alerting on those SLOs instead of cause-based alerting on infrastructure metrics.
- Add synthetic transaction monitoring from outside AWS in both regions plus a third location, on a one-minute cadence for money-moving paths.
- Add ledger and shared-PostgreSQL specific signals: replication lag, connection saturation, write latency, lock waits, transaction ID exhaustion, checkpoint pressure and bloat.
- Open a customer-report intake: Support and account managers can raise an incident directly, every customer report creates an incident record, and the customer-report path is counted as a detection source in all reporting.
- Apply the detection-gap rule: whenever a customer reports an incident before internal monitoring, a detection-gap ticket is opened automatically and owned by the responsible team.
- Publish a detection contract per service: every one of the 180 services needs an owner, at least one symptom alert, and a documented expected detect time.
- Measure current detect time per journey, set targets, and run a detection drill per team: break something in staging and see whether it pages before a human notices.
8. Three on-call rotations across 28 teams (after 5) new
The objection is that engineers will not carry a pager for another team's code. The answer is not to argue with it, but to build three rotations so the objection becomes structurally impossible.
- Run a Service On-Call rotation per team, covering only that team's own services. No engineer is ever paged for code their team does not own.
- Run a Platform Duty rotation for genuinely shared infrastructure: the shared PostgreSQL cluster, Kubernetes, networking, CI/CD and observability. This is nobody's product code, so it gets its own paid rotation, staffed from platform teams plus volunteers from other teams.
- Run a central Incident Commander roster of 12 to 16 certified senior engineers drawn from across all 28 teams, covering 24x7 on one-week shifts with a primary and a secondary.
- State the consequence honestly: 16 of 28 teams have no on-call today. Each must either build a rotation or formally transfer ownership of its services to a team that will, with the transfer dated and recorded.
- Merge small or low-traffic teams into shared rotations so no rotation has fewer than six engineers, the level below which coverage stops being sustainable.
- Cap load in the scheduling tool: no engineer is on-call more than two weeks per quarter, enforced by configuration rather than negotiation.
- Publish a coverage matrix for all 28 teams showing services, owners, rotation size and gaps, reviewed monthly.
- Write on-call participation into engineering job levels and hiring criteria, so the commitment is set before someone joins rather than after.
9. Compensation, rest and the price of opting out (after 8) new
Unpaid on-call is the most cited reason for resistance. The policy must be settled before rollout, not negotiated during it.
- Move immediately to paid on-call: a per-shift stipend benchmarked to the New York market, agreed with HR and Finance, and published with an effective date before any team is asked to join a new rotation.
- Pay event-based compensation for out-of-hours callouts, with a minimum call-out block and a 1.5x rate for time actually spent mitigating.
- Provide documented compensatory rest: no engineer works a normal day after a night incident, and the rest day is policy rather than a favour granted by a manager.
- Cap intrusion with a maximum number of off-hours pages per shift, and trigger a mandatory review whenever the cap is exceeded.
- Allow opt-out but put a price on it: an engineer may step out of a rotation, and their team must buy coverage from the paid pool at a published internal rate. This turns a cultural argument into a visible budget decision.
- Publish an explicit amnesty: incident records, near-miss reports and false declarations are never used in performance reviews. Only failure to report is a performance issue.
- Protect against gaming: pay attaches to shifts and callouts, never to incident counts.
- Review the policy every six months against real page volumes, attrition and survey results.
10. Paging contract and alert quality (after 2, 7) new
3,400 alerts a month at 85% noise is why engineers resent the pager. Fixing that is the price of admission for everything else in this plan.
- Publish the paging contract: every page must be symptom-based, actionable, owned, mapped to a severity and a class, and linked to a runbook. No runbook, no page, enforced by a CI check on the alert definition.
- Separate paging from ticketing: only customer-impacting or imminent-impact signals may wake a human; everything else becomes a ticket or a dashboard entry.
- Set a page budget per team and per service: a maximum number of pages per on-call shift. Breaching it auto-opens a remediation ticket with the engineering manager as owner.
- Put new alerts on two-week probation: a new rule runs as a ticket only and becomes a pager only after it has proved actionable, so teams stop being woken by untested rules.
- Define the default action for a noisy alert: fix, tune or delete within ten working days, with deletion treated as a legitimate and celebrated outcome and the deleted-alert count published.
- Correlate and deduplicate at the ingest pipeline so one root cause produces one page instead of forty.
- Require an expiry date on every silencing rule and temporary threshold, so suppression cannot quietly become permanent blindness.
- Run a 90-day noise sprint with a public burn-down of the top 100 noisiest rules, each owned by a named manager, and report page-to-action ratio per team monthly.
11. Incident tooling consolidation (after 5, 10)
Six alerting tools and no single incident record is a structural cause of the 22-minute detection and the three-hour mitigation.
- Select one incident platform for paging, schedules, escalation policies, incident records, postmortem workflow and status-page integration, and choose it within four weeks.
- Consolidate the six alert sources into a single event pipeline feeding that platform, with deduplication, correlation and severity and class mapping applied at ingest.
- Integrate observability, runbooks, chat and voice bridges so responders see dashboards and runbooks inside the incident record and the timeline is captured automatically.
- Implement the golden incident file export defined in S3, so audit evidence is a single click rather than a reconstruction.
- Run a dual-run period alongside the legacy tools with a defined rollback, then switch the legacy tools off on a published date.
- Host the status page outside the production failure domain so it survives a total platform outage, and test that during a game day.
- Make the platform usable from a phone: an Incident Commander must be able to run a SEV-1 from a mobile device at 3am.
12. Communications: internal, customer and regulator (after 4, 5) new
Today the status page is written by whoever is around. This step replaces improvisation with a clock, a named owner and a pre-cleared template.
- Set the internal cadence: first update within 15 minutes of declaration, then every 30 minutes for SEV-1 and hourly for SEV-2, whether or not there is progress.
- Never let the status page be how an employee learns of an incident: internal communication leads, external follows.
- Set the customer cadence: status page updated within 30 minutes of a SEV-1 and 60 minutes of a SEV-2, with a no-new-information update still mandatory.
- Define the channel hierarchy: status page for all 2,100 customers, direct email to affected customers on SEV-1, and named account-manager calls for the top 50 accounts.
- Prepare templates per severity and class in advance, pre-approved by Legal and Compliance, each with the next-update time built in.
- Forbid speculation: customer communications never guess at cause, never assign blame, and never commit to a root cause before the postmortem.
- Define the closing communication: resolution notice, how SLA credits will be handled, and the committed date for the written report.
- Build a regulator clock matrix: for each event type, which regulator, which window, who signs off, and the shortest clock that drives the first action. Cover NYDFS Part 500, money-transmitter and banking notification, security breach notification, card-network rules, and public-company disclosure where applicable.
- Route every regulatory notification through Compliance, never Engineering, and pre-clear the templates.
- Publish a customer-facing root-cause report for SEV-1 incidents, especially for regulated and top-tier accounts.
- Assign a named Customer Communications Lead plus a trained deputy on every SEV-1.
13. Customer-impact ledger and SLA credit automation (after 12) new
The $1.3M in credits is a symptom of having no single record of customer impact. This step creates one, and makes it do four jobs at once.
- Maintain one durable customer-impact ledger per incident: which customers, which journeys, from when to when, and the estimated credit.
- Use that single record for customer communications, SLA credit computation, regulatory reporting and the annual review, so nobody reconciles two versions of the same outage.
- Automate the credit path end to end — computation, approval, customer notification and finance treatment — so credits stop being a manual scramble three weeks later.
- Reconcile accrued against paid credits monthly and report the result in the executive review.
- Track credit avoidance against programme cost, so the funding case is a number rather than an argument.
14. Postmortem policy with three artifact levels (after 5) new
Postmortems currently happen for some incidents, in various formats. This step makes them mandatory where they matter and light where they do not.
- Make postmortems mandatory for every SEV-0, SEV-1 and SEV-2, every SEV-3 with customer impact or a repeat pattern, every near-miss the IC flags, and every incident in which the process itself failed.
- Use three levels so the ritual matches the weight: a lightweight async incident review for SEV-3 and SEV-4, a standard facilitated postmortem for SEV-2, and a full review with an executive sponsor for SEV-1.
- Set deadlines: draft within five business days, blameless review within ten, internal publication within fifteen.
- Adopt one template: impact, timeline, detection, response, contributing factors across tooling, process, organisation and human factors, what went well, what went badly, and actions.
- Train a pool of blameless facilitators and require a trained facilitator for every SEV-1 review.
- Prohibit counterfactual and blame language in the template, and specifically ban the phrase human error as a root cause — the question is always what made the error possible.
- Publish all postmortems internally by default, with security review only for genuinely sensitive material, and produce a customer-facing root-cause variant for SEV-1.
- Treat postmortems as the learning product of the process, not as a compliance artifact.
15. Action items: capped, verifiable, with a repeat-incident rule (after 14) new
Eleven of 64 action items closed is not a tracking problem. It is a generation problem: the process produces more actions than the organisation can absorb.
- Cap each postmortem at three action items. Anything beyond three goes into a ranked reliability backlog, not into the postmortem.
- Require each action to have a named human owner, a due date, and a definition of done that is an artifact: a merged change, a new alert that fired in a drill, or a test that fails without the fix.
- Prohibit self-reported closure. Closure requires the artifact, signed off by the process owner or the Incident Commander.
- Keep one reliability backlog in the engineering tracker with a mandatory label and a link back to the originating incident.
- Reserve a fixed percentage of every team's sprint capacity for reliability work, with unspent capacity visible to the vice presidents.
- Apply the repeat-incident rule: a second incident with the same contributing factor triggers a design review owned by an architect, not another action item.
- Run a weekly ageing review of open actions and escalate anything more than 30 days overdue to the VP of Engineering.
- Target more than 90% of actions closed within 60 days and a median age under 30 days, reported monthly by team.
16. Training, certification and drills (after 5, 6, 12, 14)
A process that exists only on a wiki page fails on the first real page. Skills have to be built and tested before they are needed.
- Build a short practical curriculum: how to be on call, how to triage and declare, how to run an incident as IC, how to communicate, and how to write a postmortem.
- Require certification before joining the IC roster: a written assessment plus a live simulated incident.
- Certify at least two ICs per team group, so the central roster has depth across all 28 teams and no holiday week is uncovered.
- Run monthly tabletops drawn from the last 12 months of real incidents, including region loss, ledger write failure and a missed settlement cut-off.
- Run quarterly game days with deliberately injected failure, including shared-PostgreSQL failover, status-page outage and alerting-pipeline outage.
- Drill the process's own failure modes, not just technical ones: IC unreachable, comms lead on PTO, two simultaneous SEV-1s, a paging storm, and a false alarm that burns an hour.
- Audit the incident process for single points of failure: who is the only person who can do each critical task, and what happens in their holiday week.
- Add a mandatory onboarding module for every engineer joining or transferring in, with completion records kept in an audit-ready form.
- Track training completion by team and publish it in the monthly reliability review.
17. Metrics and review cadence (after 2, 4, 15)
Establish what good looks like, and measure it in a way that makes people report more incidents rather than fewer.
- Define outcome metrics: time to detect by source, time to mitigate by severity and class, percentage of incidents detected by customers, incidents by severity and class, SLA credits paid, and customer-impact minutes.
- Define process metrics, because the process itself has SLOs: IC assigned within five minutes, page acknowledgement rate, first-update timeliness, update-cadence adherence, postmortem timeliness, action closure rate, action median age and IC roster coverage.
- Define health metrics: alert volume and noise ratio per team, off-hours pages per engineer, on-call satisfaction, training completion, and percentage of services with a detection contract.
- Never publish incident count as a team metric. It rewards hiding incidents. Publish reporting metrics instead — near-misses filed, detection gaps closed, false declarations made — alongside the outcome metrics.
- Publish live dashboards visible to every engineer, refreshed daily, with each metric baselined against the S2 evidence pack and given a 90-day and a 12-month target.
- Run a 30-minute weekly operational review incident by incident, a monthly reliability review of trends and actions, and a quarterly executive review with the CEO's office covering customer impact, credits and the top five systemic causes.
- Hold a quarterly review of the process itself: what in the process wasted time, what confused responders, and what should be deleted.
- Require every review to end with decisions and named owners, never with numbers alone.
18. Pilot with three to four teams, using real incidents (after 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17) new
Do not roll out to 28 teams untested, and do not rely on drills alone when real incidents are available and better.
- Recruit three to four volunteer teams spanning criticality: a payment-path team, a ledger-adjacent team, the platform team that owns the shared PostgreSQL cluster, and a low-traffic team.
- Run the entire process end to end in the pilot: the new severity and class scale, roles, escalation, communications, postmortems, action tracking and paid on-call.
- Treat real incidents during the pilot as the primary training material. Hold a retro within 48 hours of each pilot incident, run by the process owner, while the friction is fresh.
- Instrument the pilot against the S2 baseline and publish before-and-after numbers.
- Hold weekly retrospectives with the pilot teams and change the written policies, the tooling and the training in response, documenting what changed and why.
- Set explicit exit criteria: rotation coverage achieved, zero unacknowledged pages over a defined period, postmortems delivered on time, and actions tracked to closure.
- Produce a pilot report that every later rollout conversation starts from.
19. Phased rollout sequenced by cost of failure (after 15, 18) new
Rollout is a staged migration with readiness gates, not an email announcement. The sequencing matters more than the schedule.
- Sequence the 28 teams into four waves of roughly seven, ordered by incident density and customer-journey ownership, not by size or alphabet: the teams whose incidents cost the most SLA credits go first, because their improvement is the visible proof.
- Leave three weeks between waves and define a per-team readiness checklist: services mapped and owned, alerts cleaned to the paging contract, runbooks written, rotation staffed to at least six engineers, training complete, manager briefed, and compensation in effect.
- Hold a gate review with the process owner before each team joins, and move an unready team to the next wave with a dated remediation plan rather than granting an exception.
- Give each wave a named champion and run an internal communications cadence that explains the why using the pilot's numbers.
- Handle resistance directly and early: publish the own-your-code-own-your-pager rule and the paid on-call mechanics before each wave, never after.
- Hold a retro within 48 hours of each wave's first incident under the new process, and feed accepted changes back through change control.
- Retire legacy tools, escalation lists and the informal status-page process at the end of each wave on a published cutover date, so two processes never run in parallel for long.
- Sequence the waves so no team faces both a rollout and the audit dry run in the same week.
20. Audit dry run and evidence review (after 3, 19) new
This step converts a good process into a provable one, a few months before the auditors arrive.
- Run an internal dry run six weeks before the audit window, sampling real incidents from the pilot and the early waves against each control's evidence requirements.
- Remediate every gap found, prioritising anything that risks a qualified opinion, and re-test the remediated control against the same sample.
- Brief every engineer who may be interviewed so they can describe the process as they actually practise it, not as it is written.
- Prepare the auditor package: process documentation, sample incident records, paging logs, communication logs, training records, on-call schedules and the action tracking register.
- Designate one audit liaison and a small evidence-request team, so requests do not land on 28 teams simultaneously.
- Rehearse the walkthrough with an incident commander and a communications lead, because auditors probe realism under pressure, not documents.
- Keep the audit liaison and the process owner as the same person where possible, so accountability for the control is also accountability for the evidence.
21. Standing governance, resilience roadmap and continuous improvement (after 19, 20)
The classic post-audit failure is that the process freezes and then decays. This step keeps it alive and separates incident response from the engineering that prevents incidents.
- Establish a standing Incident Management Council chaired by the process owner, meeting monthly with engineering, support, compliance and product representation.
- Give the process owner a documented mandate to change standards, and require formal change control for any change to severity, classes, roles, communication timings or compensation.
- Re-validate the severity and class taxonomy quarterly against real declarations, and re-baseline the whole metric set annually.
- Fund a resilience roadmap separate from incident response: reduce the shared PostgreSQL blast radius, improve multi-region failover readiness, and tighten deploy safety. Better incident management cannot save a single shared ledger during a corruption event.
- Report quarterly to the executive team on the metric set and the top five systemic causes of incidents, and feed those themes into architecture review, release engineering and hiring.
- Keep a public backlog of process improvement ideas so any engineer can propose one, and change the process from it every quarter.
- Share wins publicly each quarter — faster detection, fewer customer-detected incidents, credits avoided — to sustain momentum after the audit.
- Refresh the training and tabletop programme annually and immediately after any SEV-1.
- Median time to detect falls from 22 minutes to under 8 minutes by month 6, and under 5 minutes for money-moving journeys by month 9.
- Share of customer-impacting incidents first detected by customers falls from 40% to under 10% by month 9.
- Median time to mitigate for SEV-1 falls from 3 h 10 min to under 60 minutes by month 9.
- 100% of SEV-1 and SEV-2 incidents have a named Triage Owner within 5 minutes and a named Incident Commander within 15 minutes, from month 4.
- Zero incidents in which command authority is unclear for more than 15 minutes, measured monthly from month 4.
- Monthly page volume falls from 3,400 to under 600, with a false-positive rate under 15%, by month 5.
- No service exceeds two pages per on-call shift for three consecutive months, measured by month 6.
- All 28 teams have a documented rotation, a service ownership map and at least one trained on-call engineer by month 6.
- Incident Commander roster holds at least 12 certified ICs covering 24x7 with no uncovered week, from month 4.
- 100% of on-call shifts are paid under a published policy from month 2.
- 100% of SEV-0, SEV-1 and SEV-2 postmortems are published internally within 15 business days, from month 5.
- Postmortem action items closed within 60 days rise from 17% to over 90%, with median age under 30 days, by month 6.
- At least one incident review or near-miss report is filed per team per quarter, from month 6.
- Status page first update is published within 30 minutes on at least 95% of SEV-1 incidents, from month 4.
- Zero missed regulatory notification windows on any incident requiring notification.
- The customer-impact ledger is complete for 100% of customer-impacting incidents from month 5, with SLA credits reconciled monthly.
- Annual SLA credits paid fall at least 50% from the $1.3M baseline within 12 months.
- Customer-impacting incidents fall at least 40% year over year from the 31-incident baseline.
- On-call satisfaction is at or above 7 out of 10, with zero voluntary attrition attributed to on-call, measured quarterly from month 6.
- Review cadence is sustained: weekly operational review in at least 90% of weeks, 12 of 12 monthly reliability reviews, 4 of 4 quarterly executive reviews.
- SOC 2 Type II is passed with zero findings related to incident response.
- 100% of new engineers complete the incident-response onboarding module within 30 days of joining.
- 100% of the 180 services have a named owner, a detection contract and at least one symptom-based alert held to the paging contract by month 7.
[SYSTEM]
You are an expert assistant in complex project planning.
Your task is to generate a detailed and structured action plan to reach the main objective in the most professional and most detailed way, no matter how much work or steps will be needed to perform.
Use your internal reasoning processes to think deeply about the problem and create the most comprehensive plan possible.
Take as much time and space as you need to think through all aspects of the problem.
After your thorough analysis, answer with the plan in the requested structure.
Every text field you write will be read by a busy person, so write for them: short sentences, one idea each, in short paragraphs. Text fields accept Markdown: separate paragraphs with a blank line, use a bullet list when you enumerate things and plain paragraphs when you explain or argue, and bold at most one key phrase per paragraph. A text of more than three sentences must be split into paragraphs; never deliver a single unbroken block of text.
[HUMAN]
Main Objective: "A B2B payments platform in New York (260 engineers in 28 teams, 2,100 customers, $4B processed a month, a 99.95% SLA with service credits) runs about 180 services on Kubernetes across two AWS regions, with a shared PostgreSQL cluster for the ledger. In the last 12 months: 31 customer-impacting incidents, median time to detect 22 minutes (customers detected 40% of them first), median time to mitigate 3 h 10 min, $1.3M paid in SLA credits, two incidents in which nobody was sure who was in charge for over an hour, and a CEO email about "outages we hear about from clients". On-call exists in 12 of the 28 teams, unpaid, fed by alerts from six different tools — 3,400 alerts a month, 85% of them noise; there is no severity scale, status-page updates are written by whoever is around, and postmortems happen for some incidents, in various formats, with action items that are rarely tracked (11 of 64 closed). A SOC 2 Type II audit in eight months will test incident response. Engineers push back against "carrying a pager for other teams' code".
Define the incident management process: severity levels and what each one triggers; roles (incident commander, communications, scribe, subject-matter responders) and how they are staffed 24x7 across 28 teams; the on-call structure, rotations, compensation and the rules for alert quality; detection and escalation paths; internal and customer communications (status page, account managers, regulators when required) with their timings; postmortems (when mandatory, format, blameless review, ownership and tracking of actions); the metrics and reviews that show whether it works; and how it is introduced across the teams without waiting for the audit."
For your consideration and refinement, here are proposals from the previous round:
Previous Proposal 1 (ID: 895aa761-b883-4277-9609-6e76ca669198, Agent: claudeHaiku4.5_initial_1, LLM: anthropic/claude-haiku-4-5):
Estimated Complexity: high
Success Metrics: - Median time to detect reduced from 22 minutes to <8 minutes within 6 months
- Customer-detected incidents drop from 40% to <5% within 6 months
- Median time to mitigate (MTTR) reduced from 3h 10min to <45 minutes for SEV-1 incidents within 6 months
- Annual SLA credits decrease from $1.3M to <$100k within 12 months
- Alert noise reduced from 3,400 per month (85% false positive) to <400 per month (>95% signal) within 3 months
- Zero incidents with command-and-control ambiguity (>1 hour without clear IC) within 2 months
- Postmortem action item completion rate reaches >80% (from 17%) within 4 months
- On-call satisfaction score (survey) reaches >7/10 for on-call engineers within 3 months
- All 28 teams integrated into incident management system with active on-call rotations by week 20
- SOC 2 Type II audit passes incident response controls with no findings 8 months from start
- Incident commander certification: 100% of active ICs trained and drilled within 2 months
- Monthly incident review meeting established and attended by leadership; trends documented
- New incident system integration complete: single alert tool, single dashboard, all 180 services feeding in, <5 min deployment
Steps (21):
1. Define severity levels and decision criteria
Create a four-tier severity framework (SEV-1 through SEV-4) that guides all downstream decisions about response, escalation, and communications.
Each level must specify: customer impact (revenue at risk, customers affected, data loss risk); financial threshold triggering service credits; whether an incident commander is required; response time SLA (e.g., SEV-1 < 5 min notification, SEV-4 < 2 hours); and the go-live decision tree (when to declare and when to resolve).
- SEV-1: Complete service down or critical path broken for >5% of customers; every minute costs money; IC required; 99.99% uptime threatened
- SEV-2: Significant degradation, features unavailable, affecting 1–5% of customers; IC typically required
- SEV-3: Minor impact, limited customer footprint or workaround exists; escalation path but not automatic IC
- SEV-4: Observations or minor issues; alert-driven, no escalation unless pattern emerges
2. Define incident roles and responsibilities (depends on: 1)
Create the organizational roles that operate during an incident: who is in charge, who talks to customers, who writes down what happened, who fixes the system, and how decisions are made under pressure.
Each role must have a single-sentence mission, decision authority, and escalation upward.
- Incident Commander: owns decision-making and timeline; declares severity; resolves conflicts; may or may not be technical
- Deputy IC: shadow to IC and takes over if IC becomes unavailable
- Communications Lead: writes status page, notifies account teams, manages customer perception
- Scribe: records decisions, who did what, key timestamps; not responsible for fixing
- SME Responders: engineers with context on the failing service(s); take IC's direction without debate
3. Design 24x7 on-call rotation structure (depends on: 2)
Build a rotation model that covers all 28 teams with primary and backup on-call engineers every hour across weekdays, evenings, weekends, and holidays; addresses the pager-carrying resistance.
Key design decisions: Is coverage per-team (each team owns its services) or pooled (shared responder pool handles anything)? How many people per rotation? How long are shifts (one week, two weeks)? When can engineers opt out without leaving the team exposed? Which roles are on-call (IC, communications, SME)?
- Recommend: dedicated IC pool (4–6 people in fast rotation) + per-team SME on-call for each team's own services
- Recommend: two-week rotation blocks to reduce handoff friction
- Recommend: one primary, one secondary per slot; secondary handles during primary's escalation
- Provide swaps, blackout dates, and a rule that no engineer is on-call more than 2 weeks per quarter
4. Define on-call compensation and incentives (depends on: 3)
Create a pay model that makes on-call acceptable and rewards engineers who carry the pager; ties compensation to real business risk.
- Base on-call stipend: e.g., $500–1,000 per week while on-call (regardless of incidents)
- Callback pay: 1.5× hourly rate for time spent mitigating incidents during off-hours
- Incident bonus: $50–100 extra per SEV-1 or SEV-2 incident mitigated (recognition)
- Comp time: full business day off after an incident that required >2 hours mitigation during night/weekend
- Annual bonus tie-in: 10–20% bonus multiplier for flawless on-call reviews
- Communicate: position as investment in reliability, not punishment for being online
5. Consolidate alert routing infrastructure
Replace six alert tools with a single ingestion and routing system; stop engineers from being woken by duplicate alerts, and make escalation automated instead of manual.
Evaluate existing tools (likely candidates: PagerDuty, Opsgenie, or Incident.io) or build a lightweight wrapper. The system must: accept alerts from all 180 services; deduplicate and correlate (same outage, different monitoring source); route to correct on-call engineer; expose an API for playbook automation; log every alert for postmortem analysis.
- Choose tool by week 2 of S5
- Migrate alerting endpoints from 6 sources to 1 by week 4
- Set up audit trail and retention
- Ensure mobile app works (on-call engineers need to engage from phone)
6. Build alert quality rules to cut noise (depends on: 1, 5)
Implement rules that automatically suppress the 85% of alerts that are noise (flapping, transient errors, auto-recovered conditions). Target: <400 actionable alerts per month.
Rules to implement: suppress alerts if service auto-recovered within 30 seconds; deduplicate same alert from multiple monitoring sources; suppress alerts for known maintenance windows; group flapping alerts (same service, >5 occurrences in 2 minutes) into a single page to on-call; rate-limit alerts from noisy services (e.g., max 1 alert per 5 minutes per service until silence clears).
- Audit existing 3,400 alerts per month: which are true signals, which are noise
- Tag each alert source with severity level (S1, S2, S3, S4 from S1)
- Create exceptions list: services known to be noisy, require different rules
- Weekly review: alert teams that trigger >50 alerts per week for reduction strategies
7. Implement automated detection and escalation paths (depends on: 1, 5, 6)
Wire the alert system to automatically escalate based on time or severity; removes the need for manual judgment calls during chaos.
Logic: SEV-1 alert arrives → IC notified instantly via phone call + SMS + Slack + mobile; if IC does not acknowledge within 2 minutes, page deputy IC; Communications Lead pinged simultaneously. SEV-2: on-call SME for that service + IC notify via Slack and mobile, escalate to IC's manager if not acknowledged in 10 min. SEV-3/4: on-call SME only, escalate after 30 min.
- Implement in alert routing system (S5)
- Test all paths weekly via synthetic page to on-call
- Track escalation metrics: how many pages reach secondary, how many hit manager
- Adjust timing based on first month of operations
8. Build incident dashboard and status tracking (depends on: 5)
Create a single source of truth during an incident that every responder sees in real time: who is on-call, incident timeline, who said what, current status, next steps.
Dashboard displays: active incidents and their severity; who is the IC and communications lead; timeline of all events (alert fired, IC assigned, customer notified, mitigation started, resolved); Slack channel and mobile notification status; on-call rosters (who is on-call right now for each team); postmortem link as soon as incident closes.
- Integrate with alert tool (S5) to auto-populate incident creation and initial severity
- Push updates to status page and customer account managers automatically
- Log all timeline entries for audit and postmortem completeness
- Mobile-optimized so IC can work from any device
9. Write incident playbooks for each severity (depends on: 1, 2, 3)
Create a one-page (or one-screen) reference for the IC and SMEs during an incident; sequences the steps and removes ambiguity.
Each severity level gets its own playbook: who gets paged (roles, order); first questions to ask (is it real, how big, who knows); what the IC should declare in first message (status page text, account manager notification, regulatory trigger); how long before escalating to executive team; decision rules for going dark vs. continuing to update customers.
- SEV-1 playbook: immediate IC + comms + CTO notification; customer status every 5 minutes
- SEV-2 playbook: IC + comms + tech lead notification; status every 15 minutes
- SEV-3 playbook: on-call SME + comms if customer-visible; status every 30 min or as resolved
- SEV-4 playbook: on-call SME only; update customers only if promised SLA is at risk
- Include decision trees: is this SEV-1 or SEV-2? Is it our code or dependency? Escalate or containment?
10. Define internal communication workflows (depends on: 2, 3)
Specify who informs whom, in what order, via what channel (call, Slack, email) during an incident; prevents gaps like "nobody knew who was in charge for an hour."
Workflow for SEV-1: IC assigned → IC calls CTO/VP Eng and incident channel lead within 1 minute; incident declared in #incidents Slack channel with severity, IC name, service affected; SME on-call for that service joins call automatically; IC pushes updates to #incidents every 5 minutes or when material change occurs. For SEV-2: IC notifies team leads via Slack, updates #incidents every 15 min. Define escalation: if IC is unreachable, deputy IC takes over and announces it.
- Create a phone tree or on-call list accessible to responders
- Set expectations: "If you don't hear from IC in 2 minutes, call them"
- Use a single incident Slack channel per incident (auto-created by incident tool)
- Log all comms in the incident dashboard for postmortem review
11. Design customer communication and status page process (depends on: 1, 2)
Plan when and how to inform customers, account managers, and regulators; ensure 2,100 customers are not learning about outages from Twitter before you tell them.
Rules by severity: SEV-1 detected → status page updated within 3 minutes (even if root cause unknown; post "Investigating"); account managers of affected customers called within 5 minutes; regulatory notification (if payment processing down) queued for approval; customer email within 10 minutes with ETA for next update. SEV-2: status page within 10 min, account managers called within 15 min, email if affecting >10 customers. SEV-3/4: no customer communication unless SLA at risk.
- Empower Communications Lead to update status page without IC approval if delay >3 min
- Prepare templated messages for common scenarios (database failover, data pipeline stuck, service crashed)
- Route regulatory notifications through legal/compliance; don't wait for perfect root cause
- Track customer impact in real time: how many customers affected by severity
12. Establish blameless postmortem process and format (depends on: 1, 2)
Build a systematic way to learn from incidents so the same failure does not happen twice; counter the fear that admitting a mistake leads to being blamed.
Mandatory postmortems: all SEV-1 and SEV-2 incidents, within 48 hours of resolution. Optional but encouraged: SEV-3 if interesting or if >3 of same type in 30 days. Format: what was the user-visible impact and for how long; what was the root cause (not "human error" but the system condition that made error possible); timeline of discovery and response; action items with owner and deadline; blameless tone (focus on process and system design, not individual mistakes).
- Assign a facilitator (not the on-call IC) to run postmortem
- Attendees: IC, comms lead, SMEs involved, team lead, customer success if customer-facing
- Write postmortem in shared doc; make it findable (searchable, linked from incident)
- No discussion of "who screwed up"; only "why did the system allow this to happen"
13. Build action item tracking and accountability (depends on: 12)
Create a system that tracks postmortem action items so they are not forgotten; currently 11 of 64 (17%) are being tracked, leaving 53 unfinished improvements.
System: each postmortem generates action items (e.g., "add monitoring for X," "update runbook for Y," "write test for Z"). Each item gets: clear description, owner (engineer's name), due date (1–4 weeks based on priority), severity (critical = must do before similar incident happens again; important = improve next month; nice-to-have = backlog). Action items live in a dedicated Jira project visible to all teams; owners are accountable (their manager reviews quarterly). Weekly: incident commander reviews open items due that week. Monthly: each team's postmortem items reviewed in their standup.
- Export action items from postmortem document to tracking system automatically
- Require IC to sign off that an action is complete before closing
- Report on completion rate as a metric (target: >80% by month 3)
14. Define incident metrics and KPIs
Establish what "good" looks like; measure so you can improve. Target metrics for 12 months out: mean time to detect 8 minutes (vs. 22 now), customers detect first <5% of incidents (vs. 40%), MTTR 45 minutes for SEV-1 (vs. 190), SLA credits <$100k/year.
Metrics to track: (1) MTTD = time from incident start to first alert/report; disaggregate: external report vs. internal detection. (2) MTTR = time from first report to full mitigation; track by severity and by service. (3) Customer-reported incidents per month (should drop to <2 per month). (4) Alert signal-to-noise ratio (goal: <5% false positive after S6 rules). (5) On-call satisfaction (survey: would you do this again?). (6) Postmortem action completion rate. (7) Incident commander and responder utilization (hours per week per person).
- Dashboard: auto-populated from incident tool, updated daily
- Disaggregate by team and service: which teams have bad MTTR? Which service is most incident-prone?
15. Create review cadence and governance process (depends on: 14)
Establish regular rhythm to inspect the metrics, spot trends, and adjust the process itself; prevent the system from calcifying.
Weekly: incident commander and on-call lead review prior week—number of incidents, any escalations, any communication gaps. Monthly: director-level incident review—trends by service, top causes of incidents, action item status, whether severity classification is working. Quarterly: full leadership review—MTTD, MTTR, customer impact, SLA credit spend, on-call satisfaction score, any systemic changes needed. Annually: audit the entire process for SOC 2 compliance.
- Assign meeting owners: weekly = on-call lead; monthly = director of reliability; quarterly = VP Eng + CFO (SLA cost) + customer success
- Use same data dashboard (S14) for all reviews
- Publish a monthly "incident newsletter" to all engineers: what happened, what we learned, what's improving
16. Prepare SOC 2 Type II audit checklist (depends on: 1, 2, 9, 12, 13, 14, 15)
Document that the incident management system meets the control requirements for a SOC 2 audit; audit happens in 8 months, so this work builds confidence in coverage.
Audit will test: (1) Is there a defined incident response process? (2) Are roles and responsibilities clear? (3) Are incidents logged and tracked? (4) Is root cause analysis performed? (5) Are action items tracked and completed? (6) Is on-call staffing adequate? (7) Are communications timely? (8) Are postmortems documented and blameless? Create a control mapping document that links each SOC 2 requirement to your process (S1–S15). Collect evidence: incident logs, postmortem documents, action item tickets, metrics reports, training records.
- Designate a compliance owner (often a reliability lead or security engineer)
- Run a mock audit at month 6 to identify gaps
- Ensure all postmortems and incidents are retained and searchable for auditor review
17. Develop implementation and rollout plan (depends on: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16)
Create a phased timeline to roll out the incident management system across all 28 teams; avoids big-bang failure and builds credibility.
Recommended structure: Phase 1 (weeks 1–4): build and test infrastructure (S5, S8, alerting); deploy severity levels and roles (S1, S2); pick pilot teams (2–3 high-traffic teams). Phase 2 (weeks 5–12): train pilot teams, run incident drills, refine playbooks based on learning; expand to half of remaining teams. Phase 3 (weeks 13–20): full rollout to all 28 teams; continue drills; track metrics. Phase 4 (weeks 21–28): stabilize, iterate on metrics, prepare for audit.
- Assign a release manager to coordinate across teams
- Create a detailed Gantt chart with swim lanes (infra, process, training, rollout)
- Identify risks: competing priorities, engineers worried about pager burden, tool adoption friction
- Plan stakeholder engagement: weekly updates to eng leadership, monthly town halls for all engineers
18. Build training and documentation (depends on: 17)
Create role-specific education so engineers understand the new system and are confident executing during an incident.
Training tracks: (1) All engineers: 30-min video on severity levels, communication expectations, postmortem format, action items (mandatory); (2) On-call responders: 2-hour workshop on playbooks, alert tool, escalation paths, when to call manager, case studies of real incidents, mobile app walkthrough; (3) Incident commanders: 4-hour IC bootcamp covering leadership under pressure, decision-making, communicating with executives, status page discipline, postmortem facilitation, practiced drills; (4) Communications leads: templates, when to update, how to talk to customers, regulatory notification rules.
- Record videos so async teams can learn on their schedule
- Create runbooks and quick-reference cards for each role (print + digital)
- Pair new on-call engineers with experienced responder for first week
- Require IC certification before anyone joins IC rotation (pass a practical drill)
19. Execute staged rollout across teams (depends on: 18)
Progressively activate the incident management system with feedback loops at each stage; reduces risk of system-wide failure.
Wave 1 (week 6–8): 3–4 pilot teams begin on-call rotations and incident response using new system; capture feedback daily. Wave 2 (week 10–14): 8–10 additional teams, incorporating lessons from Wave 1; ensure diversity of team types (payment processing, monitoring, data pipeline, auth, etc.). Wave 3 (week 15–20): remaining teams; by now, the system is proven and less hand-holding needed.
- Daily retros with Wave 1 teams: what worked, what was confusing, what broke
- Each wave produces a "lessons learned" document that informs the next
- Track adoption metrics: how many incidents reported per team, alert quality, MTTD/MTTR
- Address resistance: engineers who are skeptical of the system, on-call burden, tool friction; assign a "change champion" in each team
20. Run incident response drills and simulations (depends on: 19)
Practice incidents in a controlled setting so responders gain confidence and gaps are found before a real crisis.
Drill 1 (week 8, pilot teams only): simulated database failover; test alert routing, IC assignment, on-call notification, status page update, postmortem. Drill 2 (week 12, all teams): full incident spanning multiple services; test escalation, communications, decision-making. Drill 3 (week 18): chaos drill with intentional confusion—what if IC is unreachable, what if comms lead is on PTO, what if an alert tool goes down? Quarterly thereafter: repeat drills with new on-call engineers.
- Schedule drills at a time when key people are available (not Friday afternoon)
- Run drill with live observers (reliability lead, manager) who provide feedback
- Measure drill success: all steps completed on time, communications clear, postmortem written
- Debrief immediately: what did we learn, what do we change?
21. Establish continuous improvement process (depends on: 15, 20)
Build a feedback loop so the incident management system evolves based on what you learn.
Monthly improvement cycle: in the monthly incident review (S15), leadership identifies one or two friction points from recent incidents (e.g., "alert tool response is slow," "postmortem meetings are boring," "action items keep slipping"). Assign an owner to investigate and propose a fix (change to playbook, tweak to tool config, additional training, process change). Test the fix with pilot team, gather feedback, roll out if successful. Annual review: can we reduce on-call burden, simplify communication, improve MTTD further? Do any teams need special handling (new services, difficult customers)?
- Keep a public backlog of improvement ideas (Jira, shared doc) so any engineer can suggest
- Quarterly all-hands: share wins ("MTTD improved 50%," "on-call satisfaction up," "zero customer-detected incidents") to build momentum
- Link improvements back to metrics (S14): show that changes actually worked
Previous Proposal 2 (ID: bdeac435-cbbd-4c22-bd5c-08a0ab80adec, Agent: deepseek-flash_initial_2, LLM: deepseek/deepseek-flash):
Estimated Complexity: high
Success Metrics: - Median time to detect reduced from 22 minutes to 5 minutes or less by month 6.
- Customer-detected incidents reduced from 40 percent to 10 percent or less of all customer-impacting incidents by month 9.
- Median time to mitigate reduced from 3 h 10 min to 60 minutes or less for SEV1 and SEV2 by month 9.
- Customer-impacting incidents reduced by at least 40 percent year over year from the 31 incident baseline.
- SLA credits paid reduced by at least 50 percent year over year from the $1.3M baseline.
- 100 percent of SEV1 incidents have a named Incident Commander within 5 minutes and a scribe within 15 minutes.
- Zero incidents in which command authority is unclear for more than 15 minutes.
- 28 of 28 teams have a documented on-call rotation, an owning service map and at least one trained on-call engineer by month 6.
- Central Incident Commander roster holds at least 40 certified ICs covering 24x7 with no single point of failure.
- 100 percent of on-call rotations are paid under a published policy by month 5.
- Monthly alert volume reduced from 3,400 to below 700, with a false-positive rate below 20 percent.
- No service exceeds 2 pages per on-call shift, measured monthly for three consecutive months.
- 100 percent of SEV1 and SEV2 postmortems published internally within 15 business days.
- At least 90 percent of postmortem action items closed within 60 days, up from 17 percent (11 of 64).
- Status-page first update published within 30 minutes on at least 95 percent of SEV1 incidents.
- Zero missed regulatory notification windows on any incident requiring notification.
- SOC 2 Type II audit passed with no findings related to incident response.
- Review cadence sustained: weekly operational review in at least 90 percent of weeks, 12 of 12 monthly reliability reviews, 4 of 4 quarterly executive reviews.
- On-call satisfaction at 70 percent or higher on the quarterly survey, with zero on-call-attributed voluntary attrition.
- 100 percent of new engineers complete the incident-response onboarding module within 30 days of joining.
Steps (21):
1. Programme charter, ownership and executive mandate
This step turns the CEO's email into a funded programme with a named owner and explicit authority. Without it, every downstream decision stalls in cross-team negotiation.
- Appoint a single accountable process owner (for example a Director of Incident Management) reporting to the CTO, with a dotted line to the COO for customer and SLA matters.
- Publish a one-page charter covering scope (all customer-impacting and money-moving incidents), decision rights, and the power to override team preferences during an active incident.
- Define the funding envelope: tooling licences, training time, exercise time and on-call compensation, with an indicative annual figure.
- Set the timeline against the SOC 2 date: a working process in four months, evidence accumulating from month two, audit-ready by month seven.
- Stand up a steering group with CTO, VP Engineering, Head of Support, Head of Compliance and one engineering manager per region.
- Agree that incident-process participation is a documented performance expectation for engineering managers, not an optional extra.
2. Baseline measurement and evidence pack (depends on: 1)
You cannot demonstrate improvement without a defensible starting point, and you cannot win the argument about noise without numbers.
- Build a 12-month incident register: date, retro-assigned severity, detection source, time to detect, time to mitigate, customer impact, services involved and SLA credits paid.
- Quantify the alert estate per tool, per team and per service; compute page-to-action ratio, list the 50 noisiest rules and count off-hours interruptions per engineer.
- Survey on-call engineers and managers on burden, fairness and clarity of escalation, targeting a response rate above 70 percent.
- Interview Support, Account Management and Sales about how customers learn of incidents and what they complain about.
- Document exactly where the current process breaks: unclear command in the two known incidents, postmortem action closure at 11 of 64, and ad-hoc status-page authorship.
- Publish the pack internally as the problem statement and retain it as management-review evidence for the SOC 2 audit.
3. Severity taxonomy and trigger matrix (depends on: 1, 2)
Severity is the keystone of the whole process. Every other rule, from paging to communications timing to postmortems, is keyed off it.
- Define four levels plus a special SEV0 for security or regulatory events: SEV1 for total or material loss of a payment path, SEV2 for degradation or single-region loss, SEV3 for limited impact with a workaround, SEV4 for internal-only issues and near-misses.
- Anchor each level in customer and money terms: customers affected, dollars at risk, regions lost, ledger integrity, settlement delay.
- Specify automatic triggers, for example loss of one AWS region, ledger write failures, a missed settlement cut-off, or payment success rate below threshold for five minutes.
- State who may declare each level (any engineer, Support or account manager may declare) and who may only recommend a downgrade (the Incident Commander alone).
- Map each level to SLA credit exposure and to the customer-visible status-page state.
- Include worked examples from the last 12 months so teams recognise their own incidents in the definitions.
- Add a review clause: the taxonomy is re-validated quarterly against real declarations.
4. Incident roles, command structure and decision rights (depends on: 2, 3)
The two incidents where nobody knew who was in charge for over an hour are the direct brief for this step.
- Define roles with one-page role cards: Incident Commander, Deputy IC, Operations Lead, Internal Communications Lead, Customer Communications Lead, Scribe, Subject-Matter Responders and an Executive Sponsor for SEV1 only.
- State the core rule plainly: the IC owns the incident, not the fix, and does not debug.
- Give the IC explicit decision rights: declaring and escalating severity, freezing changes, halting deploys, approving customer messaging and calling additional responders.
- Define minimum viable role coverage per severity: SEV1 staffs every role, SEV3 staffs an IC and a scribe only.
- Define handover discipline: maximum four-hour IC shifts on SEV1, a written handover template, and a Deputy IC nominated within 15 minutes.
- Set behaviour standards for responders: one incident channel, one bridge, no side channels, and explicit asks with a named owner and a time.
- Publish role cards on the internal wiki and link them from every paging notification.
5. Escalation, paging and incident lifecycle policy (depends on: 3, 4)
This step defines the mechanical path from an alert to a declared incident and back to normal service.
- Define lifecycle states with entry and exit criteria: Detected, Triaged, Declared, Mitigated, Resolved, Postmortem, Closed.
- Set acknowledgement targets: page acknowledged within five minutes, triage decision within fifteen, declaration within thirty for anything customer-facing.
- Define escalation ladders per layer (responder, service owner, team manager, IC on-call, VP Engineering) each with an automatic timer.
- Make escalation blameless and automatic: no responder is ever criticised for escalating, and timers fire whether or not a human asks.
- Define change freeze and rollback authority during SEV1 and SEV2, and the single condition that lifts the freeze.
- Enforce one incident, one record: the incident record is the sole source of truth for timeline, roles and communications.
- Require every SEV1 and SEV2 to produce an automatically captured timeline from channel and bridge, never one written from memory afterwards.
6. Detection strategy: SLOs, signals and customer-journey monitoring (depends on: 2, 3)
Customers detected 40 percent of incidents first. That number is the reason this step exists.
- Define SLIs and SLOs for the top 20 customer journeys, including payment initiation, settlement, ledger read and write, API availability and webhook delivery, measured per region.
- Require symptom-based alerting on those SLOs rather than cause-based alerting on infrastructure metrics.
- Add synthetic transaction monitoring from outside AWS in both regions and a third location, on a one-minute cadence for money-moving paths.
- Add ledger and shared-PostgreSQL specific signals: replication lag, connection saturation, write latency, transaction ID exhaustion and checkpoint pressure.
- Open a customer-reported path so Support and account managers can raise an incident directly, and count that path as a detection source in reporting.
- Publish a detection contract per service: every one of the 180 services needs an owner, at least one symptom alert and a documented expected detect time.
- Fund a separate resilience track to reduce shared-cluster blast radius, because better detection will not save a single shared ledger during a corruption event.
7. Alert quality standard and noise-reduction programme (depends on: 2, 3, 6)
3,400 alerts a month with 85 percent noise is the reason engineers resent the pager. Fixing it is the price of admission for everything else.
- Publish alert standards: every page must be symptom-based, actionable, owned, linked to a runbook and mapped to a severity. No runbook, no page.
- Separate paging from ticketing: only customer-impacting or imminent-impact signals may page; everything else becomes a ticket or a dashboard entry.
- Set a noise budget per team and per service, for example no service may exceed two pages per on-call shift, measured monthly.
- Define the default action for a noisy alert: fix, tune or delete within ten working days, with deletion treated as a legitimate and successful outcome.
- Introduce correlation and deduplication at the event pipeline so a single root cause produces one page instead of forty.
- Require expiry dates on every silencing rule and temporary threshold so suppression cannot become permanent blindness.
- Run a 90-day noise sprint with a visible burn-down of the top 100 noisiest rules, each owned by a named manager.
- Report page-to-action ratio per team in the monthly reliability review.
8. On-call architecture and 24x7 coverage model across 28 teams (depends on: 3, 4)
This is the hardest political step. The answer to carrying a pager for another team's code is that every team carries its own, and the platform carries the shared risk.
- Adopt a federated model: every service has exactly one owning team, and that team's primary on-call carries its own pager. No team is paged for code it does not own.
- State the consequence honestly: 16 of 28 teams currently have no on-call. They must build one or formally transfer ownership of their services to a team that will.
- Merge small or low-traffic teams into shared rotations so no rotation has fewer than six engineers, below which coverage is not sustainable.
- Create a separate, centrally staffed Incident Commander on-call roster drawn from trained senior engineers across all teams, covering 24x7.
- Define primary and secondary per rotation, with the secondary engaged only on a no-acknowledge or an explicit request.
- Define coverage across the two AWS regions and New York business hours: one global IC rotation, service on-call aligned to their service's users.
- Define the unresponsive-team path: fifteen minutes without acknowledgement escalates to the team manager, thirty minutes to the IC, who may then direct any available engineer.
- Publish a coverage matrix of all 28 teams showing services, rotation size and gaps, reviewed monthly.
- Make on-call participation an explicit expectation in engineering job levels and hiring criteria.
9. On-call compensation, wellbeing and sustainability policy (depends on: 8)
Unpaid on-call is the most cited reason for resistance. The policy must be settled before rollout, not negotiated during it.
- Move to paid on-call: a per-shift stipend or salary uplift agreed with HR and Finance and benchmarked to the New York market.
- Pay event-based compensation for incident callouts outside business hours, with a minimum call-out block.
- Provide compensatory rest: no engineer works a normal day after a night incident, and the rest day is documented, not granted as a favour.
- Cap intrusion by defining a maximum number of off-hours pages per shift, with a mandatory review triggered whenever it is exceeded.
- Define a voluntary opt-out path for engineers with genuine constraints, balanced by an explicit obligation that someone else is paid to take the shift.
- Include on-call expectation and compensation in offers and job descriptions so the commitment is set before hiring.
- Publish the policy with an effective date before any team is asked to join a new rotation.
- Review the policy every six months against actual page volumes, attrition and survey results.
10. Internal and customer communications policy with timing SLAs (depends on: 3, 4)
Today the status page is written by whoever is around. This step replaces improvisation with a clock and a named owner.
- Set internal cadence: first update within 15 minutes of declaration, then every 30 minutes for SEV1 and 60 minutes for SEV2, regardless of whether there is progress.
- Set customer cadence: status page updated within 30 minutes of a SEV1 and 60 minutes of a SEV2, and a no-new-information update is still mandatory.
- Define the channel hierarchy: status page for everyone, direct email to affected customers on SEV1, named account-manager calls for the top 50 accounts.
- Prepare templates per severity in advance with legal and compliance pre-approval, covering detection, impact, workaround, mitigation and next-update time.
- Define regulatory obligations explicitly: money transmitter and banking regulator notification windows, security breach notification, and who signs off (Compliance, not Engineering).
- Prohibit speculation: customer communications never guess at cause or blame and never commit to a root cause before the postmortem.
- Define the closing communication: resolution notice, expected SLA credit handling and the committed date for a written report.
- Assign a named Customer Communications Lead per incident with a trained deputy on every SEV1.
11. Status page, notification tooling and account-manager playbook (depends on: 10)
Policy without tooling collapses at three in the morning. This step makes publishing a five-minute action.
- Upgrade or replace the status page so components map to customer journeys rather than internal services, with subscriber control per component.
- Integrate the incident tool with the status page so the incident record drives the update and the public timeline.
- Provide one-click templates pre-filled with severity, impact language and next-update time.
- Give account managers a playbook: contact tree, what they may say, what they must not say, and how to escalate a customer question into the incident channel within minutes.
- Define the SLA credit process end to end, covering computation, approval, customer notification and finance treatment, so credits stop being a manual scramble.
- Host the status page outside the production failure domain so it survives a total platform outage.
- Test publishing during game days, including a simulated status-page outage and a simulated loss of the primary region.
12. Postmortem policy, template and blameless review process (depends on: 3, 4)
Only 11 of 64 action items closed means the postmortem ritual is currently a writing exercise. This step rebuilds it around learning and tracking.
- Make postmortems mandatory for every SEV1 and SEV2, every SEV3 with customer impact or a repeat pattern, and any near-miss the IC flags.
- Set deadlines: draft within five business days, blameless review within ten, internal publication within fifteen.
- Adopt a single template: impact, timeline, detection, response, contributing factors, what went well, what went badly, action items.
- Train a pool of blameless facilitators and require a trained facilitator for every SEV1 review.
- Prohibit counterfactual and blame language in the template, and require contributing factors across tooling, process, organisation and human factors.
- Limit action items to a small number of concrete, verifiable, owned changes with dates.
- Publish all postmortems internally by default, with security review only for genuinely sensitive material.
- Produce a customer-facing root cause report variant for SEV1 incidents, especially those affecting regulated or top-tier accounts.
13. Corrective action tracking and reliability backlog governance (depends on: 12)
A postmortem without durable action tracking is a complaint, not a control.
- Create a single reliability backlog in the engineering tracker with a mandatory label, owner, due date and link to the originating incident.
- Define closure criteria that require evidence: a merged change, a tested alert or a verified drill, never a self-reported status change.
- Protect capacity by reserving a fixed percentage of each team's sprint for reliability work, with unspent capacity visible to vice presidents.
- Run a weekly ageing review of open actions and escalate anything overdue by more than 30 days to the VP Engineering.
- Report closure rate and median age monthly, targeting more than 90 percent closed within 60 days.
- Require a repeat incident in the same area to trigger a design review rather than another action item.
14. Incident tooling consolidation and integration (depends on: 3, 5, 7, 11)
Six alerting tools and no single incident record is a structural cause of the 22-minute detection and the three-hour mitigation.
- Select one incident management platform for paging, on-call schedules, escalation policies, incident records and postmortem workflow.
- Consolidate the six alerting sources into a single event pipeline feeding that platform, with deduplication and severity mapping applied at ingest.
- Integrate with platform and ledger observability so responders see dashboards and runbooks inside the incident record.
- Integrate chat and bridge: incident channel auto-created, timeline auto-captured, decisions logged as they happen.
- Define the data model and retention required for SOC 2 evidence: who did what, when, and under whose authority.
- Run a dual-run period alongside the old tools with a defined rollback, then switch off the legacy tools on a published date.
- Budget for licences, migration effort and a two-week hardening period after cutover.
15. Training, certification and exercise programme (depends on: 4, 5, 10, 12)
A process that exists only on a wiki page fails on the first real page. Skills have to be built and tested before they are needed.
- Build a short practical curriculum: how to be on call, how to declare an incident, how to run an incident as IC, how to communicate and how to write a postmortem.
- Require certification before joining the IC on-call roster: a written assessment plus a live simulated incident.
- Train at least two certified ICs per team group so the central roster has depth across all 28 teams.
- Run monthly tabletops on realistic scenarios drawn from the last 12 months, including region loss and ledger corruption.
- Run quarterly game days with deliberately injected failure, including shared-PostgreSQL failover and status-page outage.
- Add a mandatory onboarding module for every engineer joining or transferring in, with completion records suitable for audit.
- Track training completion by team and publish it in the monthly reliability review.
16. Metrics, dashboards and review cadence (depends on: 2, 3)
The programme needs a public scoreboard, or it will quietly rot after the audit.
- Define the outcome metrics: time to detect by source, time to mitigate, percentage of incidents detected by customers (target below ten), incidents by severity and SLA credits paid.
- Define the process metrics: declaration latency, page acknowledgement rate, IC roster coverage, first-update timeliness and update-cadence adherence.
- Define the health metrics: alert volume and noise ratio per team, off-hours pages per engineer, postmortem timeliness, action closure rate and action age.
- Publish live dashboards visible to every engineer, not only to managers, refreshed daily.
- Institute a weekly operational review of 30 minutes going incident by incident, a monthly reliability review of trends and actions, and a quarterly executive review with the CEO's office.
- Baseline every metric against the S2 evidence pack and set 90-day and 12-month targets.
- Require every review to end with decisions and owners, not just numbers.
17. Pilot with volunteer teams (depends on: 5, 7, 9, 11, 12, 13, 14, 15, 16)
Do not roll out to 28 teams untested. Run the entire process end to end with a small cohort first.
- Recruit three to four volunteer teams covering a mix of criticality: one payment-path team, one ledger-adjacent team, one platform team and one low-traffic team.
- Run the complete process in the pilot: new severity scale, roles, escalation, communications, postmortems, action tracking and paid on-call.
- Instrument the pilot against the S16 metrics and compare results with the S2 baseline.
- Hold weekly retrospectives with pilot teams and iterate on the written policies, the tooling and the training.
- Fix the top issues found before any wider rollout and document what changed and why.
- Produce a pilot report with before-and-after numbers to carry into every rollout conversation.
- Set explicit pilot exit criteria: rotation coverage achieved, no unacknowledged pages over a defined period, postmortems delivered on time and actions tracked.
18. Phased rollout to all 28 teams (depends on: 13, 16, 17)
Rollout is a staged migration with readiness gates, not an email announcement.
- Sequence the 28 teams into four waves of roughly seven teams, ordered by customer impact, with three weeks between waves.
- Define a per-team readiness checklist: services mapped and owned, alerts cleaned to standard, runbooks written, rotation staffed, training complete and manager briefed.
- Hold a gate review with the process owner before each team joins, and move unready teams to the next wave with a dated remediation plan.
- Give each wave a named champion and run an internal communications cadence that explains the why using pilot numbers.
- Handle resistance directly by publishing the own-your-code-own-your-pager rule and the paid on-call mechanics before each wave, not after.
- Retire legacy tools, escalation lists and the informal status-page process at the end of each wave on a published cutover date.
- Harvest feedback formally at each wave and push accepted changes back into the policy documents through change control.
19. SOC 2 incident-response control mapping and evidence framework (depends on: 1, 3, 10, 12)
The audit will test incident response. Mapping controls now means evidence accumulates from the first real incident under the new process.
- Map the process to the relevant Trust Services Criteria for incident identification, response, evaluation of incidents and communication of security events.
- Write control statements in auditor language and name a single owner for each control.
- Specify the evidence artifact for each control: incident record, severity classification, paging log, communications log, postmortem, action tracker entry and training record.
- Set evidence retention and storage location so nothing depends on a laptop or on chat history that expires.
- Run an early walkthrough with an experienced compliance partner or the auditor's readiness team to test the design before the audit window.
- Flag controls that are blocked until rollout reaches the relevant teams, and keep a gap register with owners and dates reviewed fortnightly by the steering group.
20. SOC 2 dry run, gap remediation and audit support (depends on: 18, 19)
This step converts a good process into a provable one, a few months before the auditors arrive.
- Run an internal dry run six weeks before the audit window, sampling real incidents from the pilot and early waves against each control's evidence requirements.
- Remediate every gap found, prioritising anything that risks a qualified opinion.
- Brief every engineer who may be interviewed so they can describe the process as they actually practise it.
- Prepare the auditor package: process documentation, sample incident records, training records, on-call schedules and the action tracking register.
- Designate a single audit liaison and a small evidence-request team so requests do not land on 28 teams at once.
- Rehearse the walkthrough with the incident commander and communications roles, since auditors probe realism under pressure.
- Schedule the dry run and the audit alongside rollout waves so no team is hit with both in the same week.
21. Standing governance, process ownership and continuous improvement (depends on: 18, 20)
The classic post-audit failure is that the process freezes and then decays. This step keeps it alive.
- Establish a standing Incident Management Council chaired by the process owner, meeting monthly with engineering, support, compliance and product representation.
- Give the process owner a documented mandate to change standards, and require formal change control for any change to severity, roles, communications timings or compensation.
- Re-validate the severity taxonomy quarterly against real declarations and re-baseline the metric set annually.
- Fund a resilience roadmap separate from incident response: reducing shared-database blast radius, multi-region failover readiness and deploy safety.
- Report quarterly to the executive team on the metric set and on the top five systemic causes of incidents.
- Feed incident themes into architecture review, release engineering and hiring so the programme improves the system, not just the response.
- Refresh the training and tabletop programme annually and after any SEV1.
Previous Proposal 3 (ID: bba03d5c-7ae7-4435-9aef-990402088c21, Agent: qwen3.8-flash_initial_3, LLM: alibaba/qwen3.8-flash):
Estimated Complexity: high
Success Metrics: - MTTD reduced to under 10 minutes with >90% detection occurring before customer reporting.
- MTTR for Sev1/Sev2 incidents reduced to under 60 minutes.
- SLA credits paid reduced by 50% within 12 months.
- On-call alert volume reduced by 80% (from 3,400 to <600 actionable alerts/month).
- 100% of Sev1/Sev2 incidents have a blameless postmortem completed within 5 days.
- >90% completion rate of postmortem action items within 30 days of assignment.
- Zero ambiguity in incident command: 100% of major incidents have a designated IC logged in the timeline.
- SOC 2 Type II audit passes incident response controls with zero critical findings.
Steps (12):
1. Establish Governance and Executive Sponsorship
Secure explicit CEO and CTO backing to mandate the program, framing it as a revenue protection and compliance necessity rather than an engineering burden. Form a cross-functional Incident Management Task Force.
- Identify a single executive owner (e.g., VP of Engineering or Head of Platform) with authority over budget and staffing.
- Create a charter defining the scope: severity definitions, on-call compensation, and tool consolidation.
- Communicate the strategic priority to all 28 team leads to preempt pushback regarding pager responsibility.
2. Define Severity Matrix and Response SLAs (depends on: 1)
Create a rigorous, binary severity scale to eliminate ambiguity in escalation and communication speed. Align internal response times with the 99.95% SLA.
- **Sev1 (Critical)**: System-wide outage, data corruption, or >5% transaction failure rate. Immediate executive page.
- **Sev2 (Major)**: Single customer segment impacted or >1% transaction failure rate. High-priority team response.
- **Sev3 (Minor)**: Degraded performance or isolated bug. Standard business-hours follow-up.
- **Sev4 (Cosmetic)**: No functional impact. Backlog item.
- Define maximum detection and mitigation times for each level.
3. Restructure On-Call Organization and Compensation (depends on: 1)
Transform the current unpaid, fragmented on-call system into a professional, compensated, and centralized rotation to ensure 24x7 coverage across all 28 teams. Address the 'pushback' by making it mandatory and paid.
- Consolidate the 12 current on-call teams into a unified, rotating pool covering all 28 teams.
- Implement a fair compensation model: base stipend + hourly rate for active hours + incident bonus for Sev1/2.
- Define a strict 'Buddy System' or 'Follow-the-Sun' rotation if geographically feasible, otherwise ensure adequate team size to rotate fairly.
- Mandate that every team must participate in the rotation; no exemptions.
4. Define Incident Roles and RACI Model (depends on: 2)
Eliminate the 'who is in charge' ambiguity by pre-defining roles and establishing a clear chain of command for every incident. Roles are assigned per incident, not permanently held.
- **Incident Commander (IC)**: Single point of authority. Owns the timeline, delegates tasks, and approves communications. Must be senior.
- **Scribe**: Maintains the real-time timeline log of actions, timestamps, and decisions.
- **Communications Lead**: Owns external/status page updates and internal Slack broadcasts. Shields the IC from interruptions.
- **Technical Leads (SMEs)**: Deployed by IC to troubleshoot specific services. They do *not* coordinate with each other; they report to IC.
- Create a 'Rapid Response Team' (RRT) for high-severity incidents involving the shared PostgreSQL cluster.
5. Consolidate Monitoring and Enforce Alert Hygiene (depends on: 3)
Reduce the 3,400 monthly alerts (85% noise) to a signal-to-noise ratio that allows engineers to trust the pager. Standardize tooling.
- Consolidate the six disparate alerting tools into a single platform (e.g., PagerDuty, Opsgenie, or Datadog).
- Implement 'Actionable Alerting': If an alert pages, there must be a documented, immediate runbook action. No 'informational' pages.
- Set a hard cap on alert volume per engineer/night. Breaching the cap triggers a mandatory alert quality review.
- Deprecate legacy tools that cannot map to the severity matrix defined in S2.
6. Build Detection and Escalation Automation (depends on: 5)
Reduce the 22-minute mean time to detect (MTTD) by automating the path from metric breach to page. Ensure customers never find an issue before the internal pager does.
- Implement synthetic transactions for core payment flows to detect failures before real users do.
- Configure escalation policies: If not acknowledged in 5 minutes, auto-escalate to the team lead, then to the Incident Commander pool.
- Map specific Kubernetes and PostgreSQL metrics to the Severity Matrix (S2) to trigger automated pages.
- Create a 'Major Incident Channel' in Slack that is automatically opened and pinned when a Sev1/2 is triggered.
7. Standardize Communication Templates and Timing (depends on: 4, 6)
Codify the rules for external and internal communications to stop the 'whoever is around' approach. Link these timings to the Severity Matrix.
- **Status Page Updates**:
- Sev1: First update within 5 minutes of detection. Updates every 15 minutes.
- Sev2: First update within 30 minutes. Updates every 60 minutes.
- **Customer Direct Mail**: For Sev1 impacting specific accounts, Account Managers send pre-approved templates within 1 hour of mitigation.
- **Regulator Notification**: Define legal triggers for NY financial regulations.
- Create a 'Comms War Room' Slack channel restricted to the IC and Comms Lead.
8. Formalize the Postmortem Process (depends on: 4, 7)
Transform postmortems from optional/various formats into a mandatory, blameless, and tracked process. This directly addresses the SOC 2 audit requirement for continuous improvement.
- Mandate a postmortem for every Sev1 and Sev2 incident within 5 business days.
- Enforce a 'Blameless' charter: Focus on systemic/process failures, not individual errors.
- Adopt a standardized template: Timeline, Root Cause Analysis (5 Whys or Fishbone), Impact Analysis, and Action Items.
- Action Items must have: Owner, Due Date, and Priority. They are tracked in Jira/Linear, not the doc.
9. Implement Action Item Tracking and Governance (depends on: 8)
Solve the '11 of 64 closed' problem by integrating postmortem actions into the core engineering workflow and executive reviews. Ensure fixes are shipped.
- Integrate Jira with the Postmortem tool (e.g., OpsLevel, Rootly, or Notion API). Auto-create tickets for every AI.
- Block 'Sev1 Closure' until all high-priority Action Items are closed or explicitly deferred with executive approval.
- Weekly Engineering Leadership review of open postmortem actions older than 30 days.
- Tie action item completion rates to team performance reviews.
10. Deploy Training and Runbook Library (depends on: 6, 7, 8)
Address the skill gap and the 'pushback' by equipping engineers with the tools and knowledge to succeed on-call. Reduce anxiety through clarity.
- Audit all 180 services. Require every service to have a 'Page-Worthy' runbook.
- Run 'Game Day' (Chaos Engineering) exercises quarterly to practice role assignments (IC, Scribe, Comms) in a simulated Sev1.
- Conduct mandatory training for the new on-call rotation and compensation model.
- Provide a 'Playbook' for the Incident Commander role, including scripts for communication.
11. Phased Rollout and Change Management (depends on: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10)
Introduce the new process in phases to manage cultural resistance and technical debt, ensuring adoption before the SOC 2 audit. Do not 'big bang' the change.
- **Month 1-2**: Launch Severity Matrix and Consolidated Tooling (S2, S5). Pilot on 5 high-traffic teams.
- **Month 3-4**: Roll out new On-Call Compensation and Roles (S3, S4) to all 28 teams. Enforce new Communication Templates.
- **Month 5-6**: Enforce strict Postmortem and Action Item Tracking (S8, S9).
- **Month 7**: Final audit readiness check. Simulate SOC 2 auditor questions on incident response.
12. Define Metrics for Success and Continuous Review (depends on: 2, 6, 8, 9)
Establish a dashboard to prove the new system works, linking directly to the $1.3M savings in SLA credits. Monitor the health of the incident process itself.
- Track MTTD (Target: <5 min), MTTR (Target: <30 min for Sev2), and Detection Ratio (Target: >95% internal detection).
- Monitor 'Alert Noise' (target: <5% of alerts result in a non-actionable page).
- Track 'Action Item Completion Rate' (Target: 100% on-time closure for Sev1/2).
- Calculate 'SLA Credit Avoidance' vs. On-Call Program Cost to prove ROI to leadership.
Please, considering the previous proposals as ideas that could be considered, focus on the main objective and generate an IMPROVED proposal or a completely DIFFERENT perspective if you deem it appropriate. Only if you consider any of them is amazing and impossible to improve, answer with the same proposal.
The plan has these parts:
- "steps": the list of steps, each with "step_id" (a short label unique within this proposal, such as S1, S2, S3), "title" (a short, descriptive title), "description" (what the step does and how, in Markdown: a short opening paragraph and then, if there are several concrete things to say, a bullet list) and "dependencies" (the step_ids of the other steps of THIS proposal that must be completed first; empty if none).
- "estimated_complexity": "low", "medium" or "high".
- "success_metrics": clear and measurable success metrics, one per line as a bullet list.
Think deeply about the problem, considering all aspects and implications. Develop a comprehensive understanding of what needs to be done to achieve the objective, then produce the plan.
Answer with a single JSON object, and nothing else, that conforms to this JSON schema:
{"description": "The structure every proposing agent must answer with, in every round.", "properties": {"steps": {"items": {"description": "One step of a proposed plan. Dependencies are step_ids of the same proposal.", "properties": {"step_id": {"title": "Step Id", "type": "string"}, "title": {"title": "Title", "type": "string"}, "description": {"title": "Description", "type": "string"}, "dependencies": {"items": {"type": "string"}, "title": "Dependencies", "type": "array"}}, "required": ["step_id", "title", "description", "dependencies"], "title": "Step", "type": "object", "additionalProperties": false}, "title": "Steps", "type": "array"}, "estimated_complexity": {"title": "Estimated Complexity", "type": "string"}, "success_metrics": {"title": "Success Metrics", "type": "string"}}, "required": ["steps", "estimated_complexity", "success_metrics"], "title": "ProposalContent", "type": "object", "additionalProperties": false}
{"steps": [{"step_id": "S1", "title": "Charter, mandate and the evidence clock", "description": "This step turns the CEO's email into a funded programme with one accountable owner and explicit authority, and it starts the SOC 2 clock on day one.\n\n- Appoint a single accountable process owner — a Director of Incident Management reporting to the CTO, with a dotted line to the COO for customer and SLA matters.\n- Publish a one-page charter: scope (every customer-impacting, money-moving, security and data-integrity event across all 28 teams and both regions), decision rights during an active incident, and the power to freeze deploys and override team preferences.\n- Fix the funding envelope up front: tooling licences, training and drill time, and on-call compensation, with an indicative annual figure and the expected return in avoided SLA credits.\n- **Start the evidence clock immediately.** A SOC 2 Type II report covers a period of operating effectiveness, so evidence cannot be backfilled; an audit in eight months means the operating period begins in the first weeks of this programme.\n- Stand up a steering group of CTO, VP Engineering, Head of Support, Head of Compliance and two engineering managers, meeting fortnightly.\n- Make incident-process participation a documented performance expectation for every engineering manager, not an optional extra.\n- Agree the timeline explicitly: a working process in four months, evidence accumulating from month two, audit-ready by month seven.", "dependencies": []}, {"step_id": "S2", "title": "Baseline evidence and problem statement", "description": "You cannot demonstrate improvement without a defensible starting point, and you cannot win the argument about noise without numbers.\n\n- Build a 12-month incident register and re-classify all 31 customer-impacting incidents: date, severity, class, detection source, time to detect, time to mitigate, customers affected, services involved and SLA credits paid.\n- Run an alert census per tool, per team and per service: total volume, page-to-action ratio, off-hours interruptions per engineer, and the 50 noisiest rules with a named owner.\n- Build a silent-failure register: incidents in which no internal alert fired at all. This is the number that explains the 40% customer-detected rate.\n- Survey every on-call engineer and manager on burden, fairness, escalation clarity and pay expectations, targeting above 70% response.\n- Interview Support, Account Management and Sales about how customers learn of incidents, what they complain about, and what they are promised.\n- Reconstruct the two command-ambiguity incidents minute by minute, from alert to mitigation, to find exactly where ownership lapsed.\n- Publish the pack internally as the problem statement and retain it as management-review evidence for the audit.", "dependencies": ["S1"]}, {"step_id": "S3", "title": "Control mapping and evidence architecture", "description": "Most programmes leave compliance to the end. This one maps controls in the first month, because the mapping determines what the process must capture from day one.\n\n- Map the process to the relevant Trust Services Criteria for incident identification, evaluation, response, containment, recovery and communication of events (CC7.1 to CC7.5, CC2.2, CC2.3, CC4.1, CC3.x).\n- Write each control as a plain-language statement with one named owner and its evidence artifact: incident record, severity classification, paging log, communications log, postmortem, action tracker entry, training record.\n- Define the **golden incident file**: one auto-assembled, single-click export per incident containing timeline, roles, severity, communications, postmortem, actions and closure evidence.\n- Set retention, storage location and immutability so no control depends on a laptop, a private Slack channel or chat history that expires.\n- Keep a gap register with owners and dates, reviewed fortnightly by the steering group.\n- Run an early design walkthrough with the auditor's readiness team inside the first 90 days, to test the design before building on it.", "dependencies": ["S1"]}, {"step_id": "S4", "title": "Severity times class taxonomy", "description": "Severity alone is not enough. Severity decides how loud the response is; class decides which playbook runs and what must never be traded away.\n\n- Define severity by impact in customer and money terms: customers affected, dollars at risk, regions lost, ledger integrity, settlement delay, data loss and SLA exposure.\n- Use SEV-0 for security, privacy and regulatory events; SEV-1 for total or material loss of a payment path; SEV-2 for degradation or single-region loss; SEV-3 for limited impact with a workaround; SEV-4 for internal-only issues and near-misses.\n- Define response classes: Availability, Performance, Data integrity and Ledger, Security and Privacy, Third-party dependency, Settlement and Financial, and Process failure.\n- **Class can raise a response, never lower it.** A SEV-2 data-integrity incident gets SEV-1 posture, because integrity failures are not recoverable by moving faster.\n- Specify automatic triggers: loss of one AWS region, ledger write failures, replication lag above threshold, a missed settlement cut-off, payment success rate below threshold for five minutes.\n- State who may declare: any engineer, Support agent or account manager. State who may downgrade: the Incident Commander alone.\n- Map every level to its SLA credit exposure and to the customer-visible status page state.\n- Include worked examples from the last 12 months so each of the 28 teams recognises its own incidents, and re-validate the taxonomy quarterly against real declarations.", "dependencies": ["S2"]}, {"step_id": "S5", "title": "Roles, command and the never-without-an-owner rule", "description": "The two incidents where nobody was in charge for over an hour did not fail at declaration. They failed in the gap before it, when an alert had fired and no one owned it.\n\n- Create one-page role cards: Triage Owner, Incident Commander, Deputy IC, Operations Lead, Internal Communications Lead, Customer Communications Lead, Scribe, Subject-Matter Responders and Executive Sponsor for SEV-1 only.\n- Introduce the **Triage Owner** rule: from the moment a page is acknowledged, that person owns the incident until an IC takes over or the incident is stood down. There is never an unowned minute between first page and close.\n- State the core rule plainly: the IC owns the incident, not the fix, and does not debug. An IC who starts troubleshooting has abandoned command.\n- Give the IC explicit decision rights: declare and escalate severity, freeze changes, halt deploys, approve customer messaging, and pull in any engineer in the company.\n- Define minimum viable staffing per severity: SEV-1 staffs every role; SEV-2 staffs IC, scribe, comms and responders; SEV-3 staffs an IC and a scribe only.\n- Set handover discipline: four-hour maximum IC shifts on SEV-1, a written handover template, and a deputy named within 15 minutes of declaration.\n- Set behaviour standards for responders: one incident channel, one bridge, no side channels, and asks phrased with a named owner and a time.\n- Publish the role cards on the internal wiki and link them from every paging notification.", "dependencies": ["S4"]}, {"step_id": "S6", "title": "Declaration, lifecycle and escalation policy", "description": "This step defines the mechanical path from an alert to a declared incident and back to normal service, and it removes judgment calls from the worst moments.\n\n- Define lifecycle states with entry and exit criteria: Detected, Triaged, Declared, Mitigated, Resolved, Postmortem, Closed — plus a Watch state with a hard 30-minute timer, after which the incident is either declared or stood down.\n- **Make declaring free.** A declaration that turns out to be a false alarm is closed as a false declaration, with no blame and no follow-up, and it is tracked as a metric so the cost of caution stays visible.\n- Apply the ambiguity rule: if two responders disagree about whether this is an incident, it is an incident.\n- Set acknowledgement and declaration targets: page acknowledged within five minutes, triage decision within fifteen, declaration within thirty for anything customer-facing.\n- Build automatic escalation ladders with timers at every layer: responder, service owner, team manager, IC on-call, VP Engineering. Escalation never requires a human decision and is never criticised.\n- Define the unresponsive-team path: fifteen minutes without acknowledgement escalates to the team manager, thirty minutes to the IC, who may then direct any available engineer.\n- Define change freeze and rollback authority during SEV-1 and SEV-2, and the single named condition that lifts the freeze.\n- Enforce one incident, one record, with the timeline auto-captured from the channel and bridge rather than written from memory afterwards.\n- Add the two-simultaneous-incidents rule: a second IC is designated, and no IC ever runs two incidents at once.", "dependencies": ["S5"]}, {"step_id": "S7", "title": "Detection strategy: journeys, synthetic signals and customer-report intake", "description": "Customers detected 40% of incidents first. That number is the reason this step exists, and it is fixed by measuring customer journeys rather than infrastructure.\n\n- Define SLIs and SLOs for the top 20 customer journeys — payment initiation, authorisation, settlement, ledger read and write, API availability, webhook delivery and payout — measured per region.\n- Require symptom-based alerting on those SLOs instead of cause-based alerting on infrastructure metrics.\n- Add synthetic transaction monitoring from outside AWS in both regions plus a third location, on a one-minute cadence for money-moving paths.\n- Add ledger and shared-PostgreSQL specific signals: replication lag, connection saturation, write latency, lock waits, transaction ID exhaustion, checkpoint pressure and bloat.\n- Open a customer-report intake: Support and account managers can raise an incident directly, every customer report creates an incident record, and the customer-report path is counted as a detection source in all reporting.\n- Apply the detection-gap rule: whenever a customer reports an incident before internal monitoring, a detection-gap ticket is opened automatically and owned by the responsible team.\n- Publish a detection contract per service: every one of the 180 services needs an owner, at least one symptom alert, and a documented expected detect time.\n- Measure current detect time per journey, set targets, and run a detection drill per team: break something in staging and see whether it pages before a human notices.", "dependencies": ["S4"]}, {"step_id": "S8", "title": "Three on-call rotations across 28 teams", "description": "The objection is that engineers will not carry a pager for another team's code. The answer is not to argue with it, but to build three rotations so the objection becomes structurally impossible.\n\n- Run a **Service On-Call** rotation per team, covering only that team's own services. No engineer is ever paged for code their team does not own.\n- Run a **Platform Duty** rotation for genuinely shared infrastructure: the shared PostgreSQL cluster, Kubernetes, networking, CI/CD and observability. This is nobody's product code, so it gets its own paid rotation, staffed from platform teams plus volunteers from other teams.\n- Run a central **Incident Commander** roster of 12 to 16 certified senior engineers drawn from across all 28 teams, covering 24x7 on one-week shifts with a primary and a secondary.\n- State the consequence honestly: 16 of 28 teams have no on-call today. Each must either build a rotation or formally transfer ownership of its services to a team that will, with the transfer dated and recorded.\n- Merge small or low-traffic teams into shared rotations so no rotation has fewer than six engineers, the level below which coverage stops being sustainable.\n- Cap load in the scheduling tool: no engineer is on-call more than two weeks per quarter, enforced by configuration rather than negotiation.\n- Publish a coverage matrix for all 28 teams showing services, owners, rotation size and gaps, reviewed monthly.\n- Write on-call participation into engineering job levels and hiring criteria, so the commitment is set before someone joins rather than after.", "dependencies": ["S5"]}, {"step_id": "S9", "title": "Compensation, rest and the price of opting out", "description": "Unpaid on-call is the most cited reason for resistance. The policy must be settled before rollout, not negotiated during it.\n\n- Move immediately to paid on-call: a per-shift stipend benchmarked to the New York market, agreed with HR and Finance, and published with an effective date before any team is asked to join a new rotation.\n- Pay event-based compensation for out-of-hours callouts, with a minimum call-out block and a 1.5x rate for time actually spent mitigating.\n- Provide documented compensatory rest: no engineer works a normal day after a night incident, and the rest day is policy rather than a favour granted by a manager.\n- Cap intrusion with a maximum number of off-hours pages per shift, and trigger a mandatory review whenever the cap is exceeded.\n- Allow opt-out but **put a price on it**: an engineer may step out of a rotation, and their team must buy coverage from the paid pool at a published internal rate. This turns a cultural argument into a visible budget decision.\n- Publish an explicit amnesty: incident records, near-miss reports and false declarations are never used in performance reviews. Only failure to report is a performance issue.\n- Protect against gaming: pay attaches to shifts and callouts, never to incident counts.\n- Review the policy every six months against real page volumes, attrition and survey results.", "dependencies": ["S8"]}, {"step_id": "S10", "title": "Paging contract and alert quality", "description": "3,400 alerts a month at 85% noise is why engineers resent the pager. Fixing that is the price of admission for everything else in this plan.\n\n- Publish the paging contract: every page must be symptom-based, actionable, owned, mapped to a severity and a class, and linked to a runbook. **No runbook, no page**, enforced by a CI check on the alert definition.\n- Separate paging from ticketing: only customer-impacting or imminent-impact signals may wake a human; everything else becomes a ticket or a dashboard entry.\n- Set a page budget per team and per service: a maximum number of pages per on-call shift. Breaching it auto-opens a remediation ticket with the engineering manager as owner.\n- Put new alerts on two-week probation: a new rule runs as a ticket only and becomes a pager only after it has proved actionable, so teams stop being woken by untested rules.\n- Define the default action for a noisy alert: fix, tune or delete within ten working days, with deletion treated as a legitimate and celebrated outcome and the deleted-alert count published.\n- Correlate and deduplicate at the ingest pipeline so one root cause produces one page instead of forty.\n- Require an expiry date on every silencing rule and temporary threshold, so suppression cannot quietly become permanent blindness.\n- Run a 90-day noise sprint with a public burn-down of the top 100 noisiest rules, each owned by a named manager, and report page-to-action ratio per team monthly.", "dependencies": ["S2", "S7"]}, {"step_id": "S11", "title": "Incident tooling consolidation", "description": "Six alerting tools and no single incident record is a structural cause of the 22-minute detection and the three-hour mitigation.\n\n- Select one incident platform for paging, schedules, escalation policies, incident records, postmortem workflow and status-page integration, and choose it within four weeks.\n- Consolidate the six alert sources into a single event pipeline feeding that platform, with deduplication, correlation and severity and class mapping applied at ingest.\n- Integrate observability, runbooks, chat and voice bridges so responders see dashboards and runbooks inside the incident record and the timeline is captured automatically.\n- Implement the golden incident file export defined in S3, so audit evidence is a single click rather than a reconstruction.\n- Run a dual-run period alongside the legacy tools with a defined rollback, then switch the legacy tools off on a published date.\n- Host the status page outside the production failure domain so it survives a total platform outage, and test that during a game day.\n- Make the platform usable from a phone: an Incident Commander must be able to run a SEV-1 from a mobile device at 3am.", "dependencies": ["S5", "S10"]}, {"step_id": "S12", "title": "Communications: internal, customer and regulator", "description": "Today the status page is written by whoever is around. This step replaces improvisation with a clock, a named owner and a pre-cleared template.\n\n- Set the internal cadence: first update within 15 minutes of declaration, then every 30 minutes for SEV-1 and hourly for SEV-2, whether or not there is progress.\n- Never let the status page be how an employee learns of an incident: internal communication leads, external follows.\n- Set the customer cadence: status page updated within 30 minutes of a SEV-1 and 60 minutes of a SEV-2, with a no-new-information update still mandatory.\n- Define the channel hierarchy: status page for all 2,100 customers, direct email to affected customers on SEV-1, and named account-manager calls for the top 50 accounts.\n- Prepare templates per severity and class in advance, pre-approved by Legal and Compliance, each with the next-update time built in.\n- Forbid speculation: customer communications never guess at cause, never assign blame, and never commit to a root cause before the postmortem.\n- Define the closing communication: resolution notice, how SLA credits will be handled, and the committed date for the written report.\n- Build a **regulator clock matrix**: for each event type, which regulator, which window, who signs off, and the shortest clock that drives the first action. Cover NYDFS Part 500, money-transmitter and banking notification, security breach notification, card-network rules, and public-company disclosure where applicable.\n- Route every regulatory notification through Compliance, never Engineering, and pre-clear the templates.\n- Publish a customer-facing root-cause report for SEV-1 incidents, especially for regulated and top-tier accounts.\n- Assign a named Customer Communications Lead plus a trained deputy on every SEV-1.", "dependencies": ["S4", "S5"]}, {"step_id": "S13", "title": "Customer-impact ledger and SLA credit automation", "description": "The $1.3M in credits is a symptom of having no single record of customer impact. This step creates one, and makes it do four jobs at once.\n\n- Maintain one durable customer-impact ledger per incident: which customers, which journeys, from when to when, and the estimated credit.\n- Use that single record for customer communications, SLA credit computation, regulatory reporting and the annual review, so nobody reconciles two versions of the same outage.\n- Automate the credit path end to end — computation, approval, customer notification and finance treatment — so credits stop being a manual scramble three weeks later.\n- Reconcile accrued against paid credits monthly and report the result in the executive review.\n- Track credit avoidance against programme cost, so the funding case is a number rather than an argument.", "dependencies": ["S12"]}, {"step_id": "S14", "title": "Postmortem policy with three artifact levels", "description": "Postmortems currently happen for some incidents, in various formats. This step makes them mandatory where they matter and light where they do not.\n\n- Make postmortems mandatory for every SEV-0, SEV-1 and SEV-2, every SEV-3 with customer impact or a repeat pattern, every near-miss the IC flags, and every incident in which the process itself failed.\n- Use three levels so the ritual matches the weight: a lightweight async incident review for SEV-3 and SEV-4, a standard facilitated postmortem for SEV-2, and a full review with an executive sponsor for SEV-1.\n- Set deadlines: draft within five business days, blameless review within ten, internal publication within fifteen.\n- Adopt one template: impact, timeline, detection, response, contributing factors across tooling, process, organisation and human factors, what went well, what went badly, and actions.\n- Train a pool of blameless facilitators and require a trained facilitator for every SEV-1 review.\n- Prohibit counterfactual and blame language in the template, and specifically ban the phrase human error as a root cause — the question is always what made the error possible.\n- Publish all postmortems internally by default, with security review only for genuinely sensitive material, and produce a customer-facing root-cause variant for SEV-1.\n- Treat postmortems as the learning product of the process, not as a compliance artifact.", "dependencies": ["S5"]}, {"step_id": "S15", "title": "Action items: capped, verifiable, with a repeat-incident rule", "description": "Eleven of 64 action items closed is not a tracking problem. It is a generation problem: the process produces more actions than the organisation can absorb.\n\n- **Cap each postmortem at three action items.** Anything beyond three goes into a ranked reliability backlog, not into the postmortem.\n- Require each action to have a named human owner, a due date, and a definition of done that is an artifact: a merged change, a new alert that fired in a drill, or a test that fails without the fix.\n- Prohibit self-reported closure. Closure requires the artifact, signed off by the process owner or the Incident Commander.\n- Keep one reliability backlog in the engineering tracker with a mandatory label and a link back to the originating incident.\n- Reserve a fixed percentage of every team's sprint capacity for reliability work, with unspent capacity visible to the vice presidents.\n- Apply the repeat-incident rule: a second incident with the same contributing factor triggers a design review owned by an architect, not another action item.\n- Run a weekly ageing review of open actions and escalate anything more than 30 days overdue to the VP of Engineering.\n- Target more than 90% of actions closed within 60 days and a median age under 30 days, reported monthly by team.", "dependencies": ["S14"]}, {"step_id": "S16", "title": "Training, certification and drills", "description": "A process that exists only on a wiki page fails on the first real page. Skills have to be built and tested before they are needed.\n\n- Build a short practical curriculum: how to be on call, how to triage and declare, how to run an incident as IC, how to communicate, and how to write a postmortem.\n- Require certification before joining the IC roster: a written assessment plus a live simulated incident.\n- Certify at least two ICs per team group, so the central roster has depth across all 28 teams and no holiday week is uncovered.\n- Run monthly tabletops drawn from the last 12 months of real incidents, including region loss, ledger write failure and a missed settlement cut-off.\n- Run quarterly game days with deliberately injected failure, including shared-PostgreSQL failover, status-page outage and alerting-pipeline outage.\n- **Drill the process's own failure modes**, not just technical ones: IC unreachable, comms lead on PTO, two simultaneous SEV-1s, a paging storm, and a false alarm that burns an hour.\n- Audit the incident process for single points of failure: who is the only person who can do each critical task, and what happens in their holiday week.\n- Add a mandatory onboarding module for every engineer joining or transferring in, with completion records kept in an audit-ready form.\n- Track training completion by team and publish it in the monthly reliability review.", "dependencies": ["S5", "S6", "S12", "S14"]}, {"step_id": "S17", "title": "Metrics and review cadence", "description": "Establish what good looks like, and measure it in a way that makes people report more incidents rather than fewer.\n\n- Define outcome metrics: time to detect by source, time to mitigate by severity and class, percentage of incidents detected by customers, incidents by severity and class, SLA credits paid, and customer-impact minutes.\n- Define process metrics, because the process itself has SLOs: IC assigned within five minutes, page acknowledgement rate, first-update timeliness, update-cadence adherence, postmortem timeliness, action closure rate, action median age and IC roster coverage.\n- Define health metrics: alert volume and noise ratio per team, off-hours pages per engineer, on-call satisfaction, training completion, and percentage of services with a detection contract.\n- **Never publish incident count as a team metric.** It rewards hiding incidents. Publish reporting metrics instead — near-misses filed, detection gaps closed, false declarations made — alongside the outcome metrics.\n- Publish live dashboards visible to every engineer, refreshed daily, with each metric baselined against the S2 evidence pack and given a 90-day and a 12-month target.\n- Run a 30-minute weekly operational review incident by incident, a monthly reliability review of trends and actions, and a quarterly executive review with the CEO's office covering customer impact, credits and the top five systemic causes.\n- Hold a quarterly review of the process itself: what in the process wasted time, what confused responders, and what should be deleted.\n- Require every review to end with decisions and named owners, never with numbers alone.", "dependencies": ["S2", "S4", "S15"]}, {"step_id": "S18", "title": "Pilot with three to four teams, using real incidents", "description": "Do not roll out to 28 teams untested, and do not rely on drills alone when real incidents are available and better.\n\n- Recruit three to four volunteer teams spanning criticality: a payment-path team, a ledger-adjacent team, the platform team that owns the shared PostgreSQL cluster, and a low-traffic team.\n- Run the entire process end to end in the pilot: the new severity and class scale, roles, escalation, communications, postmortems, action tracking and paid on-call.\n- Treat real incidents during the pilot as the primary training material. Hold a retro within 48 hours of each pilot incident, run by the process owner, while the friction is fresh.\n- Instrument the pilot against the S2 baseline and publish before-and-after numbers.\n- Hold weekly retrospectives with the pilot teams and change the written policies, the tooling and the training in response, documenting what changed and why.\n- Set explicit exit criteria: rotation coverage achieved, zero unacknowledged pages over a defined period, postmortems delivered on time, and actions tracked to closure.\n- Produce a pilot report that every later rollout conversation starts from.", "dependencies": ["S6", "S7", "S8", "S9", "S10", "S11", "S12", "S13", "S14", "S15", "S16", "S17"]}, {"step_id": "S19", "title": "Phased rollout sequenced by cost of failure", "description": "Rollout is a staged migration with readiness gates, not an email announcement. The sequencing matters more than the schedule.\n\n- Sequence the 28 teams into four waves of roughly seven, ordered by **incident density and customer-journey ownership**, not by size or alphabet: the teams whose incidents cost the most SLA credits go first, because their improvement is the visible proof.\n- Leave three weeks between waves and define a per-team readiness checklist: services mapped and owned, alerts cleaned to the paging contract, runbooks written, rotation staffed to at least six engineers, training complete, manager briefed, and compensation in effect.\n- Hold a gate review with the process owner before each team joins, and move an unready team to the next wave with a dated remediation plan rather than granting an exception.\n- Give each wave a named champion and run an internal communications cadence that explains the why using the pilot's numbers.\n- Handle resistance directly and early: publish the own-your-code-own-your-pager rule and the paid on-call mechanics before each wave, never after.\n- Hold a retro within 48 hours of each wave's first incident under the new process, and feed accepted changes back through change control.\n- Retire legacy tools, escalation lists and the informal status-page process at the end of each wave on a published cutover date, so two processes never run in parallel for long.\n- Sequence the waves so no team faces both a rollout and the audit dry run in the same week.", "dependencies": ["S15", "S18"]}, {"step_id": "S20", "title": "Audit dry run and evidence review", "description": "This step converts a good process into a provable one, a few months before the auditors arrive.\n\n- Run an internal dry run six weeks before the audit window, sampling real incidents from the pilot and the early waves against each control's evidence requirements.\n- Remediate every gap found, prioritising anything that risks a qualified opinion, and re-test the remediated control against the same sample.\n- Brief every engineer who may be interviewed so they can describe the process as they actually practise it, not as it is written.\n- Prepare the auditor package: process documentation, sample incident records, paging logs, communication logs, training records, on-call schedules and the action tracking register.\n- Designate one audit liaison and a small evidence-request team, so requests do not land on 28 teams simultaneously.\n- Rehearse the walkthrough with an incident commander and a communications lead, because auditors probe realism under pressure, not documents.\n- Keep the audit liaison and the process owner as the same person where possible, so accountability for the control is also accountability for the evidence.", "dependencies": ["S3", "S19"]}, {"step_id": "S21", "title": "Standing governance, resilience roadmap and continuous improvement", "description": "The classic post-audit failure is that the process freezes and then decays. This step keeps it alive and separates incident response from the engineering that prevents incidents.\n\n- Establish a standing Incident Management Council chaired by the process owner, meeting monthly with engineering, support, compliance and product representation.\n- Give the process owner a documented mandate to change standards, and require formal change control for any change to severity, classes, roles, communication timings or compensation.\n- Re-validate the severity and class taxonomy quarterly against real declarations, and re-baseline the whole metric set annually.\n- Fund a **resilience roadmap separate from incident response**: reduce the shared PostgreSQL blast radius, improve multi-region failover readiness, and tighten deploy safety. Better incident management cannot save a single shared ledger during a corruption event.\n- Report quarterly to the executive team on the metric set and the top five systemic causes of incidents, and feed those themes into architecture review, release engineering and hiring.\n- Keep a public backlog of process improvement ideas so any engineer can propose one, and change the process from it every quarter.\n- Share wins publicly each quarter — faster detection, fewer customer-detected incidents, credits avoided — to sustain momentum after the audit.\n- Refresh the training and tabletop programme annually and immediately after any SEV-1.", "dependencies": ["S19", "S20"]}], "estimated_complexity": "high", "success_metrics": "- Median time to detect falls from 22 minutes to under 8 minutes by month 6, and under 5 minutes for money-moving journeys by month 9.\n- Share of customer-impacting incidents first detected by customers falls from 40% to under 10% by month 9.\n- Median time to mitigate for SEV-1 falls from 3 h 10 min to under 60 minutes by month 9.\n- 100% of SEV-1 and SEV-2 incidents have a named Triage Owner within 5 minutes and a named Incident Commander within 15 minutes, from month 4.\n- Zero incidents in which command authority is unclear for more than 15 minutes, measured monthly from month 4.\n- Monthly page volume falls from 3,400 to under 600, with a false-positive rate under 15%, by month 5.\n- No service exceeds two pages per on-call shift for three consecutive months, measured by month 6.\n- All 28 teams have a documented rotation, a service ownership map and at least one trained on-call engineer by month 6.\n- Incident Commander roster holds at least 12 certified ICs covering 24x7 with no uncovered week, from month 4.\n- 100% of on-call shifts are paid under a published policy from month 2.\n- 100% of SEV-0, SEV-1 and SEV-2 postmortems are published internally within 15 business days, from month 5.\n- Postmortem action items closed within 60 days rise from 17% to over 90%, with median age under 30 days, by month 6.\n- At least one incident review or near-miss report is filed per team per quarter, from month 6.\n- Status page first update is published within 30 minutes on at least 95% of SEV-1 incidents, from month 4.\n- Zero missed regulatory notification windows on any incident requiring notification.\n- The customer-impact ledger is complete for 100% of customer-impacting incidents from month 5, with SLA credits reconciled monthly.\n- Annual SLA credits paid fall at least 50% from the $1.3M baseline within 12 months.\n- Customer-impacting incidents fall at least 40% year over year from the 31-incident baseline.\n- On-call satisfaction is at or above 7 out of 10, with zero voluntary attrition attributed to on-call, measured quarterly from month 6.\n- Review cadence is sustained: weekly operational review in at least 90% of weeks, 12 of 12 monthly reliability reviews, 4 of 4 quarterly executive reviews.\n- SOC 2 Type II is passed with zero findings related to incident response.\n- 100% of new engineers complete the incident-response onboarding module within 30 days of joining.\n- 100% of the 180 services have a named owner, a detection contract and at least one symptom-based alert held to the paging contract by month 7."}P3 grew from 12 to 20 steps by adopting P2's skeleton — charter, baseline, detection SLOs, compensation, control mapping, dry run, culture — which fills most of the prompt's requirements it previously skipped. It remains the thinnest plan in mechanics, dropped its own escalation automation, and its dependency order pushes drills and metrics past full rollout.
- Added step 1 charter with a Director of Incident Management and the "Own Your Code, Own Your Pager" principle, plus step 2 baseline register, alert census and immutable evidence store.
- Added step 8 detection strategy with SLIs for transaction success, settlement lag, external synthetics in both regions and error budgets — round 0 had only synthetic transactions buried in an escalation step.
- Step 5 now separates a federated SME rotation from a central IC rotation and adds the "Unowned Service" rule (transfer or decommission), replacing round-0's implausible single unified pool.
- Added step 12 SOC 2 control mapping with "evidence of operation" per control and step 19 dry run sampling 10 incidents and interviewing engineers.
- Added step 18 culture and change management: publicly correcting blameful leadership language and rotating engineers off when call-out thresholds break.
- Lost round-0 step 6's auto-escalation (5-minute no-ack to team lead, then IC pool); the round-1 plan has no escalation ladder, no acknowledgement timers and no unresponsive-team path — a direct miss on the prompt's "escalation paths" and on the hour-long ownership failures.
- Step 16 (game days, comms drills) depends on step 15 (full rollout), so nothing is rehearsed during the pilot; step 17's metrics dashboard likewise lands after rollout although step 14 claims to measure the pilot against baseline.
- Success metrics carry no dates or checkpoints ("MTTD < 10 minutes") and are the shortest list of the three.
- No SLA credit process, no regulator windows (only "define triggers for NY financial regulations"), no reserved reliability capacity and no postmortem-action cap — the three levers the peers use to fix $1.3M in credits and the 17% closure rate.
- Step 11 keeps "SEV-1 cannot be closed until high-priority actions are completed", which conflates incident closure with remediation and will keep incidents nominally open for weeks.
- Proposal 2 : Named process owner with a charter and steering committee, plus a 12-month baseline register and top-50 noisiest rules.
- Proposal 2 : Paging contract (symptom-based, runbook-linked) and a 90-day noise sprint over the top 100 rules.
- Proposal 2 : Federated per-team rotation plus a central IC roster, paid stipend, call-out pay and mandatory rest.
- Proposal 2 : Status page hosted outside production, control mapping to Trust Services Criteria, and an audit dry run with engineer interviews.
- Proposal 1 : Change champions, celebration of wins and a published reliability newsletter.
- Proposal 2 : Escalation, paging and incident lifecycle policy with acknowledgement targets and automatic timers at every layer.
- Proposal 2 : Reserve a fixed percentage of every team's sprint for reliability work.
- Proposal 1 : Run the first drill during the pilot (week 8) before wider rollout.
+ On-Call Architecture & Compensation Policy+ Detection Strategy & SLOs+ Communication Protocols & Templates+ Postmortem Framework (Blameless)+ Action Item Tracking & Governance+ SOC 2 Control Mapping+ Pilot Implementation (Wave 1)+ Full Rollout Strategy (Waves 2-4)+ Simulations & Game Days+ Culture & Change Management+ SOC 2 Dry Run & Evidence Prep+ Continuous Improvement LoopEscalation, paging and incident lifecycle policyDetection strategy: SLOs, signals and customer-journey monitoringOn-call architecture and 24x7 coverage model across 28 teamsOn-call compensation, wellbeing and sustainability policyInternal and customer communications policy with timing SLAsStatus page, notification tooling and account-manager playbookPostmortem policy, template and blameless review processCorrective action tracking and reliability backlog governancePilot with volunteer teamsPhased rollout to all 28 teamsSOC 2 incident-response control mapping and evidence frameworkSOC 2 dry run, gap remediation and audit supportStanding governance, process ownership and continuous improvement
The plan produced
1. Program Charter & Executive Mandate from P2 step 1
Secure explicit authority to mandate the incident management program, transforming it from an engineering initiative into a business-critical operational standard.
- Appoint a Director of Incident Management (DIM) reporting to the CTO, with a dotted line to the Head of Compliance for audit readiness.
- Publish a charter defining scope: all customer-impacting incidents, payment path failures, and data integrity issues.
- Establish a steering committee (CTO, VP Eng, Head of Support, Legal) to resolve cross-team conflicts and approve tooling budgets.
- Define the 'Own Your Code, Own Your Pager' principle as a non-negotiable engineering value, directly addressing the 28-team resistance.
2. Baseline Audit & Evidence Framework (after 1) from P2 step 2
Create a defensible starting point for metrics and a structured evidence repository for the SOC 2 Type II audit.
- Consolidate the last 12 months of incident data into a single register, flagging the 31 customer-impacting incidents.
- Map current alert sources to specific services to quantify the 3,400 monthly alerts and identify the top 50 noisiest rules.
- Conduct a survey of on-call engineers to measure current sentiment and identify specific pain points in the unpaid rotations.
- Designate a secure, immutable storage location for all future incident artifacts (timelines, logs, postmortems) to satisfy audit requirements.
3. Severity Taxonomy & Trigger Matrix (after 1) from P2 step 3
Define unambiguous severity levels tied to specific triggers, communication SLAs, and executive notifications to eliminate classification debates.
- SEV-1 (Critical): System-wide outage, data loss risk, or >5% revenue-impacting transactions. Triggers immediate IC, VP Eng, and legal notification.
- SEV-2 (Major): Significant degradation or region-wide latency. Triggers IC and Team Lead.
- SEV-3 (Minor): Localized issue with workaround. Triggers on-call engineer.
- SEV-4 (Internal): Low-priority noise. Triggers ticket only.
- Map each severity to specific SLA credit exposures to align technical response with financial risk.
4. Incident Roles & Command Structure (after 3) from P2 step 4
Standardize the Incident Command System (ICS) to prevent the 'ambiguity of ownership' that caused previous hour-long delays.
- Incident Commander (IC): Single point of decision authority; does not debug. Required for SEV-1/2.
- Scribe: Maintains the real-time timeline log; mandatory for SEV-1.
- Comms Lead: Owns status page and internal broadcasts; shields IC from external noise.
- SMEs: Technical responders focused solely on diagnosis/mitigation under IC direction.
- Create role cards with clear 'Do' and 'Don't' responsibilities for quick reference during chaos.
5. On-Call Architecture & Compensation Policy (after 4) new
Establish a sustainable, paid 24x7 on-call model that respects ownership boundaries and compensates for disruption.
- Adopt a federated model: each of the 28 teams maintains its own SME rotation for owned services.
- Create a central IC rotation staffed by senior engineers across teams to ensure command coverage.
- Implement a paid policy: base stipend per shift, enhanced pay for call-outs, and mandatory recovery time after SEV-1 night shifts.
- Define the 'Unowned Service' rule: services without a dedicated team rotation must be transferred to an owning team or decommissioned.
6. Tooling Consolidation & Integration (after 2, 5) from P2 step 14
Replace the six fragmented alerting tools with a single Incident Management Platform to ensure a unified source of truth.
- Select a platform (e.g., PagerDuty + Statuspage + Slack/Teams) capable of handling 180 services.
- Integrate Kubernetes and PostgreSQL monitoring to feed alerts into the central platform.
- Automate the creation of incident channels and timelines upon alert acknowledgment.
- Ensure the status page is decoupled from the production environment to remain available during outages.
7. Alert Quality & Noise Reduction Program (after 6) from P2 step 7
Implement strict alert hygiene rules to reduce the 3,400 monthly alerts and restore trust in the pager.
- Rule: All paging alerts must be symptom-based (customer impact) and have a linked runbook.
- Rule: Implement deduplication and rate-limiting at the ingestion layer.
- Sprint: Conduct a 90-day noise sprint to fix, tune, or delete the top 100 noisiest rules.
- Metric: Enforce a 'Page-to-Action' ratio target; if an alert does not result in human action, it must be demoted to ticket-only.
8. Detection Strategy & SLOs (after 3, 6) new
Shift from infrastructure metrics to Service Level Objectives (SLOs) to detect customer-impacting issues faster than clients can.
- Define SLIs for critical payment paths: Transaction Success Rate, API Latency, Settlement Lag.
- Implement synthetic transaction monitoring from external vantage points in both AWS regions.
- Establish error budgets; breach of error budget triggers automatic review and potential feature freezes.
- Integrate customer-reported signals (support tickets, social media mentions) into the detection workflow as low-priority alerts.
9. Communication Protocols & Templates (after 4, 6) new
Define who talks, when, and what they say to stop ad-hoc status updates and protect brand reputation.
- Status Page: SEV-1 initial update within 15 mins; subsequent updates every 30 mins. Use legal-approved templates.
- Internal: IC broadcasts to #exec-leadership for SEV-1 every hour.
- Regulatory: Define triggers for NY financial regulator notifications; assign Legal/Compliance responsibility for filing.
- Client Success: Provide Account Managers with a 'Known Issues' playbook and real-time access to status page APIs for high-value accounts.
10. Postmortem Framework (Blameless) (after 9) new
Standardize the learning process to ensure consistent quality and adherence to the 'Blameless' culture.
- Mandatory: All SEV-1 and SEV-2 incidents require a postmortem within 5 business days.
- Format: Fixed template including Impact, Timeline, Root Cause, Contributing Factors, and Action Items.
- Review: Facilitated by a neutral party (not the IC) to enforce blameless language and prevent finger-pointing.
- Publication: All postmortems published internally on the Wiki with full searchability.
11. Action Item Tracking & Governance (after 10)
Close the loop on incident resolution by enforcing rigorous tracking of corrective actions to fix the 17% completion rate.
- Automatically create Jira tickets for every action item identified in the postmortem.
- Enforcement: SEV-1 incident cannot be marked 'Closed' until high-priority actions are completed or deferred with VP approval.
- Review: Weekly review of overdue actions in the Engineering Leadership standup.
- Metric: Track 'Mean Time to Remediation' for action items as a key health indicator.
12. SOC 2 Control Mapping (after 2, 10) new
Proactively map the new incident processes to SOC 2 Trust Services Criteria to ensure audit readiness.
- Map S4 (Roles), S9 (Comms), and S10 (Postmortems) to Security and Availability criteria.
- Define 'Evidence of Operation' for each control (e.g., automated timeline logs, signed-off postmortems).
- Identify gaps between current state and audit requirements; assign remediation tasks to the DIM.
- Establish a quarterly internal compliance review to test control effectiveness before the Type II audit.
13. Training & Certification Curriculum (after 4, 6) from P2 step 15
Equip all engineers with the skills to operate within the new framework, reducing anxiety and improving response quality.
- Universal Training: 1-hour module on severity levels and tools for all engineers.
- IC Certification: Mandatory workshop and simulation for engineers joining the central IC rotation.
- Runbook Review: Each team must update and validate runbooks for their top 3 critical alerts.
- Onboarding: Include incident response basics in the engineering onboarding checklist.
14. Pilot Implementation (Wave 1) (after 7, 8, 9, 13) new
Deploy the new process to a controlled subset of high-traffic teams to validate assumptions before broad rollout.
- Select 3 teams: Payments Core, Ledger/API, and one Infrastructure team.
- Run the full cycle for 6 weeks: Alerts, IC handover, Comms, Postmortems.
- Collect feedback daily and adjust playbooks and tooling configurations based on real-world friction.
- Measure success using the baseline metrics from S2 (MTTD, Noise, Satisfaction).
15. Full Rollout Strategy (Waves 2-4) (after 14) new
Gradually migrate the remaining 25 teams using lessons learned from the pilot to manage change resistance.
- Wave 2 (Month 3): Deploy to 8 remaining critical customer-facing teams.
- Wave 3 (Month 4): Deploy to internal platform and data teams.
- Wave 4 (Month 5): Deploy to remaining low-traffic teams and legacy services.
- Maintain a 'Help Desk' for new on-call engineers to troubleshoot process questions during their first shift.
16. Simulations & Game Days (after 15) new
Test the resilience of the process and the tools under controlled failure conditions.
- Tabletop Exercises: Quarterly scenario walks-throughs for leadership and ICs (e.g., 'What if the primary region fails?').
- Chaos Engineering: Inject failures in non-production or canary environments to test alert accuracy and runbook validity.
- Communication Drills: Simulate SEV-1 to test the speed of status page updates and internal notification paths.
- Document findings in postmortems and create action items for identified weaknesses.
17. Metrics Dashboard & Executive Review (after 2, 15) from P2 step 16
Establish a continuous feedback loop to monitor the health of the incident management system.
- Build a real-time dashboard tracking MTTD, MTTR, Alert Volume, and SLA Credits.
- Weekly: Operational review of new incidents and action items with the DIM and Team Leads.
- Monthly: Executive review with CTO/VP Eng focusing on trends, noise ratios, and on-call fatigue.
- Publish 'Reliability Reports' to all engineers to celebrate wins and share learnings.
18. Culture & Change Management (after 5, 15)
Address the cultural shift required to move from 'hero mode' to 'process mode' and mitigate pushback.
- Highlight success stories where the new process reduced toil or prevented customer churn.
- Enforce the 'Blameless' mantra by publicly correcting any leadership communication that implies blame.
- Recognize and reward effective ICs and engineers who improve runbooks or alert quality.
- Address 'Pager Fatigue' proactively by monitoring call-out volume and rotating engineers off if thresholds are breached.
19. SOC 2 Dry Run & Evidence Prep (after 12, 17) new
Conduct a mock audit six months out to identify gaps in evidence retention or process execution.
- Sample 10 recent incidents and verify all required artifacts (timeline, logs, postmortem, actions) are present and accessible.
- Interview on-call engineers to ensure they can describe the process and their roles without hesitation.
- Remediate any 'Control Failures' identified during the dry run.
- Prepare the 'Audit Readiness' package for the external auditors.
20. Continuous Improvement Loop (after 17, 19) from P1 step 21
Institutionalize the evolution of the incident process to prevent stagnation.
- Review the Severity Taxonomy quarterly to ensure it reflects current system behavior and customer expectations.
- Update Tooling integrations annually based on vendor roadmaps and new AWS/K8s features.
- Revise Compensation Policy annually based on market data and internal fairness reviews.
- Set annual 'Reliability Goals' for the engineering organization to drive long-term focus.
- Median Time to Detect (MTTD) < 10 minutes.
- Median Time to Mitigate (MTTR) < 45 minutes for SEV-1.
- Customer-detected incidents < 5% of total incidents.
- Monthly alert volume < 400 actionable alerts (90% reduction in noise).
- SLA credit payouts < $100k annually.
- Postmortem action item completion rate > 90% within 30 days.
- 100% of SEV-1 incidents have a designated IC and Scribe.
- On-call engineer satisfaction score > 4.0/5.0.
- Zero critical findings in SOC 2 Type II audit regarding incident response.
[SYSTEM]
You are an expert assistant in complex project planning.
Your task is to generate a detailed and structured action plan to reach the main objective in the most professional and most detailed way, no matter how much work or steps will be needed to perform.
Use your internal reasoning processes to think deeply about the problem and create the most comprehensive plan possible.
Take as much time and space as you need to think through all aspects of the problem.
After your thorough analysis, answer with the plan in the requested structure.
Every text field you write will be read by a busy person, so write for them: short sentences, one idea each, in short paragraphs. Text fields accept Markdown: separate paragraphs with a blank line, use a bullet list when you enumerate things and plain paragraphs when you explain or argue, and bold at most one key phrase per paragraph. A text of more than three sentences must be split into paragraphs; never deliver a single unbroken block of text.
[HUMAN]
Main Objective: "A B2B payments platform in New York (260 engineers in 28 teams, 2,100 customers, $4B processed a month, a 99.95% SLA with service credits) runs about 180 services on Kubernetes across two AWS regions, with a shared PostgreSQL cluster for the ledger. In the last 12 months: 31 customer-impacting incidents, median time to detect 22 minutes (customers detected 40% of them first), median time to mitigate 3 h 10 min, $1.3M paid in SLA credits, two incidents in which nobody was sure who was in charge for over an hour, and a CEO email about "outages we hear about from clients". On-call exists in 12 of the 28 teams, unpaid, fed by alerts from six different tools — 3,400 alerts a month, 85% of them noise; there is no severity scale, status-page updates are written by whoever is around, and postmortems happen for some incidents, in various formats, with action items that are rarely tracked (11 of 64 closed). A SOC 2 Type II audit in eight months will test incident response. Engineers push back against "carrying a pager for other teams' code".
Define the incident management process: severity levels and what each one triggers; roles (incident commander, communications, scribe, subject-matter responders) and how they are staffed 24x7 across 28 teams; the on-call structure, rotations, compensation and the rules for alert quality; detection and escalation paths; internal and customer communications (status page, account managers, regulators when required) with their timings; postmortems (when mandatory, format, blameless review, ownership and tracking of actions); the metrics and reviews that show whether it works; and how it is introduced across the teams without waiting for the audit."
For your consideration and refinement, here are proposals from the previous round:
Previous Proposal 1 (ID: 895aa761-b883-4277-9609-6e76ca669198, Agent: claudeHaiku4.5_initial_1, LLM: anthropic/claude-haiku-4-5):
Estimated Complexity: high
Success Metrics: - Median time to detect reduced from 22 minutes to <8 minutes within 6 months
- Customer-detected incidents drop from 40% to <5% within 6 months
- Median time to mitigate (MTTR) reduced from 3h 10min to <45 minutes for SEV-1 incidents within 6 months
- Annual SLA credits decrease from $1.3M to <$100k within 12 months
- Alert noise reduced from 3,400 per month (85% false positive) to <400 per month (>95% signal) within 3 months
- Zero incidents with command-and-control ambiguity (>1 hour without clear IC) within 2 months
- Postmortem action item completion rate reaches >80% (from 17%) within 4 months
- On-call satisfaction score (survey) reaches >7/10 for on-call engineers within 3 months
- All 28 teams integrated into incident management system with active on-call rotations by week 20
- SOC 2 Type II audit passes incident response controls with no findings 8 months from start
- Incident commander certification: 100% of active ICs trained and drilled within 2 months
- Monthly incident review meeting established and attended by leadership; trends documented
- New incident system integration complete: single alert tool, single dashboard, all 180 services feeding in, <5 min deployment
Steps (21):
1. Define severity levels and decision criteria
Create a four-tier severity framework (SEV-1 through SEV-4) that guides all downstream decisions about response, escalation, and communications.
Each level must specify: customer impact (revenue at risk, customers affected, data loss risk); financial threshold triggering service credits; whether an incident commander is required; response time SLA (e.g., SEV-1 < 5 min notification, SEV-4 < 2 hours); and the go-live decision tree (when to declare and when to resolve).
- SEV-1: Complete service down or critical path broken for >5% of customers; every minute costs money; IC required; 99.99% uptime threatened
- SEV-2: Significant degradation, features unavailable, affecting 1–5% of customers; IC typically required
- SEV-3: Minor impact, limited customer footprint or workaround exists; escalation path but not automatic IC
- SEV-4: Observations or minor issues; alert-driven, no escalation unless pattern emerges
2. Define incident roles and responsibilities (depends on: 1)
Create the organizational roles that operate during an incident: who is in charge, who talks to customers, who writes down what happened, who fixes the system, and how decisions are made under pressure.
Each role must have a single-sentence mission, decision authority, and escalation upward.
- Incident Commander: owns decision-making and timeline; declares severity; resolves conflicts; may or may not be technical
- Deputy IC: shadow to IC and takes over if IC becomes unavailable
- Communications Lead: writes status page, notifies account teams, manages customer perception
- Scribe: records decisions, who did what, key timestamps; not responsible for fixing
- SME Responders: engineers with context on the failing service(s); take IC's direction without debate
3. Design 24x7 on-call rotation structure (depends on: 2)
Build a rotation model that covers all 28 teams with primary and backup on-call engineers every hour across weekdays, evenings, weekends, and holidays; addresses the pager-carrying resistance.
Key design decisions: Is coverage per-team (each team owns its services) or pooled (shared responder pool handles anything)? How many people per rotation? How long are shifts (one week, two weeks)? When can engineers opt out without leaving the team exposed? Which roles are on-call (IC, communications, SME)?
- Recommend: dedicated IC pool (4–6 people in fast rotation) + per-team SME on-call for each team's own services
- Recommend: two-week rotation blocks to reduce handoff friction
- Recommend: one primary, one secondary per slot; secondary handles during primary's escalation
- Provide swaps, blackout dates, and a rule that no engineer is on-call more than 2 weeks per quarter
4. Define on-call compensation and incentives (depends on: 3)
Create a pay model that makes on-call acceptable and rewards engineers who carry the pager; ties compensation to real business risk.
- Base on-call stipend: e.g., $500–1,000 per week while on-call (regardless of incidents)
- Callback pay: 1.5× hourly rate for time spent mitigating incidents during off-hours
- Incident bonus: $50–100 extra per SEV-1 or SEV-2 incident mitigated (recognition)
- Comp time: full business day off after an incident that required >2 hours mitigation during night/weekend
- Annual bonus tie-in: 10–20% bonus multiplier for flawless on-call reviews
- Communicate: position as investment in reliability, not punishment for being online
5. Consolidate alert routing infrastructure
Replace six alert tools with a single ingestion and routing system; stop engineers from being woken by duplicate alerts, and make escalation automated instead of manual.
Evaluate existing tools (likely candidates: PagerDuty, Opsgenie, or Incident.io) or build a lightweight wrapper. The system must: accept alerts from all 180 services; deduplicate and correlate (same outage, different monitoring source); route to correct on-call engineer; expose an API for playbook automation; log every alert for postmortem analysis.
- Choose tool by week 2 of S5
- Migrate alerting endpoints from 6 sources to 1 by week 4
- Set up audit trail and retention
- Ensure mobile app works (on-call engineers need to engage from phone)
6. Build alert quality rules to cut noise (depends on: 1, 5)
Implement rules that automatically suppress the 85% of alerts that are noise (flapping, transient errors, auto-recovered conditions). Target: <400 actionable alerts per month.
Rules to implement: suppress alerts if service auto-recovered within 30 seconds; deduplicate same alert from multiple monitoring sources; suppress alerts for known maintenance windows; group flapping alerts (same service, >5 occurrences in 2 minutes) into a single page to on-call; rate-limit alerts from noisy services (e.g., max 1 alert per 5 minutes per service until silence clears).
- Audit existing 3,400 alerts per month: which are true signals, which are noise
- Tag each alert source with severity level (S1, S2, S3, S4 from S1)
- Create exceptions list: services known to be noisy, require different rules
- Weekly review: alert teams that trigger >50 alerts per week for reduction strategies
7. Implement automated detection and escalation paths (depends on: 1, 5, 6)
Wire the alert system to automatically escalate based on time or severity; removes the need for manual judgment calls during chaos.
Logic: SEV-1 alert arrives → IC notified instantly via phone call + SMS + Slack + mobile; if IC does not acknowledge within 2 minutes, page deputy IC; Communications Lead pinged simultaneously. SEV-2: on-call SME for that service + IC notify via Slack and mobile, escalate to IC's manager if not acknowledged in 10 min. SEV-3/4: on-call SME only, escalate after 30 min.
- Implement in alert routing system (S5)
- Test all paths weekly via synthetic page to on-call
- Track escalation metrics: how many pages reach secondary, how many hit manager
- Adjust timing based on first month of operations
8. Build incident dashboard and status tracking (depends on: 5)
Create a single source of truth during an incident that every responder sees in real time: who is on-call, incident timeline, who said what, current status, next steps.
Dashboard displays: active incidents and their severity; who is the IC and communications lead; timeline of all events (alert fired, IC assigned, customer notified, mitigation started, resolved); Slack channel and mobile notification status; on-call rosters (who is on-call right now for each team); postmortem link as soon as incident closes.
- Integrate with alert tool (S5) to auto-populate incident creation and initial severity
- Push updates to status page and customer account managers automatically
- Log all timeline entries for audit and postmortem completeness
- Mobile-optimized so IC can work from any device
9. Write incident playbooks for each severity (depends on: 1, 2, 3)
Create a one-page (or one-screen) reference for the IC and SMEs during an incident; sequences the steps and removes ambiguity.
Each severity level gets its own playbook: who gets paged (roles, order); first questions to ask (is it real, how big, who knows); what the IC should declare in first message (status page text, account manager notification, regulatory trigger); how long before escalating to executive team; decision rules for going dark vs. continuing to update customers.
- SEV-1 playbook: immediate IC + comms + CTO notification; customer status every 5 minutes
- SEV-2 playbook: IC + comms + tech lead notification; status every 15 minutes
- SEV-3 playbook: on-call SME + comms if customer-visible; status every 30 min or as resolved
- SEV-4 playbook: on-call SME only; update customers only if promised SLA is at risk
- Include decision trees: is this SEV-1 or SEV-2? Is it our code or dependency? Escalate or containment?
10. Define internal communication workflows (depends on: 2, 3)
Specify who informs whom, in what order, via what channel (call, Slack, email) during an incident; prevents gaps like "nobody knew who was in charge for an hour."
Workflow for SEV-1: IC assigned → IC calls CTO/VP Eng and incident channel lead within 1 minute; incident declared in #incidents Slack channel with severity, IC name, service affected; SME on-call for that service joins call automatically; IC pushes updates to #incidents every 5 minutes or when material change occurs. For SEV-2: IC notifies team leads via Slack, updates #incidents every 15 min. Define escalation: if IC is unreachable, deputy IC takes over and announces it.
- Create a phone tree or on-call list accessible to responders
- Set expectations: "If you don't hear from IC in 2 minutes, call them"
- Use a single incident Slack channel per incident (auto-created by incident tool)
- Log all comms in the incident dashboard for postmortem review
11. Design customer communication and status page process (depends on: 1, 2)
Plan when and how to inform customers, account managers, and regulators; ensure 2,100 customers are not learning about outages from Twitter before you tell them.
Rules by severity: SEV-1 detected → status page updated within 3 minutes (even if root cause unknown; post "Investigating"); account managers of affected customers called within 5 minutes; regulatory notification (if payment processing down) queued for approval; customer email within 10 minutes with ETA for next update. SEV-2: status page within 10 min, account managers called within 15 min, email if affecting >10 customers. SEV-3/4: no customer communication unless SLA at risk.
- Empower Communications Lead to update status page without IC approval if delay >3 min
- Prepare templated messages for common scenarios (database failover, data pipeline stuck, service crashed)
- Route regulatory notifications through legal/compliance; don't wait for perfect root cause
- Track customer impact in real time: how many customers affected by severity
12. Establish blameless postmortem process and format (depends on: 1, 2)
Build a systematic way to learn from incidents so the same failure does not happen twice; counter the fear that admitting a mistake leads to being blamed.
Mandatory postmortems: all SEV-1 and SEV-2 incidents, within 48 hours of resolution. Optional but encouraged: SEV-3 if interesting or if >3 of same type in 30 days. Format: what was the user-visible impact and for how long; what was the root cause (not "human error" but the system condition that made error possible); timeline of discovery and response; action items with owner and deadline; blameless tone (focus on process and system design, not individual mistakes).
- Assign a facilitator (not the on-call IC) to run postmortem
- Attendees: IC, comms lead, SMEs involved, team lead, customer success if customer-facing
- Write postmortem in shared doc; make it findable (searchable, linked from incident)
- No discussion of "who screwed up"; only "why did the system allow this to happen"
13. Build action item tracking and accountability (depends on: 12)
Create a system that tracks postmortem action items so they are not forgotten; currently 11 of 64 (17%) are being tracked, leaving 53 unfinished improvements.
System: each postmortem generates action items (e.g., "add monitoring for X," "update runbook for Y," "write test for Z"). Each item gets: clear description, owner (engineer's name), due date (1–4 weeks based on priority), severity (critical = must do before similar incident happens again; important = improve next month; nice-to-have = backlog). Action items live in a dedicated Jira project visible to all teams; owners are accountable (their manager reviews quarterly). Weekly: incident commander reviews open items due that week. Monthly: each team's postmortem items reviewed in their standup.
- Export action items from postmortem document to tracking system automatically
- Require IC to sign off that an action is complete before closing
- Report on completion rate as a metric (target: >80% by month 3)
14. Define incident metrics and KPIs
Establish what "good" looks like; measure so you can improve. Target metrics for 12 months out: mean time to detect 8 minutes (vs. 22 now), customers detect first <5% of incidents (vs. 40%), MTTR 45 minutes for SEV-1 (vs. 190), SLA credits <$100k/year.
Metrics to track: (1) MTTD = time from incident start to first alert/report; disaggregate: external report vs. internal detection. (2) MTTR = time from first report to full mitigation; track by severity and by service. (3) Customer-reported incidents per month (should drop to <2 per month). (4) Alert signal-to-noise ratio (goal: <5% false positive after S6 rules). (5) On-call satisfaction (survey: would you do this again?). (6) Postmortem action completion rate. (7) Incident commander and responder utilization (hours per week per person).
- Dashboard: auto-populated from incident tool, updated daily
- Disaggregate by team and service: which teams have bad MTTR? Which service is most incident-prone?
15. Create review cadence and governance process (depends on: 14)
Establish regular rhythm to inspect the metrics, spot trends, and adjust the process itself; prevent the system from calcifying.
Weekly: incident commander and on-call lead review prior week—number of incidents, any escalations, any communication gaps. Monthly: director-level incident review—trends by service, top causes of incidents, action item status, whether severity classification is working. Quarterly: full leadership review—MTTD, MTTR, customer impact, SLA credit spend, on-call satisfaction score, any systemic changes needed. Annually: audit the entire process for SOC 2 compliance.
- Assign meeting owners: weekly = on-call lead; monthly = director of reliability; quarterly = VP Eng + CFO (SLA cost) + customer success
- Use same data dashboard (S14) for all reviews
- Publish a monthly "incident newsletter" to all engineers: what happened, what we learned, what's improving
16. Prepare SOC 2 Type II audit checklist (depends on: 1, 2, 9, 12, 13, 14, 15)
Document that the incident management system meets the control requirements for a SOC 2 audit; audit happens in 8 months, so this work builds confidence in coverage.
Audit will test: (1) Is there a defined incident response process? (2) Are roles and responsibilities clear? (3) Are incidents logged and tracked? (4) Is root cause analysis performed? (5) Are action items tracked and completed? (6) Is on-call staffing adequate? (7) Are communications timely? (8) Are postmortems documented and blameless? Create a control mapping document that links each SOC 2 requirement to your process (S1–S15). Collect evidence: incident logs, postmortem documents, action item tickets, metrics reports, training records.
- Designate a compliance owner (often a reliability lead or security engineer)
- Run a mock audit at month 6 to identify gaps
- Ensure all postmortems and incidents are retained and searchable for auditor review
17. Develop implementation and rollout plan (depends on: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16)
Create a phased timeline to roll out the incident management system across all 28 teams; avoids big-bang failure and builds credibility.
Recommended structure: Phase 1 (weeks 1–4): build and test infrastructure (S5, S8, alerting); deploy severity levels and roles (S1, S2); pick pilot teams (2–3 high-traffic teams). Phase 2 (weeks 5–12): train pilot teams, run incident drills, refine playbooks based on learning; expand to half of remaining teams. Phase 3 (weeks 13–20): full rollout to all 28 teams; continue drills; track metrics. Phase 4 (weeks 21–28): stabilize, iterate on metrics, prepare for audit.
- Assign a release manager to coordinate across teams
- Create a detailed Gantt chart with swim lanes (infra, process, training, rollout)
- Identify risks: competing priorities, engineers worried about pager burden, tool adoption friction
- Plan stakeholder engagement: weekly updates to eng leadership, monthly town halls for all engineers
18. Build training and documentation (depends on: 17)
Create role-specific education so engineers understand the new system and are confident executing during an incident.
Training tracks: (1) All engineers: 30-min video on severity levels, communication expectations, postmortem format, action items (mandatory); (2) On-call responders: 2-hour workshop on playbooks, alert tool, escalation paths, when to call manager, case studies of real incidents, mobile app walkthrough; (3) Incident commanders: 4-hour IC bootcamp covering leadership under pressure, decision-making, communicating with executives, status page discipline, postmortem facilitation, practiced drills; (4) Communications leads: templates, when to update, how to talk to customers, regulatory notification rules.
- Record videos so async teams can learn on their schedule
- Create runbooks and quick-reference cards for each role (print + digital)
- Pair new on-call engineers with experienced responder for first week
- Require IC certification before anyone joins IC rotation (pass a practical drill)
19. Execute staged rollout across teams (depends on: 18)
Progressively activate the incident management system with feedback loops at each stage; reduces risk of system-wide failure.
Wave 1 (week 6–8): 3–4 pilot teams begin on-call rotations and incident response using new system; capture feedback daily. Wave 2 (week 10–14): 8–10 additional teams, incorporating lessons from Wave 1; ensure diversity of team types (payment processing, monitoring, data pipeline, auth, etc.). Wave 3 (week 15–20): remaining teams; by now, the system is proven and less hand-holding needed.
- Daily retros with Wave 1 teams: what worked, what was confusing, what broke
- Each wave produces a "lessons learned" document that informs the next
- Track adoption metrics: how many incidents reported per team, alert quality, MTTD/MTTR
- Address resistance: engineers who are skeptical of the system, on-call burden, tool friction; assign a "change champion" in each team
20. Run incident response drills and simulations (depends on: 19)
Practice incidents in a controlled setting so responders gain confidence and gaps are found before a real crisis.
Drill 1 (week 8, pilot teams only): simulated database failover; test alert routing, IC assignment, on-call notification, status page update, postmortem. Drill 2 (week 12, all teams): full incident spanning multiple services; test escalation, communications, decision-making. Drill 3 (week 18): chaos drill with intentional confusion—what if IC is unreachable, what if comms lead is on PTO, what if an alert tool goes down? Quarterly thereafter: repeat drills with new on-call engineers.
- Schedule drills at a time when key people are available (not Friday afternoon)
- Run drill with live observers (reliability lead, manager) who provide feedback
- Measure drill success: all steps completed on time, communications clear, postmortem written
- Debrief immediately: what did we learn, what do we change?
21. Establish continuous improvement process (depends on: 15, 20)
Build a feedback loop so the incident management system evolves based on what you learn.
Monthly improvement cycle: in the monthly incident review (S15), leadership identifies one or two friction points from recent incidents (e.g., "alert tool response is slow," "postmortem meetings are boring," "action items keep slipping"). Assign an owner to investigate and propose a fix (change to playbook, tweak to tool config, additional training, process change). Test the fix with pilot team, gather feedback, roll out if successful. Annual review: can we reduce on-call burden, simplify communication, improve MTTD further? Do any teams need special handling (new services, difficult customers)?
- Keep a public backlog of improvement ideas (Jira, shared doc) so any engineer can suggest
- Quarterly all-hands: share wins ("MTTD improved 50%," "on-call satisfaction up," "zero customer-detected incidents") to build momentum
- Link improvements back to metrics (S14): show that changes actually worked
Previous Proposal 2 (ID: bdeac435-cbbd-4c22-bd5c-08a0ab80adec, Agent: deepseek-flash_initial_2, LLM: deepseek/deepseek-flash):
Estimated Complexity: high
Success Metrics: - Median time to detect reduced from 22 minutes to 5 minutes or less by month 6.
- Customer-detected incidents reduced from 40 percent to 10 percent or less of all customer-impacting incidents by month 9.
- Median time to mitigate reduced from 3 h 10 min to 60 minutes or less for SEV1 and SEV2 by month 9.
- Customer-impacting incidents reduced by at least 40 percent year over year from the 31 incident baseline.
- SLA credits paid reduced by at least 50 percent year over year from the $1.3M baseline.
- 100 percent of SEV1 incidents have a named Incident Commander within 5 minutes and a scribe within 15 minutes.
- Zero incidents in which command authority is unclear for more than 15 minutes.
- 28 of 28 teams have a documented on-call rotation, an owning service map and at least one trained on-call engineer by month 6.
- Central Incident Commander roster holds at least 40 certified ICs covering 24x7 with no single point of failure.
- 100 percent of on-call rotations are paid under a published policy by month 5.
- Monthly alert volume reduced from 3,400 to below 700, with a false-positive rate below 20 percent.
- No service exceeds 2 pages per on-call shift, measured monthly for three consecutive months.
- 100 percent of SEV1 and SEV2 postmortems published internally within 15 business days.
- At least 90 percent of postmortem action items closed within 60 days, up from 17 percent (11 of 64).
- Status-page first update published within 30 minutes on at least 95 percent of SEV1 incidents.
- Zero missed regulatory notification windows on any incident requiring notification.
- SOC 2 Type II audit passed with no findings related to incident response.
- Review cadence sustained: weekly operational review in at least 90 percent of weeks, 12 of 12 monthly reliability reviews, 4 of 4 quarterly executive reviews.
- On-call satisfaction at 70 percent or higher on the quarterly survey, with zero on-call-attributed voluntary attrition.
- 100 percent of new engineers complete the incident-response onboarding module within 30 days of joining.
Steps (21):
1. Programme charter, ownership and executive mandate
This step turns the CEO's email into a funded programme with a named owner and explicit authority. Without it, every downstream decision stalls in cross-team negotiation.
- Appoint a single accountable process owner (for example a Director of Incident Management) reporting to the CTO, with a dotted line to the COO for customer and SLA matters.
- Publish a one-page charter covering scope (all customer-impacting and money-moving incidents), decision rights, and the power to override team preferences during an active incident.
- Define the funding envelope: tooling licences, training time, exercise time and on-call compensation, with an indicative annual figure.
- Set the timeline against the SOC 2 date: a working process in four months, evidence accumulating from month two, audit-ready by month seven.
- Stand up a steering group with CTO, VP Engineering, Head of Support, Head of Compliance and one engineering manager per region.
- Agree that incident-process participation is a documented performance expectation for engineering managers, not an optional extra.
2. Baseline measurement and evidence pack (depends on: 1)
You cannot demonstrate improvement without a defensible starting point, and you cannot win the argument about noise without numbers.
- Build a 12-month incident register: date, retro-assigned severity, detection source, time to detect, time to mitigate, customer impact, services involved and SLA credits paid.
- Quantify the alert estate per tool, per team and per service; compute page-to-action ratio, list the 50 noisiest rules and count off-hours interruptions per engineer.
- Survey on-call engineers and managers on burden, fairness and clarity of escalation, targeting a response rate above 70 percent.
- Interview Support, Account Management and Sales about how customers learn of incidents and what they complain about.
- Document exactly where the current process breaks: unclear command in the two known incidents, postmortem action closure at 11 of 64, and ad-hoc status-page authorship.
- Publish the pack internally as the problem statement and retain it as management-review evidence for the SOC 2 audit.
3. Severity taxonomy and trigger matrix (depends on: 1, 2)
Severity is the keystone of the whole process. Every other rule, from paging to communications timing to postmortems, is keyed off it.
- Define four levels plus a special SEV0 for security or regulatory events: SEV1 for total or material loss of a payment path, SEV2 for degradation or single-region loss, SEV3 for limited impact with a workaround, SEV4 for internal-only issues and near-misses.
- Anchor each level in customer and money terms: customers affected, dollars at risk, regions lost, ledger integrity, settlement delay.
- Specify automatic triggers, for example loss of one AWS region, ledger write failures, a missed settlement cut-off, or payment success rate below threshold for five minutes.
- State who may declare each level (any engineer, Support or account manager may declare) and who may only recommend a downgrade (the Incident Commander alone).
- Map each level to SLA credit exposure and to the customer-visible status-page state.
- Include worked examples from the last 12 months so teams recognise their own incidents in the definitions.
- Add a review clause: the taxonomy is re-validated quarterly against real declarations.
4. Incident roles, command structure and decision rights (depends on: 2, 3)
The two incidents where nobody knew who was in charge for over an hour are the direct brief for this step.
- Define roles with one-page role cards: Incident Commander, Deputy IC, Operations Lead, Internal Communications Lead, Customer Communications Lead, Scribe, Subject-Matter Responders and an Executive Sponsor for SEV1 only.
- State the core rule plainly: the IC owns the incident, not the fix, and does not debug.
- Give the IC explicit decision rights: declaring and escalating severity, freezing changes, halting deploys, approving customer messaging and calling additional responders.
- Define minimum viable role coverage per severity: SEV1 staffs every role, SEV3 staffs an IC and a scribe only.
- Define handover discipline: maximum four-hour IC shifts on SEV1, a written handover template, and a Deputy IC nominated within 15 minutes.
- Set behaviour standards for responders: one incident channel, one bridge, no side channels, and explicit asks with a named owner and a time.
- Publish role cards on the internal wiki and link them from every paging notification.
5. Escalation, paging and incident lifecycle policy (depends on: 3, 4)
This step defines the mechanical path from an alert to a declared incident and back to normal service.
- Define lifecycle states with entry and exit criteria: Detected, Triaged, Declared, Mitigated, Resolved, Postmortem, Closed.
- Set acknowledgement targets: page acknowledged within five minutes, triage decision within fifteen, declaration within thirty for anything customer-facing.
- Define escalation ladders per layer (responder, service owner, team manager, IC on-call, VP Engineering) each with an automatic timer.
- Make escalation blameless and automatic: no responder is ever criticised for escalating, and timers fire whether or not a human asks.
- Define change freeze and rollback authority during SEV1 and SEV2, and the single condition that lifts the freeze.
- Enforce one incident, one record: the incident record is the sole source of truth for timeline, roles and communications.
- Require every SEV1 and SEV2 to produce an automatically captured timeline from channel and bridge, never one written from memory afterwards.
6. Detection strategy: SLOs, signals and customer-journey monitoring (depends on: 2, 3)
Customers detected 40 percent of incidents first. That number is the reason this step exists.
- Define SLIs and SLOs for the top 20 customer journeys, including payment initiation, settlement, ledger read and write, API availability and webhook delivery, measured per region.
- Require symptom-based alerting on those SLOs rather than cause-based alerting on infrastructure metrics.
- Add synthetic transaction monitoring from outside AWS in both regions and a third location, on a one-minute cadence for money-moving paths.
- Add ledger and shared-PostgreSQL specific signals: replication lag, connection saturation, write latency, transaction ID exhaustion and checkpoint pressure.
- Open a customer-reported path so Support and account managers can raise an incident directly, and count that path as a detection source in reporting.
- Publish a detection contract per service: every one of the 180 services needs an owner, at least one symptom alert and a documented expected detect time.
- Fund a separate resilience track to reduce shared-cluster blast radius, because better detection will not save a single shared ledger during a corruption event.
7. Alert quality standard and noise-reduction programme (depends on: 2, 3, 6)
3,400 alerts a month with 85 percent noise is the reason engineers resent the pager. Fixing it is the price of admission for everything else.
- Publish alert standards: every page must be symptom-based, actionable, owned, linked to a runbook and mapped to a severity. No runbook, no page.
- Separate paging from ticketing: only customer-impacting or imminent-impact signals may page; everything else becomes a ticket or a dashboard entry.
- Set a noise budget per team and per service, for example no service may exceed two pages per on-call shift, measured monthly.
- Define the default action for a noisy alert: fix, tune or delete within ten working days, with deletion treated as a legitimate and successful outcome.
- Introduce correlation and deduplication at the event pipeline so a single root cause produces one page instead of forty.
- Require expiry dates on every silencing rule and temporary threshold so suppression cannot become permanent blindness.
- Run a 90-day noise sprint with a visible burn-down of the top 100 noisiest rules, each owned by a named manager.
- Report page-to-action ratio per team in the monthly reliability review.
8. On-call architecture and 24x7 coverage model across 28 teams (depends on: 3, 4)
This is the hardest political step. The answer to carrying a pager for another team's code is that every team carries its own, and the platform carries the shared risk.
- Adopt a federated model: every service has exactly one owning team, and that team's primary on-call carries its own pager. No team is paged for code it does not own.
- State the consequence honestly: 16 of 28 teams currently have no on-call. They must build one or formally transfer ownership of their services to a team that will.
- Merge small or low-traffic teams into shared rotations so no rotation has fewer than six engineers, below which coverage is not sustainable.
- Create a separate, centrally staffed Incident Commander on-call roster drawn from trained senior engineers across all teams, covering 24x7.
- Define primary and secondary per rotation, with the secondary engaged only on a no-acknowledge or an explicit request.
- Define coverage across the two AWS regions and New York business hours: one global IC rotation, service on-call aligned to their service's users.
- Define the unresponsive-team path: fifteen minutes without acknowledgement escalates to the team manager, thirty minutes to the IC, who may then direct any available engineer.
- Publish a coverage matrix of all 28 teams showing services, rotation size and gaps, reviewed monthly.
- Make on-call participation an explicit expectation in engineering job levels and hiring criteria.
9. On-call compensation, wellbeing and sustainability policy (depends on: 8)
Unpaid on-call is the most cited reason for resistance. The policy must be settled before rollout, not negotiated during it.
- Move to paid on-call: a per-shift stipend or salary uplift agreed with HR and Finance and benchmarked to the New York market.
- Pay event-based compensation for incident callouts outside business hours, with a minimum call-out block.
- Provide compensatory rest: no engineer works a normal day after a night incident, and the rest day is documented, not granted as a favour.
- Cap intrusion by defining a maximum number of off-hours pages per shift, with a mandatory review triggered whenever it is exceeded.
- Define a voluntary opt-out path for engineers with genuine constraints, balanced by an explicit obligation that someone else is paid to take the shift.
- Include on-call expectation and compensation in offers and job descriptions so the commitment is set before hiring.
- Publish the policy with an effective date before any team is asked to join a new rotation.
- Review the policy every six months against actual page volumes, attrition and survey results.
10. Internal and customer communications policy with timing SLAs (depends on: 3, 4)
Today the status page is written by whoever is around. This step replaces improvisation with a clock and a named owner.
- Set internal cadence: first update within 15 minutes of declaration, then every 30 minutes for SEV1 and 60 minutes for SEV2, regardless of whether there is progress.
- Set customer cadence: status page updated within 30 minutes of a SEV1 and 60 minutes of a SEV2, and a no-new-information update is still mandatory.
- Define the channel hierarchy: status page for everyone, direct email to affected customers on SEV1, named account-manager calls for the top 50 accounts.
- Prepare templates per severity in advance with legal and compliance pre-approval, covering detection, impact, workaround, mitigation and next-update time.
- Define regulatory obligations explicitly: money transmitter and banking regulator notification windows, security breach notification, and who signs off (Compliance, not Engineering).
- Prohibit speculation: customer communications never guess at cause or blame and never commit to a root cause before the postmortem.
- Define the closing communication: resolution notice, expected SLA credit handling and the committed date for a written report.
- Assign a named Customer Communications Lead per incident with a trained deputy on every SEV1.
11. Status page, notification tooling and account-manager playbook (depends on: 10)
Policy without tooling collapses at three in the morning. This step makes publishing a five-minute action.
- Upgrade or replace the status page so components map to customer journeys rather than internal services, with subscriber control per component.
- Integrate the incident tool with the status page so the incident record drives the update and the public timeline.
- Provide one-click templates pre-filled with severity, impact language and next-update time.
- Give account managers a playbook: contact tree, what they may say, what they must not say, and how to escalate a customer question into the incident channel within minutes.
- Define the SLA credit process end to end, covering computation, approval, customer notification and finance treatment, so credits stop being a manual scramble.
- Host the status page outside the production failure domain so it survives a total platform outage.
- Test publishing during game days, including a simulated status-page outage and a simulated loss of the primary region.
12. Postmortem policy, template and blameless review process (depends on: 3, 4)
Only 11 of 64 action items closed means the postmortem ritual is currently a writing exercise. This step rebuilds it around learning and tracking.
- Make postmortems mandatory for every SEV1 and SEV2, every SEV3 with customer impact or a repeat pattern, and any near-miss the IC flags.
- Set deadlines: draft within five business days, blameless review within ten, internal publication within fifteen.
- Adopt a single template: impact, timeline, detection, response, contributing factors, what went well, what went badly, action items.
- Train a pool of blameless facilitators and require a trained facilitator for every SEV1 review.
- Prohibit counterfactual and blame language in the template, and require contributing factors across tooling, process, organisation and human factors.
- Limit action items to a small number of concrete, verifiable, owned changes with dates.
- Publish all postmortems internally by default, with security review only for genuinely sensitive material.
- Produce a customer-facing root cause report variant for SEV1 incidents, especially those affecting regulated or top-tier accounts.
13. Corrective action tracking and reliability backlog governance (depends on: 12)
A postmortem without durable action tracking is a complaint, not a control.
- Create a single reliability backlog in the engineering tracker with a mandatory label, owner, due date and link to the originating incident.
- Define closure criteria that require evidence: a merged change, a tested alert or a verified drill, never a self-reported status change.
- Protect capacity by reserving a fixed percentage of each team's sprint for reliability work, with unspent capacity visible to vice presidents.
- Run a weekly ageing review of open actions and escalate anything overdue by more than 30 days to the VP Engineering.
- Report closure rate and median age monthly, targeting more than 90 percent closed within 60 days.
- Require a repeat incident in the same area to trigger a design review rather than another action item.
14. Incident tooling consolidation and integration (depends on: 3, 5, 7, 11)
Six alerting tools and no single incident record is a structural cause of the 22-minute detection and the three-hour mitigation.
- Select one incident management platform for paging, on-call schedules, escalation policies, incident records and postmortem workflow.
- Consolidate the six alerting sources into a single event pipeline feeding that platform, with deduplication and severity mapping applied at ingest.
- Integrate with platform and ledger observability so responders see dashboards and runbooks inside the incident record.
- Integrate chat and bridge: incident channel auto-created, timeline auto-captured, decisions logged as they happen.
- Define the data model and retention required for SOC 2 evidence: who did what, when, and under whose authority.
- Run a dual-run period alongside the old tools with a defined rollback, then switch off the legacy tools on a published date.
- Budget for licences, migration effort and a two-week hardening period after cutover.
15. Training, certification and exercise programme (depends on: 4, 5, 10, 12)
A process that exists only on a wiki page fails on the first real page. Skills have to be built and tested before they are needed.
- Build a short practical curriculum: how to be on call, how to declare an incident, how to run an incident as IC, how to communicate and how to write a postmortem.
- Require certification before joining the IC on-call roster: a written assessment plus a live simulated incident.
- Train at least two certified ICs per team group so the central roster has depth across all 28 teams.
- Run monthly tabletops on realistic scenarios drawn from the last 12 months, including region loss and ledger corruption.
- Run quarterly game days with deliberately injected failure, including shared-PostgreSQL failover and status-page outage.
- Add a mandatory onboarding module for every engineer joining or transferring in, with completion records suitable for audit.
- Track training completion by team and publish it in the monthly reliability review.
16. Metrics, dashboards and review cadence (depends on: 2, 3)
The programme needs a public scoreboard, or it will quietly rot after the audit.
- Define the outcome metrics: time to detect by source, time to mitigate, percentage of incidents detected by customers (target below ten), incidents by severity and SLA credits paid.
- Define the process metrics: declaration latency, page acknowledgement rate, IC roster coverage, first-update timeliness and update-cadence adherence.
- Define the health metrics: alert volume and noise ratio per team, off-hours pages per engineer, postmortem timeliness, action closure rate and action age.
- Publish live dashboards visible to every engineer, not only to managers, refreshed daily.
- Institute a weekly operational review of 30 minutes going incident by incident, a monthly reliability review of trends and actions, and a quarterly executive review with the CEO's office.
- Baseline every metric against the S2 evidence pack and set 90-day and 12-month targets.
- Require every review to end with decisions and owners, not just numbers.
17. Pilot with volunteer teams (depends on: 5, 7, 9, 11, 12, 13, 14, 15, 16)
Do not roll out to 28 teams untested. Run the entire process end to end with a small cohort first.
- Recruit three to four volunteer teams covering a mix of criticality: one payment-path team, one ledger-adjacent team, one platform team and one low-traffic team.
- Run the complete process in the pilot: new severity scale, roles, escalation, communications, postmortems, action tracking and paid on-call.
- Instrument the pilot against the S16 metrics and compare results with the S2 baseline.
- Hold weekly retrospectives with pilot teams and iterate on the written policies, the tooling and the training.
- Fix the top issues found before any wider rollout and document what changed and why.
- Produce a pilot report with before-and-after numbers to carry into every rollout conversation.
- Set explicit pilot exit criteria: rotation coverage achieved, no unacknowledged pages over a defined period, postmortems delivered on time and actions tracked.
18. Phased rollout to all 28 teams (depends on: 13, 16, 17)
Rollout is a staged migration with readiness gates, not an email announcement.
- Sequence the 28 teams into four waves of roughly seven teams, ordered by customer impact, with three weeks between waves.
- Define a per-team readiness checklist: services mapped and owned, alerts cleaned to standard, runbooks written, rotation staffed, training complete and manager briefed.
- Hold a gate review with the process owner before each team joins, and move unready teams to the next wave with a dated remediation plan.
- Give each wave a named champion and run an internal communications cadence that explains the why using pilot numbers.
- Handle resistance directly by publishing the own-your-code-own-your-pager rule and the paid on-call mechanics before each wave, not after.
- Retire legacy tools, escalation lists and the informal status-page process at the end of each wave on a published cutover date.
- Harvest feedback formally at each wave and push accepted changes back into the policy documents through change control.
19. SOC 2 incident-response control mapping and evidence framework (depends on: 1, 3, 10, 12)
The audit will test incident response. Mapping controls now means evidence accumulates from the first real incident under the new process.
- Map the process to the relevant Trust Services Criteria for incident identification, response, evaluation of incidents and communication of security events.
- Write control statements in auditor language and name a single owner for each control.
- Specify the evidence artifact for each control: incident record, severity classification, paging log, communications log, postmortem, action tracker entry and training record.
- Set evidence retention and storage location so nothing depends on a laptop or on chat history that expires.
- Run an early walkthrough with an experienced compliance partner or the auditor's readiness team to test the design before the audit window.
- Flag controls that are blocked until rollout reaches the relevant teams, and keep a gap register with owners and dates reviewed fortnightly by the steering group.
20. SOC 2 dry run, gap remediation and audit support (depends on: 18, 19)
This step converts a good process into a provable one, a few months before the auditors arrive.
- Run an internal dry run six weeks before the audit window, sampling real incidents from the pilot and early waves against each control's evidence requirements.
- Remediate every gap found, prioritising anything that risks a qualified opinion.
- Brief every engineer who may be interviewed so they can describe the process as they actually practise it.
- Prepare the auditor package: process documentation, sample incident records, training records, on-call schedules and the action tracking register.
- Designate a single audit liaison and a small evidence-request team so requests do not land on 28 teams at once.
- Rehearse the walkthrough with the incident commander and communications roles, since auditors probe realism under pressure.
- Schedule the dry run and the audit alongside rollout waves so no team is hit with both in the same week.
21. Standing governance, process ownership and continuous improvement (depends on: 18, 20)
The classic post-audit failure is that the process freezes and then decays. This step keeps it alive.
- Establish a standing Incident Management Council chaired by the process owner, meeting monthly with engineering, support, compliance and product representation.
- Give the process owner a documented mandate to change standards, and require formal change control for any change to severity, roles, communications timings or compensation.
- Re-validate the severity taxonomy quarterly against real declarations and re-baseline the metric set annually.
- Fund a resilience roadmap separate from incident response: reducing shared-database blast radius, multi-region failover readiness and deploy safety.
- Report quarterly to the executive team on the metric set and on the top five systemic causes of incidents.
- Feed incident themes into architecture review, release engineering and hiring so the programme improves the system, not just the response.
- Refresh the training and tabletop programme annually and after any SEV1.
Previous Proposal 3 (ID: bba03d5c-7ae7-4435-9aef-990402088c21, Agent: qwen3.8-flash_initial_3, LLM: alibaba/qwen3.8-flash):
Estimated Complexity: high
Success Metrics: - MTTD reduced to under 10 minutes with >90% detection occurring before customer reporting.
- MTTR for Sev1/Sev2 incidents reduced to under 60 minutes.
- SLA credits paid reduced by 50% within 12 months.
- On-call alert volume reduced by 80% (from 3,400 to <600 actionable alerts/month).
- 100% of Sev1/Sev2 incidents have a blameless postmortem completed within 5 days.
- >90% completion rate of postmortem action items within 30 days of assignment.
- Zero ambiguity in incident command: 100% of major incidents have a designated IC logged in the timeline.
- SOC 2 Type II audit passes incident response controls with zero critical findings.
Steps (12):
1. Establish Governance and Executive Sponsorship
Secure explicit CEO and CTO backing to mandate the program, framing it as a revenue protection and compliance necessity rather than an engineering burden. Form a cross-functional Incident Management Task Force.
- Identify a single executive owner (e.g., VP of Engineering or Head of Platform) with authority over budget and staffing.
- Create a charter defining the scope: severity definitions, on-call compensation, and tool consolidation.
- Communicate the strategic priority to all 28 team leads to preempt pushback regarding pager responsibility.
2. Define Severity Matrix and Response SLAs (depends on: 1)
Create a rigorous, binary severity scale to eliminate ambiguity in escalation and communication speed. Align internal response times with the 99.95% SLA.
- **Sev1 (Critical)**: System-wide outage, data corruption, or >5% transaction failure rate. Immediate executive page.
- **Sev2 (Major)**: Single customer segment impacted or >1% transaction failure rate. High-priority team response.
- **Sev3 (Minor)**: Degraded performance or isolated bug. Standard business-hours follow-up.
- **Sev4 (Cosmetic)**: No functional impact. Backlog item.
- Define maximum detection and mitigation times for each level.
3. Restructure On-Call Organization and Compensation (depends on: 1)
Transform the current unpaid, fragmented on-call system into a professional, compensated, and centralized rotation to ensure 24x7 coverage across all 28 teams. Address the 'pushback' by making it mandatory and paid.
- Consolidate the 12 current on-call teams into a unified, rotating pool covering all 28 teams.
- Implement a fair compensation model: base stipend + hourly rate for active hours + incident bonus for Sev1/2.
- Define a strict 'Buddy System' or 'Follow-the-Sun' rotation if geographically feasible, otherwise ensure adequate team size to rotate fairly.
- Mandate that every team must participate in the rotation; no exemptions.
4. Define Incident Roles and RACI Model (depends on: 2)
Eliminate the 'who is in charge' ambiguity by pre-defining roles and establishing a clear chain of command for every incident. Roles are assigned per incident, not permanently held.
- **Incident Commander (IC)**: Single point of authority. Owns the timeline, delegates tasks, and approves communications. Must be senior.
- **Scribe**: Maintains the real-time timeline log of actions, timestamps, and decisions.
- **Communications Lead**: Owns external/status page updates and internal Slack broadcasts. Shields the IC from interruptions.
- **Technical Leads (SMEs)**: Deployed by IC to troubleshoot specific services. They do *not* coordinate with each other; they report to IC.
- Create a 'Rapid Response Team' (RRT) for high-severity incidents involving the shared PostgreSQL cluster.
5. Consolidate Monitoring and Enforce Alert Hygiene (depends on: 3)
Reduce the 3,400 monthly alerts (85% noise) to a signal-to-noise ratio that allows engineers to trust the pager. Standardize tooling.
- Consolidate the six disparate alerting tools into a single platform (e.g., PagerDuty, Opsgenie, or Datadog).
- Implement 'Actionable Alerting': If an alert pages, there must be a documented, immediate runbook action. No 'informational' pages.
- Set a hard cap on alert volume per engineer/night. Breaching the cap triggers a mandatory alert quality review.
- Deprecate legacy tools that cannot map to the severity matrix defined in S2.
6. Build Detection and Escalation Automation (depends on: 5)
Reduce the 22-minute mean time to detect (MTTD) by automating the path from metric breach to page. Ensure customers never find an issue before the internal pager does.
- Implement synthetic transactions for core payment flows to detect failures before real users do.
- Configure escalation policies: If not acknowledged in 5 minutes, auto-escalate to the team lead, then to the Incident Commander pool.
- Map specific Kubernetes and PostgreSQL metrics to the Severity Matrix (S2) to trigger automated pages.
- Create a 'Major Incident Channel' in Slack that is automatically opened and pinned when a Sev1/2 is triggered.
7. Standardize Communication Templates and Timing (depends on: 4, 6)
Codify the rules for external and internal communications to stop the 'whoever is around' approach. Link these timings to the Severity Matrix.
- **Status Page Updates**:
- Sev1: First update within 5 minutes of detection. Updates every 15 minutes.
- Sev2: First update within 30 minutes. Updates every 60 minutes.
- **Customer Direct Mail**: For Sev1 impacting specific accounts, Account Managers send pre-approved templates within 1 hour of mitigation.
- **Regulator Notification**: Define legal triggers for NY financial regulations.
- Create a 'Comms War Room' Slack channel restricted to the IC and Comms Lead.
8. Formalize the Postmortem Process (depends on: 4, 7)
Transform postmortems from optional/various formats into a mandatory, blameless, and tracked process. This directly addresses the SOC 2 audit requirement for continuous improvement.
- Mandate a postmortem for every Sev1 and Sev2 incident within 5 business days.
- Enforce a 'Blameless' charter: Focus on systemic/process failures, not individual errors.
- Adopt a standardized template: Timeline, Root Cause Analysis (5 Whys or Fishbone), Impact Analysis, and Action Items.
- Action Items must have: Owner, Due Date, and Priority. They are tracked in Jira/Linear, not the doc.
9. Implement Action Item Tracking and Governance (depends on: 8)
Solve the '11 of 64 closed' problem by integrating postmortem actions into the core engineering workflow and executive reviews. Ensure fixes are shipped.
- Integrate Jira with the Postmortem tool (e.g., OpsLevel, Rootly, or Notion API). Auto-create tickets for every AI.
- Block 'Sev1 Closure' until all high-priority Action Items are closed or explicitly deferred with executive approval.
- Weekly Engineering Leadership review of open postmortem actions older than 30 days.
- Tie action item completion rates to team performance reviews.
10. Deploy Training and Runbook Library (depends on: 6, 7, 8)
Address the skill gap and the 'pushback' by equipping engineers with the tools and knowledge to succeed on-call. Reduce anxiety through clarity.
- Audit all 180 services. Require every service to have a 'Page-Worthy' runbook.
- Run 'Game Day' (Chaos Engineering) exercises quarterly to practice role assignments (IC, Scribe, Comms) in a simulated Sev1.
- Conduct mandatory training for the new on-call rotation and compensation model.
- Provide a 'Playbook' for the Incident Commander role, including scripts for communication.
11. Phased Rollout and Change Management (depends on: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10)
Introduce the new process in phases to manage cultural resistance and technical debt, ensuring adoption before the SOC 2 audit. Do not 'big bang' the change.
- **Month 1-2**: Launch Severity Matrix and Consolidated Tooling (S2, S5). Pilot on 5 high-traffic teams.
- **Month 3-4**: Roll out new On-Call Compensation and Roles (S3, S4) to all 28 teams. Enforce new Communication Templates.
- **Month 5-6**: Enforce strict Postmortem and Action Item Tracking (S8, S9).
- **Month 7**: Final audit readiness check. Simulate SOC 2 auditor questions on incident response.
12. Define Metrics for Success and Continuous Review (depends on: 2, 6, 8, 9)
Establish a dashboard to prove the new system works, linking directly to the $1.3M savings in SLA credits. Monitor the health of the incident process itself.
- Track MTTD (Target: <5 min), MTTR (Target: <30 min for Sev2), and Detection Ratio (Target: >95% internal detection).
- Monitor 'Alert Noise' (target: <5% of alerts result in a non-actionable page).
- Track 'Action Item Completion Rate' (Target: 100% on-time closure for Sev1/2).
- Calculate 'SLA Credit Avoidance' vs. On-Call Program Cost to prove ROI to leadership.
Please, considering the previous proposals as ideas that could be considered, focus on the main objective and generate an IMPROVED proposal or a completely DIFFERENT perspective if you deem it appropriate. Only if you consider any of them is amazing and impossible to improve, answer with the same proposal.
The plan has these parts:
- "steps": the list of steps, each with "step_id" (a short label unique within this proposal, such as S1, S2, S3), "title" (a short, descriptive title), "description" (what the step does and how, in Markdown: a short opening paragraph and then, if there are several concrete things to say, a bullet list) and "dependencies" (the step_ids of the other steps of THIS proposal that must be completed first; empty if none).
- "estimated_complexity": "low", "medium" or "high".
- "success_metrics": clear and measurable success metrics, one per line as a bullet list.
Think deeply about the problem, considering all aspects and implications. Develop a comprehensive understanding of what needs to be done to achieve the objective, then produce the plan.
{"steps": [{"step_id": "S1", "title": "Program Charter & Executive Mandate", "description": "Secure explicit authority to mandate the incident management program, transforming it from an engineering initiative into a business-critical operational standard.\n\n* Appoint a Director of Incident Management (DIM) reporting to the CTO, with a dotted line to the Head of Compliance for audit readiness.\n* Publish a charter defining scope: all customer-impacting incidents, payment path failures, and data integrity issues.\n* Establish a steering committee (CTO, VP Eng, Head of Support, Legal) to resolve cross-team conflicts and approve tooling budgets.\n* Define the 'Own Your Code, Own Your Pager' principle as a non-negotiable engineering value, directly addressing the 28-team resistance.", "dependencies": []}, {"step_id": "S2", "title": "Baseline Audit & Evidence Framework", "description": "Create a defensible starting point for metrics and a structured evidence repository for the SOC 2 Type II audit.\n\n* Consolidate the last 12 months of incident data into a single register, flagging the 31 customer-impacting incidents.\n* Map current alert sources to specific services to quantify the 3,400 monthly alerts and identify the top 50 noisiest rules.\n* Conduct a survey of on-call engineers to measure current sentiment and identify specific pain points in the unpaid rotations.\n* Designate a secure, immutable storage location for all future incident artifacts (timelines, logs, postmortems) to satisfy audit requirements.", "dependencies": ["S1"]}, {"step_id": "S3", "title": "Severity Taxonomy & Trigger Matrix", "description": "Define unambiguous severity levels tied to specific triggers, communication SLAs, and executive notifications to eliminate classification debates.\n\n* **SEV-1 (Critical):** System-wide outage, data loss risk, or >5% revenue-impacting transactions. Triggers immediate IC, VP Eng, and legal notification.\n* **SEV-2 (Major):** Significant degradation or region-wide latency. Triggers IC and Team Lead.\n* **SEV-3 (Minor):** Localized issue with workaround. Triggers on-call engineer.\n* **SEV-4 (Internal):** Low-priority noise. Triggers ticket only.\n* Map each severity to specific SLA credit exposures to align technical response with financial risk.", "dependencies": ["S1"]}, {"step_id": "S4", "title": "Incident Roles & Command Structure", "description": "Standardize the Incident Command System (ICS) to prevent the 'ambiguity of ownership' that caused previous hour-long delays.\n\n* **Incident Commander (IC):** Single point of decision authority; does not debug. Required for SEV-1/2.\n* **Scribe:** Maintains the real-time timeline log; mandatory for SEV-1.\n* **Comms Lead:** Owns status page and internal broadcasts; shields IC from external noise.\n* **SMEs:** Technical responders focused solely on diagnosis/mitigation under IC direction.\n* Create role cards with clear 'Do' and 'Don't' responsibilities for quick reference during chaos.", "dependencies": ["S3"]}, {"step_id": "S5", "title": "On-Call Architecture & Compensation Policy", "description": "Establish a sustainable, paid 24x7 on-call model that respects ownership boundaries and compensates for disruption.\n\n* Adopt a federated model: each of the 28 teams maintains its own SME rotation for owned services.\n* Create a central IC rotation staffed by senior engineers across teams to ensure command coverage.\n* Implement a paid policy: base stipend per shift, enhanced pay for call-outs, and mandatory recovery time after SEV-1 night shifts.\n* Define the 'Unowned Service' rule: services without a dedicated team rotation must be transferred to an owning team or decommissioned.", "dependencies": ["S4"]}, {"step_id": "S6", "title": "Tooling Consolidation & Integration", "description": "Replace the six fragmented alerting tools with a single Incident Management Platform to ensure a unified source of truth.\n\n* Select a platform (e.g., PagerDuty + Statuspage + Slack/Teams) capable of handling 180 services.\n* Integrate Kubernetes and PostgreSQL monitoring to feed alerts into the central platform.\n* Automate the creation of incident channels and timelines upon alert acknowledgment.\n* Ensure the status page is decoupled from the production environment to remain available during outages.", "dependencies": ["S2", "S5"]}, {"step_id": "S7", "title": "Alert Quality & Noise Reduction Program", "description": "Implement strict alert hygiene rules to reduce the 3,400 monthly alerts and restore trust in the pager.\n\n* **Rule:** All paging alerts must be symptom-based (customer impact) and have a linked runbook.\n* **Rule:** Implement deduplication and rate-limiting at the ingestion layer.\n* **Sprint:** Conduct a 90-day noise sprint to fix, tune, or delete the top 100 noisiest rules.\n* **Metric:** Enforce a 'Page-to-Action' ratio target; if an alert does not result in human action, it must be demoted to ticket-only.", "dependencies": ["S6"]}, {"step_id": "S8", "title": "Detection Strategy & SLOs", "description": "Shift from infrastructure metrics to Service Level Objectives (SLOs) to detect customer-impacting issues faster than clients can.\n\n* Define SLIs for critical payment paths: Transaction Success Rate, API Latency, Settlement Lag.\n* Implement synthetic transaction monitoring from external vantage points in both AWS regions.\n* Establish error budgets; breach of error budget triggers automatic review and potential feature freezes.\n* Integrate customer-reported signals (support tickets, social media mentions) into the detection workflow as low-priority alerts.", "dependencies": ["S3", "S6"]}, {"step_id": "S9", "title": "Communication Protocols & Templates", "description": "Define who talks, when, and what they say to stop ad-hoc status updates and protect brand reputation.\n\n* **Status Page:** SEV-1 initial update within 15 mins; subsequent updates every 30 mins. Use legal-approved templates.\n* **Internal:** IC broadcasts to #exec-leadership for SEV-1 every hour.\n* **Regulatory:** Define triggers for NY financial regulator notifications; assign Legal/Compliance responsibility for filing.\n* **Client Success:** Provide Account Managers with a 'Known Issues' playbook and real-time access to status page APIs for high-value accounts.", "dependencies": ["S4", "S6"]}, {"step_id": "S10", "title": "Postmortem Framework (Blameless)", "description": "Standardize the learning process to ensure consistent quality and adherence to the 'Blameless' culture.\n\n* **Mandatory:** All SEV-1 and SEV-2 incidents require a postmortem within 5 business days.\n* **Format:** Fixed template including Impact, Timeline, Root Cause, Contributing Factors, and Action Items.\n* **Review:** Facilitated by a neutral party (not the IC) to enforce blameless language and prevent finger-pointing.\n* **Publication:** All postmortems published internally on the Wiki with full searchability.", "dependencies": ["S9"]}, {"step_id": "S11", "title": "Action Item Tracking & Governance", "description": "Close the loop on incident resolution by enforcing rigorous tracking of corrective actions to fix the 17% completion rate.\n\n* Automatically create Jira tickets for every action item identified in the postmortem.\n* **Enforcement:** SEV-1 incident cannot be marked 'Closed' until high-priority actions are completed or deferred with VP approval.\n* **Review:** Weekly review of overdue actions in the Engineering Leadership standup.\n* **Metric:** Track 'Mean Time to Remediation' for action items as a key health indicator.", "dependencies": ["S10"]}, {"step_id": "S12", "title": "SOC 2 Control Mapping", "description": "Proactively map the new incident processes to SOC 2 Trust Services Criteria to ensure audit readiness.\n\n* Map S4 (Roles), S9 (Comms), and S10 (Postmortems) to Security and Availability criteria.\n* Define 'Evidence of Operation' for each control (e.g., automated timeline logs, signed-off postmortems).\n* Identify gaps between current state and audit requirements; assign remediation tasks to the DIM.\n* Establish a quarterly internal compliance review to test control effectiveness before the Type II audit.", "dependencies": ["S2", "S10"]}, {"step_id": "S13", "title": "Training & Certification Curriculum", "description": "Equip all engineers with the skills to operate within the new framework, reducing anxiety and improving response quality.\n\n* **Universal Training:** 1-hour module on severity levels and tools for all engineers.\n* **IC Certification:** Mandatory workshop and simulation for engineers joining the central IC rotation.\n* **Runbook Review:** Each team must update and validate runbooks for their top 3 critical alerts.\n* **Onboarding:** Include incident response basics in the engineering onboarding checklist.", "dependencies": ["S4", "S6"]}, {"step_id": "S14", "title": "Pilot Implementation (Wave 1)", "description": "Deploy the new process to a controlled subset of high-traffic teams to validate assumptions before broad rollout.\n\n* Select 3 teams: Payments Core, Ledger/API, and one Infrastructure team.\n* Run the full cycle for 6 weeks: Alerts, IC handover, Comms, Postmortems.\n* Collect feedback daily and adjust playbooks and tooling configurations based on real-world friction.\n* Measure success using the baseline metrics from S2 (MTTD, Noise, Satisfaction).", "dependencies": ["S7", "S8", "S9", "S13"]}, {"step_id": "S15", "title": "Full Rollout Strategy (Waves 2-4)", "description": "Gradually migrate the remaining 25 teams using lessons learned from the pilot to manage change resistance.\n\n* **Wave 2 (Month 3):** Deploy to 8 remaining critical customer-facing teams.\n* **Wave 3 (Month 4):** Deploy to internal platform and data teams.\n* **Wave 4 (Month 5):** Deploy to remaining low-traffic teams and legacy services.\n* Maintain a 'Help Desk' for new on-call engineers to troubleshoot process questions during their first shift.", "dependencies": ["S14"]}, {"step_id": "S16", "title": "Simulations & Game Days", "description": "Test the resilience of the process and the tools under controlled failure conditions.\n\n* **Tabletop Exercises:** Quarterly scenario walks-throughs for leadership and ICs (e.g., 'What if the primary region fails?').\n* **Chaos Engineering:** Inject failures in non-production or canary environments to test alert accuracy and runbook validity.\n* **Communication Drills:** Simulate SEV-1 to test the speed of status page updates and internal notification paths.\n* Document findings in postmortems and create action items for identified weaknesses.", "dependencies": ["S15"]}, {"step_id": "S17", "title": "Metrics Dashboard & Executive Review", "description": "Establish a continuous feedback loop to monitor the health of the incident management system.\n\n* Build a real-time dashboard tracking MTTD, MTTR, Alert Volume, and SLA Credits.\n* **Weekly:** Operational review of new incidents and action items with the DIM and Team Leads.\n* **Monthly:** Executive review with CTO/VP Eng focusing on trends, noise ratios, and on-call fatigue.\n* Publish 'Reliability Reports' to all engineers to celebrate wins and share learnings.", "dependencies": ["S2", "S15"]}, {"step_id": "S18", "title": "Culture & Change Management", "description": "Address the cultural shift required to move from 'hero mode' to 'process mode' and mitigate pushback.\n\n* Highlight success stories where the new process reduced toil or prevented customer churn.\n* Enforce the 'Blameless' mantra by publicly correcting any leadership communication that implies blame.\n* Recognize and reward effective ICs and engineers who improve runbooks or alert quality.\n* Address 'Pager Fatigue' proactively by monitoring call-out volume and rotating engineers off if thresholds are breached.", "dependencies": ["S5", "S15"]}, {"step_id": "S19", "title": "SOC 2 Dry Run & Evidence Prep", "description": "Conduct a mock audit six months out to identify gaps in evidence retention or process execution.\n\n* Sample 10 recent incidents and verify all required artifacts (timeline, logs, postmortem, actions) are present and accessible.\n* Interview on-call engineers to ensure they can describe the process and their roles without hesitation.\n* Remediate any 'Control Failures' identified during the dry run.\n* Prepare the 'Audit Readiness' package for the external auditors.", "dependencies": ["S12", "S17"]}, {"step_id": "S20", "title": "Continuous Improvement Loop", "description": "Institutionalize the evolution of the incident process to prevent stagnation.\n\n* Review the Severity Taxonomy quarterly to ensure it reflects current system behavior and customer expectations.\n* Update Tooling integrations annually based on vendor roadmaps and new AWS/K8s features.\n* Revise Compensation Policy annually based on market data and internal fairness reviews.\n* Set annual 'Reliability Goals' for the engineering organization to drive long-term focus.", "dependencies": ["S17", "S19"]}], "estimated_complexity": "high", "success_metrics": "- Median Time to Detect (MTTD) < 10 minutes.\n- Median Time to Mitigate (MTTR) < 45 minutes for SEV-1.\n- Customer-detected incidents < 5% of total incidents.\n- Monthly alert volume < 400 actionable alerts (90% reduction in noise).\n- SLA credit payouts < $100k annually.\n- Postmortem action item completion rate > 90% within 30 days.\n- 100% of SEV-1 incidents have a designated IC and Scribe.\n- On-call engineer satisfaction score > 4.0/5.0.\n- Zero critical findings in SOC 2 Type II audit regarding incident response."}Round 2 — refinement 2 of 2
P1 absorbed almost the entire distinctive vocabulary of P2's round-1 plan (Triage Owner, severity×class, three rotations, golden incident file, capped action items), producing two near-twin heavyweight plans; P2 added the genuinely new ideas of the round (a two-week minimum viable process, a funded ledger-resilience track, a cost-of-downtime model). P3 went the other way and collapsed into 22 bare titles with no content, losing everything that made it assessable.
What still separates them
- Bridging the gap before tooling exists: P2 S5 defines ten day-one rules, a manual duty-IC rotation drawn from the 12 teams that already have on-call, and a daily 15-minute stand-up for month one. P1 has no interim process between the charter (S1) and platform selection (S11); P3 has none either.
- Prevention as a funded track: P2 S23 is a standalone engineering roadmap with concrete bets (ledger read-only tripwire, connection-pool isolation, PITR restore tests with published timings, rollback on SLO burn). P1 keeps resilience as two bullets inside S9 and S22; P3 does not mention it.
- Communication clocks: P1 S13 demands a status-page update within 3 minutes of SEV-1 and internal updates every 5 minutes; P2 S13 sets 15 minutes internal first, 30 minutes to status page, with a mandatory no-news update. P1's 3-minute rule contradicts its own success metric of 30 minutes for 95% of SEV-1s.
- Level of specification: P1 and P2 give thresholds, timers, dollar ranges and dates throughout; P3 R2 gives only step titles and a one-line rationale each — no severity definitions, no timings, no compensation mechanics, no dates on any metric.
Who took what from whom
- P1 took nearly all of P2's round-1 signature ideas: Triage Owner (P2 S5→P1 S5/S6), severity×class with "class can raise, never lower" (P2 S4→P1 S4), the three-rotation model (P2 S8→P1 S7), priced opt-out and amnesty (P2 S9→P1 S8), golden incident file and month-one control mapping (P2 S1/S3→P1 S1/S3), three capped action items and the repeat-incident design review (P2 S15→P1 S16).
- P2 took P1's weekly synthetic-page testing of escalation ladders (P1 R1 S9→P2 S7) and P1's status-page-component-to-customer-journey mapping plus a named status-page owner (P1 R1 S14→P2 S14).
- P2 took P3's culture and change-management step (P3 R1 S18) and turned it into S24: on-the-spot correction of blame language, pager-fatigue monitoring, public recognition for deleted alerts.
- P3 took P2's "evidence clock" framing into its S1 title but nothing else of substance; it adopted no new mechanisms this round.
- Nobody adopted P3's error budgets triggering feature freezes (P3 R1 S8) or its rule that a SEV-1 cannot close until high-priority actions are done (P3 R1 S11) — P1 S16 and P2 S16 instead cap actions at three and track them in a separate reliability backlog.
The calls of this round
Influences: who took what from whom
| Round 2 ↓ · round 1 → | Proposal 1 | Proposal 2 | Proposal 3 | New steps |
|---|---|---|---|---|
| Proposal 1 |
kept13 | same titles6 analyst sees+10 / −1 | same titles0 analyst sees+0 / −1 | new3 |
| Proposal 2 |
same titles2 analyst sees+3 / −3 | kept16 | same titles2 analyst sees+1 / −1 | new4 |
| Proposal 3 |
same titles9 analyst sees+0 / −0 | same titles4 analyst sees+2 / −2 | kept9 | new0 |
P1 rewrote itself around P2's round-1 mechanisms while keeping its own depth. New steps 4, 5, 7, 12, 14, 15, 20, 21 replace vaguer round-1 equivalents, and detection, escalation and postmortem policy are now far more concrete. A truncated step 11 and a few internal contradictions are the cost.
- S5 adds the Triage Owner rule with a 15-minute triage decision, closing the unowned gap between page and declaration that caused the two hour-long ambiguity incidents.
- S4 adds response classes orthogonal to severity, with automatic triggers (ledger write failure → SEV-1, payment success <99% for 5 min, missed settlement window) instead of round-1's generic definitions.
- S7 replaces the single "on-call architecture" step with three named rotations (Service, Platform Duty, IC roster of 12–16) and a hard gate: build a rotation of ≥6 or transfer ownership by end of month 2.
- S16 caps postmortems at three action items, requires an artifact for closure, and adds the repeat-incident design-review rule — a plausible fix for 11-of-64 rather than more tracking.
- S12 adds mechanical escalation timers, the dual-IC rule, the ambiguity rule and a 30-minute Watch state.
- Metrics now carry per-month deadlines and include credit avoidance, detection contracts on all 180 services, and golden-file completeness.
- Step 11 is cut off mid-sentence ("Define rollback criteria (e.g.,"), losing the dual-run exit criteria and cutover date.
- Success metric says SEV-1 MTTM <45 minutes by month 9, but S18 sets the target at <60 minutes — two numbers for the same thing.
- S13's 3-minute status-page update for SEV-1 contradicts the stated metric of 30 minutes for ≥95% of SEV-1s, and is barely achievable with legal-cleared templates.
- Alert target tightened to <400/month in the metrics while S10 still says <600 — inconsistent.
- No interim operating model: the plan produces nothing usable until tooling is selected, despite claiming "working process in month 2".
- Proposal 2 : The Triage Owner owns the incident from page acknowledgement until an IC takes over or stand-down.
- Proposal 2 : Severity times response class, where class can raise the response but never lower it.
- Proposal 2 : Three rotations — Service On-Call, Platform Duty, central IC roster.
- Proposal 2 : Opt-out with a price paid by the team, plus an amnesty on incident records in performance reviews.
- Proposal 2 : Start the SOC 2 evidence clock on day one and map controls in month one, with a golden incident file as the audit unit.
- Proposal 2 : Cap postmortems at three action items; a repeat contributing factor triggers a design review, not another ticket.
- Proposal 2 : Never publish incident count as a team metric; publish reporting metrics instead.
- Proposal 2 : A regulator clock matrix mapping event type to regulator, window and signer.
- Proposal 2 : Detection-gap ticket when a customer reports first, and two-week ticket-only probation for new alerts.
- Proposal 2 : Sequence rollout waves by incident density and cost of failure, not by ease.
- Proposal 2 : A 15-minute internal first update and a 30-minute status-page clock for SEV-1.
- Proposal 3 : A SEV-1 cannot be marked closed until high-priority action items are completed or deferred with VP approval.
+ Severity and response class taxonomy+ Incident lifecycle, Triage Owner rule, and escalation policy+ Three on-call rotations: Service, Platform, and Incident Commander+ Status page infrastructure and customer-impact ledger+ Postmortem policy: mandatory, blameless, three-level framework+ Phased rollout sequenced by cost of failure+ SOC 2 dry run and evidence reviewSeverity taxonomy and trigger matrixAlert consolidation and event pipelineOn-call architecture and 24x7 coverage modelPlaybooks and communication templates by severityStatus page, customer notifications, and account-manager playbookPostmortem policy, blameless process, and facilitationPhased rollout to all 28 teamsSOC 2 dry run, gap remediation, and audit support
The plan produced
1. Executive charter, governance structure, and evidence clock
Turn the CEO email into a funded, authorized program with clear ownership and documented evidence collection for SOC 2, starting today.
- Appoint a Director of Incident Management reporting to CTO, with dotted line to COO (customer impact) and Head of Compliance (audit readiness).
- Publish a one-page charter: scope (all customer-impacting, payment-path, data-integrity incidents across 28 teams and 2 regions), decision rights (IC may freeze changes, override team preferences during incidents), and authority to mandate process participation.
- Secure annual budget for tooling, training, on-call compensation ($500–800K estimated), and resilience work. Connect funding to avoided SLA credits ($1.3M baseline).
- Establish standing Incident Management Steering Group: CTO, VP Engineering, Head of Support, Head of Compliance, one engineering manager per region. Meet monthly.
- Start the SOC 2 evidence clock on day 1. An audit in eight months means operating-period evidence begins now; design the process to capture evidence continuously, not retroactively.
- Make incident-response participation a documented performance expectation for all engineering managers and team leads.
- Publish timeline: working process in month 2, all 28 teams in month 5, audit-ready in month 7.
2. Baseline measurement, incident register, and evidence pack (after 1)
Establish defensible baseline metrics and identify structural gaps that explain the 40% customer-detected rate and 22-minute detection time.
- Build a 12-month incident register with all 31 customer-impacting incidents: date, detection source, detection time, mitigation time, customer count, services involved, SLA credits paid, root cause class.
- Audit the current alert estate: total volume per tool, volume per team, volume per service, page-to-action ratio, top 50 noisiest rules, off-hours interruptions per engineer.
- Construct a silent-failure register: incidents with no internal alert fired at all. This explains the 40% customer-detected rate.
- Reconstruct the two command-ambiguity incidents minute by minute: exactly when did ownership become unclear, how long, what was the decision bottleneck.
- Survey on-call engineers (target >70% response): burden, fairness, pay expectations, escalation clarity, willingness to stay.
- Interview Support and Account Management: how do customers discover incidents, what do they complain about, how do they contact you.
- Publish the problem statement internally; retain all artifacts for SOC 2 audit evidence. This is the baseline against which all improvements are measured.
3. Control mapping and evidence architecture (after 1, 2) from P2 step 3
Design the process to generate SOC 2-compliant evidence automatically, from the first real incident, so the audit clock ticks in your favour.
- Map the new process to Trust Services Criteria CC7.1–7.5 (incident identification, response, evaluation, containment, communication), CC2.2–2.3 (authorization), CC4.1 (change management), CC3.x (information availability).
- For each control, write a one-paragraph plain-language statement, name a single owner, and specify the evidence artifact (incident record, timeline, communications log, postmortem, action tracker, training record).
- Define the golden incident file: one single-click export per incident containing severity, timeline, roles assigned, decisions made, communications sent, postmortem, and action items. This is the audit unit.
- Specify data retention, immutability, access control, and storage location (not laptops, not chat history that expires). Ensure evidence is searchable and organized by incident date.
- Keep a gap register with owners and dates; review fortnightly in the steering group. Identify which controls are blocked by incomplete rollout and when they unblock.
- Run an early design walkthrough with an experienced SOC 2 readiness partner inside month 1 to stress-test control design before building on it.
4. Severity and response class taxonomy (after 2) from P2 step 4
Define four severity levels and four response classes so every decision—paging, communications, postmortem, compensation—keys off a defensible rule, not a judgment call.
- Severity by impact scope: SEV-1 (total payment-path loss, data corruption, or >5% transaction failure for >5 min); SEV-2 (significant degradation or single region loss); SEV-3 (limited impact with workaround available); SEV-4 (internal issue or cosmetic); SEV-0 (reserved for security/regulatory/privacy events).
- Response class (orthogonal to severity): Availability, Performance, Data Integrity & Ledger, Security & Privacy. Key rule: class can raise severity, never lower it. A SEV-3 data-integrity incident gets SEV-1 response posture because integrity is not recoverable by moving faster.
- Automatic triggers: loss of one AWS region → SEV-1 or SEV-2 (class-dependent); ledger write failures → SEV-1; replication lag >10s → escalation review; payment success rate <99% for >5 min → SEV-1/2; missed settlement window → SEV-1; total external API unavailability → SEV-1.
- Who may declare: any engineer, Support, account manager (based on observed customer impact). Who may downgrade: IC only, after investigation.
- Map each level to SLA credit exposure and to customer-facing status-page state.
- Include worked examples from the last 12 months so all 28 teams recognize their own incidents in the taxonomy. Re-validate quarterly against real declarations.
5. Incident lifecycle, Triage Owner rule, and escalation policy (after 4) from P2 step 6
Eliminate the "nobody was in charge for over an hour" problem by assigning ownership the moment a page is acknowledged.
- Define lifecycle states with clear entry/exit criteria: Detected (alert fired) → Triaged (is this real and customer-impacting?) → Declared (severity assigned) → Mitigated (core issue resolved) → Resolved (all verifications done) → Postmortem (review scheduled) → Closed (action items tracked or dismissed).
- Introduce the Triage Owner rule: the person who acknowledges the page owns the incident until an IC is assigned or the incident is stood down. There is never an unowned gap between first page and declaration. Triage Owner's sole job: decide within 15 minutes whether this requires an IC or a direct stand-down.
- Set aggressive acknowledgement and declaration targets: page acknowledged in 5 min; triage decision (is this real?) in 15 min; severity declaration in 30 min for any customer-facing incident.
- Implement automatic escalation ladders with no human judgment required: if responder does not acknowledge in 5 min, escalate to service owner; if no ack in 10 min, escalate to team manager; if no ack in 15 min, escalate to IC on-call. Escalation is never criticized.
- Define unresponsive-team path: if a service's on-call is unreachable for 30 min, IC may direct any available engineer from any team to engage.
- For SEV-1 and SEV-2: change freeze until IC declares mitigation confirmed; IC unfreezes changes explicitly.
- Enforce one incident, one record. Timeline auto-captured from Slack channel and bridge; never written from memory later.
6. Incident roles, command structure, and decision rights (after 5)
Define clear roles with one-page responsibility cards published and linked from every paging notification.
- Incident Commander: owns incident outcome, not the fix. Declares severity, decides escalation, approves all customer communications, freezes changes, calls responders, hands off in shifts. Non-technical ICs are acceptable; technical depth is not required.
- Deputy IC: assigned within 15 min of declaration; shadows IC; takes over if IC unavailable or after 4-hour shift on SEV-1. Maximum IC shift: 4 hours on SEV-1, 6 hours on SEV-2.
- Triage Owner (new role): owns incident from first page acknowledgement until IC takes over or stand-down decision is made. Required for all incidents.
- Communications Lead: owns internal Slack updates and status-page messaging; shields IC from customer contact and interruptions.
- Scribe: records real-time timeline with decisions, actions, and key timestamps; not responsible for fixing.
- Subject-Matter Responders: engineers with service context; take IC direction; report only to IC; no side channels or parallel debugging.
- Operations Lead (SEV-1 only): coordinates multiple responders, manages incident bridge, maintains escalation list.
- Minimum viable staffing: SEV-1 requires all roles; SEV-2 requires IC, Deputy, Comms, Scribe, SMEs; SEV-3 requires Triage Owner and IC.
- Create laminated role cards for every on-call shift location (office, home, printed in pockets).
7. Three on-call rotations: Service, Platform, and Incident Commander (after 5, 6) new
Directly address the "carrying a pager for another team's code" objection by making it structurally impossible.
- Service On-Call rotation (federated): each of the 28 teams maintains a rotation for their own services only. No engineer is paged for code their team does not own. The answer to "why am I carrying a pager?" is now simply: "for your team's code."
- Platform Duty rotation (centrally staffed): shared PostgreSQL cluster, Kubernetes, networking, CI/CD, observability, and incident management tooling. Nobody's product code, so it gets its own dedicated rotation. Staffed from platform teams plus volunteers from other teams; paid at premium rate.
- Incident Commander roster (24x7): 12–16 certified senior engineers from across all 28 teams, on one-week primary shifts with secondary backup. Covers every hour with no single point of failure and no uncovered holiday week.
- Consequences and gates: 16 of 28 teams have no on-call today. Each must either (a) build a Service On-Call rotation of at least 6 engineers, or (b) formally transfer service ownership to a team that will, with transfer documented and dated. No exceptions, no waivers. Unowned services are decommissioned or transferred by end of month 2.
- Merge small or low-traffic teams into shared rotations where service ownership is unclear (e.g., shared analytics, testing infrastructure).
- Enforce scheduling limits in the tooling: no engineer on-call more than 2 weeks per quarter, automatically enforced by configuration, not negotiation.
- Publish a coverage matrix for all 28 teams showing services, owners, rotation size, gaps, and monthly status.
8. On-call compensation, rest policy, and sustainability (after 7)
Settle compensation before rollout, not during negotiations. Make on-call sustainable and valued.
- Paid on-call: effective immediately upon joining a rotation. Weekly stipend while on shift (benchmark to New York market: $600–1,000 per week per engineer), regardless of incident volume.
- Event-based compensation: 1.5× hourly rate for time spent mitigating out-of-hours incidents, minimum one-hour block per callout. Tracked by incident record (auto-capture from timeline).
- Compensatory rest: no engineer works a normal 8-hour business day after a night incident requiring >2 hours mitigation. Rest day is documented policy, not a favour granted by manager.
- Intrusion cap: maximum 3 unscheduled pages per week per engineer. Exceed the cap in a week and trigger an immediate review; exceed in a month and escalate to VP Engineering. Breaches are structural signal that alert quality or service stability has a problem.
- Voluntary opt-out: an engineer may exit a rotation; their team must hire or buy replacement coverage from paid pool at published internal rate ($X per shift). This converts culture debate into visible budget decision.
- Amnesty policy: incident records, near-miss reports, and false declarations are never used in performance reviews or compensation discussion. Only failure to report is a performance issue.
- Policy publication: publish compensation structure and effective date before any team is asked to join a rotation, and include on-call expectations in job descriptions and hiring conversations.
- Semi-annual review: reassess compensation and caps every six months against actual page volumes, attrition rates, and survey feedback.
9. Detection strategy: SLOs, synthetic monitoring, and customer-report intake (after 4, 7) from P2 step 7
Close the 40% customer-detected gap by monitoring customer journeys instead of infrastructure metrics.
- SLO-based alerting: Define SLIs and SLOs for the top 20 customer journeys (payment initiation, authorization, settlement, ledger read/write, API availability, webhook delivery, payout). Measure per region. Alert on SLO breach, not on infrastructure metric (e.g., alert on "payment success rate <99%" not "database CPU >80%").
- Synthetic transaction monitoring: deploy synthetic transactions from outside AWS in both regions plus a third geographic location, one-minute cadence, for all money-moving paths. These are your first alarm bell.
- Ledger-critical signals: PostgreSQL replication lag (target: <1s, alert >5s), connection saturation, write latency (p95), lock-wait time, transaction ID exhaustion proximity, checkpoint pressure, table bloat. These are separate alerts on shared-database health.
- Customer-report intake (new detection channel): Support and Account Managers can raise an incident directly in the platform. Every customer report creates an incident record automatically, and the "customer report" detection source is counted in all metrics. This is a legitimate detection method, not a failure.
- Detection-gap rule: whenever a customer reports an incident before internal monitoring fires, auto-create a ticket in the owning service's backlog with root cause: "Monitoring gap on [journey]."
- Detection contract per service: every one of the 180 services needs a named owner, at least one symptom-based alert mapped to a SLO, and a documented expected detect time (target: <5 min for payment path, <10 min for others). Published on wiki and reviewed monthly.
- Detection drills: run a quarterly drill per team: simulate a broken service in staging and verify it triggers a page before a human notices.
- Resilience roadmap separation: detection improvements do not protect against ledger corruption or multi-region failure. Fund a separate resilience roadmap to reduce shared-database blast radius and improve failover safety.
10. Alert quality standards and noise-reduction program (after 9)
Cut the 3,400 monthly alerts (85% noise) to <600 with 95% signal. This is the price of admission for on-call buy-in.
- Paging contract: every page must satisfy all of (1) symptom-based (customer impact, not infrastructure cause), (2) actionable (linked runbook with immediate next step), (3) owned (named team responsible), (4) severity-mapped (SEV-1/2/3/4), (5) SLO-linked where applicable. No runbook, no page. Enforce with CI check on alert definition.
- Separation rule: only customer-impacting or imminent-impact signals trigger pages. Everything else becomes a dashboard entry, a ticket, or a log line. Noisy infrastructure metrics go to dashboards, not pagers.
- Page budget per service: no service may exceed 2 pages per on-call shift per month. Exceeding budget auto-opens a remediation ticket in the owning team's backlog (with alert-quality review assigned to tech lead).
- Automatic suppression rules: (1) silence alerts if service auto-recovered within 30s, (2) suppress known maintenance windows, (3) group flapping alerts (>5 in 2 min) into one page, (4) rate-limit noisy services (max 1 page per 5 min until condition clears). All suppression rules must have an expiry date; no permanent silence without a ticket.
- Probation for new alerts: new alert rules run as tickets only and alert to a Slack channel; after two weeks of proving actionability (every alert resulted in human action), they graduate to pager.
- Noise sprint: run a focused 90-day program with a public burn-down of the top 100 noisiest rules. Assign each to a named manager. Default action: fix root cause, tune threshold, or delete within 10 working days. Deletion is a legitimate successful outcome (celebrate it).
- Correlation and deduplication: consolidate alert sources at ingest pipeline so one outage triggering 40 alerts produces one page, not 40.
- Alert ownership: every alert must have an owning team and a maintenance contact. Update monthly.
11. Incident tooling consolidation and integration (after 5, 10)
Replace six alert tools and ad-hoc incident records with a single source of truth that unifies paging, escalation, timeline, and audit evidence.
- Tool selection: choose an incident-management platform (e.g., PagerDuty, Incident.io, Opsgenie) that integrates paging schedules, escalation policies, incident records, postmortem workflow, and status-page APIs. Decision gate: month 1.
- Event pipeline consolidation: route all alerts from the six legacy tools into a single event pipeline that feeds the incident platform. Apply deduplication, correlation, severity/class mapping, and rate-limiting at ingest.
- Observability integration: connect the incident platform to your Kubernetes dashboards, PostgreSQL monitoring, distributed tracing, and logs so responders see context in one pane. Link runbooks directly into incident records.
- Slack and bridge integration: auto-create incident Slack channels, auto-invite roles, auto-capture timeline from channel transcript and voice-bridge recording. Timeline is not written from memory; it is auto-captured.
- Golden incident file: implement the export defined in S3. One click produces a complete, immutable, audit-ready PDF: severity, timeline, roles, decisions, communications, postmortem, action items, and closure evidence.
- Dual-run period: run both legacy and new platform in parallel for two weeks. Define rollback criteria (e.g.,
12. Escalation automation and incident lifecycle enforcement (after 5, 11)
Eliminate judgment calls from the worst moments. Escalation is automatic, mechanical, and blameless.
- Automatic escalation ladders: page responder → if no ack in 5 min, page service owner → if no ack in 10 min, page team manager → if no ack in 15 min, page IC on-call + call them immediately (phone + SMS + Slack). No human decides to escalate; timers fire escalations.
- Severity-based escalation tempo: SEV-1 uses faster timers (2 min for IC on-call), SEV-2 uses moderate timers (5–10 min), SEV-3 uses slower timers (15–30 min). Configured in tooling, reviewed quarterly.
- Dual IC rule: if a second SEV-1 incident is detected while the first is active, immediately page and assign a separate IC. ICs never run two incidents in parallel.
- Change freeze and rollback authority: SEV-1 and SEV-2 trigger automatic deploy freeze. Only the IC (with CTO/VP Eng notification) may unfreeze. Freeze lifts only when IC explicitly declares mitigation confirmed and verifies no new incident symptoms for 5 min.
- Unresponsive team escalation: if service's on-call does not acknowledge in 30 min, IC may direct any engineer from any team (volunteers first, then rotated) to engage. This is documented and reported in monthly review (escalation = signal of rotation problem).
- One incident, one record: all decisions logged in the incident platform. Auto-capture from Slack, bridge, status-page updates. Timeline is the source of truth; postmortem is written from timeline, never constructed after the fact.
- Ambiguity rule: if two responders disagree about whether an incident should be declared, it is declared. False declarations (stand-downs within 30 min of declaration) are tracked as metrics and closed without blame.
- Watch state: an unconfirmed incident can live in "Watch" state for max 30 min; after that, either declare it or stand it down explicitly.
13. Internal, customer, and regulatory communications workflows (after 6, 12)
Define who informs whom, in what order, via what channel, with explicit timings and pre-approved templates.
- Internal cadence: first update to #incidents Slack channel within 3 min of declaration (even if "Investigating"). Then updates every 5 min (SEV-1), 15 min (SEV-2), or 30 min (SEV-3), or immediately on material change (e.g., mitigation achieved, scope widened). Comms Lead owns the update; IC must not be interrupted.
- Executive notification: IC calls CTO and VP Eng within 1 min of SEV-1 declaration (not email, not Slack, call). Incident declared in Slack with severity label, IC name, and affected service. Escalation channel lead auto-pinged.
- Customer communication channels: status page (all 2,100 customers), direct email to affected customers (top-tier accounts and customers affected by SEV-1), account-manager calls (top 50 accounts on SEV-1).
- Status page timings: update within 3 min of SEV-1 declaration, 10 min of SEV-2, 30 min of SEV-3 (even if root cause unknown; use "Investigating" with next-update ETA). Updates every 5–30 min depending on severity. Always include next-update time.
- Pre-approved templates: draft customer-facing language for each severity and class in advance with Legal and Compliance. Templates specify impact language ("some of your transactions are delayed" not "our database failed"), workarounds if available, and next-update commitment. Never speculate on cause in customer communication.
- Regulatory notification path: identify incidents requiring regulator notification (NYDFS Part 500, money-transmitter rules, payment-card-network rules, securities disclosure). Build a clock matrix: event type → regulator → notification window → signer. Compliance owns all regulatory notifications (never Engineering). Pre-clear templates. Flag incidents to Compliance immediately upon declaration.
- Account manager playbook: contact tree for top 50 accounts, templated talking points (facts only, never speculation), escalation path if customer escalates, what to offer (service credit, technical deep-dive call).
- Closing communication: resolution notice, SLA credit impact, commitment date for written root-cause report, customer action required (none, or security update, etc.).
14. Status page infrastructure and customer-impact ledger (after 13) new
Make the status page reliable, customer-centric, and audit-ready. Track customer impact in a single durable record.
- Status page decoupling: host status page outside production failure domain (separate cloud, separate infrastructure, separate database). Integrate incident platform with status page so incident record drives all public updates. Status page survives total platform outage.
- Component-to-journey mapping: status page components map to customer journeys ("Payments", "Settlements", "Payouts", "Ledger API") not to internal services. Allow customers to subscribe to components; notify by email or webhook.
- One-click update templates: pre-fill status-page template with severity, impact language, next-update time, and estimated resolution. Comms Lead types minimal new info ("Root cause identified" or "Workaround available"), and updates auto-post.
- Customer-impact ledger (one record per incident): which customer accounts affected, which journey(s) impacted, exact start and end time of impact, estimated SLA-credit exposure. Use this single record for customer communications, credit computation, regulatory reporting, and annual review. No reconciliation of two versions of the same outage.
- SLA credit automation: compute credit based on duration × severity × customer tier → auto-generate customer notification → auto-post to finance system. Reconcile accrued vs. paid credits monthly and report in executive review.
- Testing during game days: simulate status-page outage and verify alerts continue to fire; test total region loss and confirm status page remains updated; drill runbook for manually updating status page if platform is down.
15. Postmortem policy: mandatory, blameless, three-level framework (after 6, 13) new
Turn postmortems from a writing exercise (11 of 64 action items closed) into the learning engine of the system.
- Mandatory postmortems: all SEV-0, SEV-1, and SEV-2 incidents; all SEV-3 with customer impact or repeat pattern; any near-miss IC flags; any incident where the process itself failed (IC unreachable, Comms Lead unavailable, false declaration, missed update SLA).
- Three-level framework (proportionate to weight): (1) lightweight async review for SEV-4 and low-impact SEV-3 (10 min template in shared doc, owner + IC review), (2) standard facilitated postmortem for SEV-2 and impactful SEV-3 (full template, facilitated by trained neutral party, published within 10 days), (3) full executive postmortem for every SEV-1 and every security incident (executive sponsor assigned, full investigation, published within 15 days, customer-facing variant prepared).
- Fixed timeline: draft postmortem within 5 business days, blameless review within 10 days, internal publication within 15 days.
- Single template: impact (who, how many, how long, financial exposure), timeline (detection through resolution), root cause (not "human error" but system condition that enabled error; what was the gap?), contributing factors (tooling, process, organization, knowledge, monitoring), what went well, what went badly, action items (≤3, rest go to reliability backlog).
- Blameless facilitation: train a pool of blameless postmortem facilitators (target: 10+ engineers). Require a trained, neutral facilitator for every SEV-1 and SEV-2 review. Prohibit counterfactual language ("if the engineer had"), blame language, and the phrase "human error" as a root cause.
- Publication rule: publish all postmortems internally by default; security review only for genuinely sensitive material (e.g., unpatched vulnerability details or customer PII in logs). Create a customer-facing root-cause report for every SEV-1, especially for regulated customers, with legal and compliance sign-off.
- Searchability: store postmortems in a searchable wiki or issue tracker with tags (service, class, root cause category) so teams can learn from similar incidents without repeating them.
16. Action item tracking, reliability backlog, and repeat-incident design rule (after 15)
Close the loop on incident learning by enforcing verifiable, tracked action items and breaking cycles of repeat incidents.
- Action item capping: each postmortem generates a maximum of 3 action items. Anything beyond 3 goes into a ranked reliability backlog, not into the postmortem, to prevent overwhelming teams.
- Action item requirements: each item must have (1) a named human owner (not a team), (2) a due date (≤60 days, target ≤30 days), (3) a definition of done (merged code change, tested alert, audit evidence, architectural decision, new runbook, training completed) not self-reported status.
- Single reliability backlog: create one backlog in your engineering tracker (Jira, Linear, etc.) with mandatory label (e.g.,
incident-action), link to originating incident, and link to postmortem. Track progress weekly. - Closure sign-off: Incident Commander or postmortem facilitator must sign off on closure, verifying artifact exists (code merged, alert tested in drill, runbook verified).
- Repeat-incident rule: if the same service or component has a second incident with the same contributing factor, do not create another action item. Instead, escalate immediately to an architect or tech lead and trigger a design review (not a task, a review). This breaks the cycle of repeated patches; the system needs a structure change.
- Capacity protection: reserve a fixed percentage of each team's sprint capacity (10–15%) for reliability work. Track unspent capacity and report to VP Engineering monthly; if a team is not spending it, work with them to identify and fix blockers.
- Ageing and escalation: run a weekly review of open actions; escalate anything >30 days overdue to team lead and VP Engineering. Monthly report: completion rate (target >90% within 60 days) and median action age (target <30 days).
17. Training, certification, and exercise program (after 6, 13, 15, 16)
Build skills before deploying the process. Run ongoing drills so the system is tested, not guessed at.
- Curriculum: (1) All engineers (30-min async video): severity taxonomy, communication expectations, postmortem format, when to declare an incident, where to find runbooks. (2) On-call responders (2-hr workshop): alert tool walkthrough, playbooks by severity, escalation paths and timers, when to call manager, mobile app walkthrough, case studies from the last 12 months. (3) Incident Commanders (4-hr bootcamp + test): leadership under pressure, decision-making (severity, escalation, rollback), communicating with executives, status-page discipline, postmortem facilitation, handling ambiguity, live simulated incident (pass/fail certification). (4) Communications Leads (2-hr training): templates per severity and class, customer-communication rules (no speculation, no blame), update timings, how to shield IC, regulatory triggers.
- IC certification: written assessment (75% pass required) plus live simulated incident (role-play with facilitator, graded on severity declaration, escalation decisions, communication, handover). Certification valid for 12 months; recertify via annual refresher or another live sim.
- Depth across teams: certify at least 2 ICs per team or team group so central roster is not siloed in one group; no holiday week is uncovered.
- Async content: record all training videos so async teams can learn on their schedule. Create quick-reference cards (laminated, pocket-sized) for roles and playbooks; distribute to on-call locations (office, home).
- Monthly tabletop exercises: drawn from real incidents from the last 12 months (region loss, ledger write failure, missed settlement window, cascading failures). Facilitator describes scenario; 3–4 responders play out response (Triage Owner, IC, Comms) as if real. Run 30 min; retro for 15 min afterward.
- Quarterly game days: deliberately inject failures into production (database failover, status-page outage, alerting-pipeline outage, dual SEV-1 incidents). All on-call roles engage. Run 2–3 hours; measure response times, decision quality, and communication. Document findings and create action items for identified gaps.
- Drill the process's own failure modes: IC unreachable (on-call unavailable, phone broken), Comms Lead on PTO, two simultaneous SEV-1s, paging storm (100+ alerts), false alarm that consumes an hour. Test escalation paths, deputy takeover, and recovery.
- New-engineer onboarding: add incident-response module to all engineering onboarding (completion tracked, audit-ready). All engineers must complete within 30 days of joining or transferring in.
18. Metrics, dashboards, and review cadence (after 2, 12, 16, 17)
Measure to prove the system works. Publish live dashboards so every engineer sees the scoreboard and the system is transparent.
- Outcome metrics: Median Time to Detect by source (target: <5 min internally detected, <10 min customer journeys); Median Time to Mitigate for SEV-1/2 (target: <60 min SEV-1); customer-detected incidents as % of total (target: <5%); incidents by severity (should be mostly SEV-3/4, few SEV-1); SLA credits paid (target: <$100K/year by month 12); annual credit avoidance vs. program cost.
- Process metrics: IC assigned within 5 min (target: >95% of incidents); page acknowledgement rate (target: >98% within 5 min); first-update timeliness (target: >95% within SLA); postmortem timeliness (target: 100% of mandatory postmortems published on time); IC roster coverage (zero uncovered hours, monitored weekly).
- Health metrics: alert volume and signal-to-noise ratio per team (trending toward target); off-hours pages per engineer per month (trend, cap enforcement); on-call satisfaction survey (target: >7/10); training completion by team (target: 100% within 30 days); % of services with active detection contract (target: 100%).
- Never publish incident count as a team metric. Reward hiding. Instead publish detection metrics (near-misses reported per team, detection gaps closed, false declarations made).
- Live dashboards: build dashboards visible to all engineers (not just managers) showing outcome, process, and health metrics. Auto-populate from incident platform and alert tool. Update daily. Link from Slack and internal wiki.
- Baseline all metrics against S2 evidence pack. Set 90-day and 12-month targets for each metric. Publish targets and progress monthly.
- Review cadence: (1) weekly 30-min operational review (incident by incident from prior week: what went well, what hurt, actions); (2) monthly 60-min reliability review (trends, top causes, action-item aging, alert quality per team); (3) quarterly 60-min executive review (CEO's office: customer impact, SLA credits, top five systemic causes, program ROI).
- Quarterly process review: what in the process wasted responder time, what confused people, what should be deleted. Solicit feedback from ICs, Comms Leads, and responders. Document changes and reasoning.
19. Pilot program with 3–4 volunteer teams (after 6, 8, 11, 12, 13, 14, 15, 16, 17, 18)
Do not roll out untested to 28 teams. Run the entire process end-to-end with a small cohort using real incidents as the primary training material.
- Team selection: recruit 3–4 volunteers spanning criticality: one payment-path team, one ledger-adjacent team, one shared infrastructure team (platform or Kubernetes), one low-traffic team. Volunteers see early adoption and influence.
- Full process in pilot: new severity and class taxonomy (S4), consolidated tooling (S11), roles and Triage Owner (S5–6), three rotations (S7), escalation automation (S12), communications (S13–14), postmortems (S15), action tracking (S16), paid on-call (S8), training (S17), metrics (S18). This is not a partial test; it is the complete system.
- Real incidents are the training: hold a retro within 48 hours of each pilot incident (while memory is fresh). Process Owner facilitates. Discuss: what worked, what hurt, how is the runbook, is the alert tuned, did Comms template work, did roles work, was timeline auto-captured correctly. Document feedback and iterate.
- Instrument against baseline: measure all S18 metrics in pilot; compare results with S2 baseline. If MTTD dropped from 22 min to 8 min, that is the number you carry into the next conversation.
- Weekly retros with pilot teams: structured 30-min sessions reviewing policies, tooling, training, templates, and process changes. Change the written policy or tooling in response; document what changed and why so later teams see the evolution, not just the finished state.
- Explicit exit criteria: (1) rotation coverage sustained (≥6 engineers per rotation, zero unacknowledged pages over 3 weeks), (2) postmortems delivered on time (100% of mandatory postmortems published within 15 days), (3) action tracking established (100% of action items in backlog with owner and date), (4) metrics live (dashboards updated daily, first weekly review completed).
- Pilot report: document before/after numbers (MTTD, MTTR, alert noise, action completion rate, on-call satisfaction) and key process learnings. This report is the foundation for every conversation in the rollout.
20. Phased rollout sequenced by cost of failure (after 16, 18, 19) from P2 step 19
Rollout is a staged migration with readiness gates and sequencing that prioritizes visible impact, not ease.
- Wave sequencing: divide 28 teams into 4 waves of ~7 teams each, ordered by incident density and customer-journey ownership (highest-cost-of-failure teams first). Teams with the most SLA credits at stake go first; their improvement is the proof.
- Wave spacing: three weeks between waves. This gives each wave time to stabilize and find problems before the next cohort joins.
- Readiness checklist per team: (1) all services mapped and owned (no unowned services), (2) alerts cleaned to paging contract (runbook linked, severity mapped), (3) playbooks updated and tested in staging, (4) rotation staffed to ≥6 engineers, (5) team completes training module, (6) manager briefed on policy, (7) on-call compensation in effect.
- Gate review before each wave: Process Owner holds gate review with target teams. Move unready teams to next wave with a dated remediation plan. No exceptions, no waivers; readiness is non-negotiable.
- Wave champion: assign a named engineer per wave to champion the rollout, answer questions, escalate issues to Process Owner. Champions are not representatives; they are advocates and feedback collectors.
- Communication cadence: weekly all-hands or newsletter for 4 weeks before each wave. Explain why (owned-code-owned-pager rule, paid on-call, SLA credit savings). Use pilot numbers. Answer FAQs. Announce champion and escalation path.
- First incident under new process: hold a retro within 48 hours. Feed accepted process changes back through change control.
- Retire legacy tools and processes: at end of each wave, retire legacy alert tools, informal escalation lists, ad-hoc status-page process. No parallel processes running for >3 weeks; this prevents confusion and half-learning.
- Sequence to avoid audit collision: ensure no team is rolling out in the same week as the audit dry run (S21).
21. SOC 2 dry run and evidence review (after 3, 20) from P2 step 20
Convert a good working process into a provable one. Test control evidence a few months before auditors arrive, when you can still fix gaps.
- Dry run timing: run 6 weeks before audit window (around month 7 of this program).
- Scope: sample 10–15 real incidents from pilot and early rollout waves. For each incident, verify evidence artifact exists and is complete: incident record, timeline (auto-captured), severity and class declaration, roles assigned and logged, communications log (Slack + status page), postmortem (if mandatory), action items in tracker with due dates, action completion evidence (code, alert test, training record).
- Control walkthrough: walk through each control statement from S3 with a checklist. Is the evidence artifact present? Is it immutable? Is it searchable? Is retention adequate? Is access logged?
- Gap remediation: for every gap found, estimate time to fix and prioritize by audit risk. Anything risking a qualified opinion (e.g., missing postmortem, no timeline evidence) must be fixed before the audit. Test the remediation against a new incident or a resample.
- Interview readiness: brief 10–15 engineers who may be interviewed by auditors (ICs, Comms Leads, Process Owner, team managers). Ask them to describe the process as they actually practice it, not as written. Listen for confusion or gaps in understanding. Correct them.
- Auditor package preparation: assemble process documentation, sample incident records (5–10 complete golden files), training records, on-call schedules, alert quality metrics, action tracker register, and status-page archive. Organize by control. Create a table of contents and index.
- Single audit liaison: designate Process Owner or a small dedicated compliance person as sole point of contact for audit requests. Prevents requests scattering across 28 teams.
- Rehearsal: conduct mock interview with an IC and a Comms Lead. Auditors ask tough questions under pressure ("How do you know the timeline is accurate?", "What happens when both ICs are unavailable?", "Show me how you proved the alert was actionable."). Practice answering.
22. Standing governance and continuous improvement (after 20, 21)
Lock in durable improvement. The classic post-audit failure is the process freezing and then decaying. This step prevents that.
- Standing Incident Management Council: chaired by Process Owner, monthly meetings, attendees: CTO, VP Engineering, Head of Support, Head of Compliance, one engineering manager per region, one IC, one team manager from recent wave. Agenda: metrics review, policy changes, gaps from recent incidents, escalation for contentious issues.
- Change control mandate: give Process Owner documented authority to change severity taxonomy, response classes, roles, communications timings, and compensation policy. Any change requires: written justification, steering group approval (monthly cadence), and documented effective date before implementation. This prevents silent drift and ensures changes are deliberate.
- Quarterly validation: re-validate severity and response class taxonomy against all incidents from the prior quarter. Ask: did our taxonomy correctly predict response posture? Did we misclassify? Update taxonomy if patterns emerge.
- Annual metric re-baselining: every 12 months, re-run measurements from S2 (alert census, incident register) to reset targets. System should improve; targets should tighten.
- Resilience roadmap separation: fund a distinct architectural or platform roadmap for incident prevention (reduce shared-database blast radius, multi-region failover, deploy safety, observability investments). Better incident response does not protect a single-ledger corruption or unplanned failover. These are separate problems.
- Quarterly executive report: CTO and VP Eng report to CEO/CFO on metric set (detection time, mitigation time, SLA credits, customer-detected %), top five systemic causes of incidents, program cost vs. credit avoidance, and strategic architecture changes in flight.
- Public backlog of improvement ideas: teams and engineers propose process improvements via Slack or wiki. Process Owner reviews quarterly and implements accepted ideas (e.g., "add a dashboard for detection gaps", "update postmortem template"). Publish what changed and why.
- Celebration and learning: share wins publicly each quarter ("We reduced MTTD from 22 min to 5 min", "Customer-detected incidents down 80%", "$800K SLA credits avoided"). Refresh training and tabletop program annually and immediately after any SEV-1 to keep the system sharp and responsive to new scenarios.
- Median Time to Detect from 22 minutes to <5 minutes for internal detection and <10 minutes for customer journeys by month 6; <2% of customer-impacting incidents first detected by customers (vs. 40% currently) by month 9.
- Median Time to Mitigate for SEV-1 from 3h 10min to <45 minutes by month 9; SEV-2 to <90 minutes by month 9.
- Monthly alert volume from 3,400 to <400 with >95% signal-to-noise ratio by month 5; all alerts satisfy paging contract (symptom-based, owned, runbook-linked, severity-mapped).
- Zero incidents with command ambiguity lasting >15 minutes from month 2 onwards; 100% of SEV-1 and SEV-2 incidents have named IC logged in timeline within 5 minutes; 100% of Triage Owners assigned within acknowledgement.
- 100% of the 28 teams have a documented Service On-Call rotation or a signed service-ownership transfer by month 4; Platform Duty rotation and IC roster both staffed and live by month 3.
- On-call compensation policy published and effective by month 1; on-call satisfaction score ≥7/10 by month 6; zero on-call-attributed voluntary attrition by month 6.
- SLA credits paid from $1.3M annually to <$100K by month 12; credit avoidance (prevented credits) tracked and reported monthly.
- 100% of mandatory postmortems (SEV-0, SEV-1, SEV-2, and repeat incidents) published internally within 15 business days by month 4.
- Postmortem action item completion rate from 17% (11 of 64) to >90% within 60 days by month 6; median action age <30 days; zero repeat incidents caused by the same contributing factor without a design review.
- 100% of the 180 services have a named owner, a detection contract, and at least one symptom-based alert by month 6.
- Status-page first update published within 30 minutes for ≥95% of SEV-1 incidents by month 4.
- 100% of new engineers complete incident-response onboarding within 30 days of joining; IC certification includes written exam and live simulation; ≥2 certified ICs per team group; zero uncovered hours in 24x7 IC roster.
- Weekly operational review held in ≥90% of weeks; 12 of 12 monthly reliability reviews; 4 of 4 quarterly executive reviews; all reviews end with documented decisions and owners.
- SOC 2 Type II audit passes all incident-response controls (CC7.1–7.5, CC2.2–2.3, CC4.1, CC3.x) with zero findings by month 8.
- All pilot and rollout incidents captured with complete golden incident files (timeline, roles, communications, postmortem, actions, closure evidence) by month 3 onwards; audit dry run identifies zero critical gaps by month 7.
- All 28 teams transitioned to new process by month 5; all legacy alert tools decommissioned; single source of truth for incidents established and sustained.
[SYSTEM]
You are an expert assistant in complex project planning.
Your task is to generate a detailed and structured action plan to reach the main objective in the most professional and most detailed way, no matter how much work or steps will be needed to perform.
Use your internal reasoning processes to think deeply about the problem and create the most comprehensive plan possible.
Take as much time and space as you need to think through all aspects of the problem.
After your thorough analysis, answer with the plan in the requested structure.
Every text field you write will be read by a busy person, so write for them: short sentences, one idea each, in short paragraphs. Text fields accept Markdown: separate paragraphs with a blank line, use a bullet list when you enumerate things and plain paragraphs when you explain or argue, and bold at most one key phrase per paragraph. A text of more than three sentences must be split into paragraphs; never deliver a single unbroken block of text.
[HUMAN]
Main Objective: "A B2B payments platform in New York (260 engineers in 28 teams, 2,100 customers, $4B processed a month, a 99.95% SLA with service credits) runs about 180 services on Kubernetes across two AWS regions, with a shared PostgreSQL cluster for the ledger. In the last 12 months: 31 customer-impacting incidents, median time to detect 22 minutes (customers detected 40% of them first), median time to mitigate 3 h 10 min, $1.3M paid in SLA credits, two incidents in which nobody was sure who was in charge for over an hour, and a CEO email about "outages we hear about from clients". On-call exists in 12 of the 28 teams, unpaid, fed by alerts from six different tools — 3,400 alerts a month, 85% of them noise; there is no severity scale, status-page updates are written by whoever is around, and postmortems happen for some incidents, in various formats, with action items that are rarely tracked (11 of 64 closed). A SOC 2 Type II audit in eight months will test incident response. Engineers push back against "carrying a pager for other teams' code".
Define the incident management process: severity levels and what each one triggers; roles (incident commander, communications, scribe, subject-matter responders) and how they are staffed 24x7 across 28 teams; the on-call structure, rotations, compensation and the rules for alert quality; detection and escalation paths; internal and customer communications (status page, account managers, regulators when required) with their timings; postmortems (when mandatory, format, blameless review, ownership and tracking of actions); the metrics and reviews that show whether it works; and how it is introduced across the teams without waiting for the audit."
For your consideration and refinement, here are proposals from the previous round:
Previous Proposal 1 (ID: 2c50b755-66c6-49f5-aea8-767330e89bc3, Agent: claudeHaiku4.5_refine_1, LLM: anthropic/claude-haiku-4-5):
Estimated Complexity: high
Success Metrics: - Median time to detect reduced from 22 minutes to under 5 minutes by month 6, with >90% internal detection (vs. 40% customer-detected now).
- Median time to mitigate reduced from 3h 10min to under 60 minutes for SEV1 and SEV2 by month 9.
- Customer-impacting incidents detected by customers drop from 40% to <5% of all customer-impacting incidents.
- Alert volume reduced from 3,400 per month to <600 per month; signal-to-noise ratio improves from 15:85 to >95:5.
- SLA credits paid reduced from $1.3M to <$100K annually by month 12.
- Zero incidents with command-and-control ambiguity lasting >15 minutes; all SEV1/2 incidents have named IC logged in timeline within 5 minutes.
- Postmortem action item completion rate reaches >80% (from 11 of 64, or 17%) by month 4.
- 100% of SEV1 and SEV2 postmortems published within 15 business days by month 5.
- All 28 teams integrated into incident management system with active on-call rotations by week 20; no team unresponsive to pages for >30 minutes.
- On-call satisfaction score reaches >7/10 on survey; zero on-call-attributed voluntary attrition by month 6.
- Incident commander roster: 40+ certified ICs covering 24x7 with no single point of failure by month 4.
- Status-page first update published within 30 minutes on ≥95% of SEV1 incidents by month 3.
- SOC 2 Type II audit passes incident response controls with zero findings by month 8.
- Weekly incident review cadence sustained in ≥90% of weeks; monthly reliability reviews 12 of 12; quarterly executive reviews 4 of 4.
- 100% of new engineers complete incident-response onboarding within 30 days of joining.
Steps (23):
1. Executive mandate and governance structure
Turn the CEO email into a funded, authorized program with clear ownership and decision rights. Without executive backing, every downstream decision stalls in negotiation.
- Appoint a single program owner (e.g., Director of Incident Management) reporting to the CTO and COO.
- Publish a one-page charter covering scope (all customer-impacting incidents), authority to override team preferences during incidents, and funding for tooling, training, and on-call compensation.
- Establish a standing Incident Management Steering Group with CTO, VP Engineering, VP Support, Head of Compliance, and one engineering manager per region meeting monthly.
- Secure budget envelope: tool licenses, training time, incident-response infrastructure, and on-call compensation (estimated $400–600K annually).
- Make incident-response participation a documented performance expectation for all engineering managers and team leads.
2. Baseline measurement and evidence pack (depends on: 1)
You cannot prove improvement without defensible baseline numbers, and you cannot win arguments about noise or impact without data.
- Build a 12-month incident register: date, detection source, impact scope, time to detect, time to mitigate, SLA credits paid, and services involved.
- Audit the current alert estate: count alerts per tool, per team, per service; compute page-to-action ratio; identify top 50 noisiest rules and off-hours interruptions per engineer.
- Survey on-call engineers and managers on burden, fairness, and escalation clarity (target >70% response rate).
- Interview Support and Account Management: how do customers discover incidents, what do they complain about?
- Document the two command-ambiguity incidents: exactly when unclear who was in charge and why, how long it lasted.
- Publish this pack internally as the problem statement and retain all artifacts for SOC 2 audit evidence.
3. Severity taxonomy and trigger matrix (depends on: 2)
Severity is the keystone. Every other rule—paging, communications, postmortems, compensation—keys off it. Define four levels plus a special SEV0 for security/regulatory events.
- **SEV1 (Critical)**: Complete service outage, data corruption, or >5% payment-path failure rate for >5 min. Every minute costs money. IC required; 99.99% uptime threatened.
- **SEV2 (Major)**: Significant degradation, single region loss, or 1–5% transaction failure. IC typically required; service credit exposure.
- **SEV3 (Minor)**: Limited customer impact with workaround available, or internal issues affecting operations. On-call SME + escalation if SLA at risk.
- **SEV4 (Cosmetic)**: Observations, non-impacting bugs, alerts. Alert-driven, no escalation unless pattern emerges.
- Specify automatic triggers: region loss, ledger write failures, missed settlement window, payment success rate thresholds.
- Define who may declare (any engineer, Support, account manager) and who may downgrade (IC only).
- Include worked examples from the last 12 months so teams recognize their incidents in the definitions.
- Review and re-validate quarterly against actual declarations.
4. Incident roles, command structure, and decision rights (depends on: 3)
The two incidents with >1 hour of command ambiguity prove this step is non-negotiable. Define clear roles with explicit decision authority.
- **Incident Commander**: Owns the incident timeline, not the fix. Declares severity, decides escalation, approves communications, freezes changes, calls responders. Non-technical ICs are acceptable.
- **Deputy IC**: Shadows IC; takes over if IC unavailable. Nominated within 5 minutes of incident declaration.
- **Communications Lead**: Owns internal Slack updates and status-page messaging. Shields IC from interruptions.
- **Scribe**: Records real-time timeline: decisions, who did what, key timestamps. Not responsible for fixing.
- **Subject-Matter Responders**: Engineers with service context. Take IC direction without debate. Report only to IC.
- **Operations Lead** (SEV1 only): Coordinates across multiple responders, manages incident bridge.
- Write one-page role cards with mission, decision authority, and escalation upward. Publish on wiki and link from every paging notification.
- Define minimum viable coverage per severity: SEV1 staffs all roles; SEV2 staffs IC, Comms, Scribe; SEV3 staffs IC + Scribe.
- Establish handover discipline: maximum 4-hour IC shifts on SEV1, written handover template required.
5. Incident tooling consolidation and integration (depends on: 1, 3)
Six alert tools and ad-hoc incident records are structural causes of the 22-minute detection and 3+ hour mitigation. Consolidate to a single incident platform that is the source of truth.
- Select an incident management platform (PagerDuty, Opsgenie, Incident.io, etc.) that supports paging, schedules, escalation, incident records, and postmortem workflow.
- Requirement: the platform must integrate with observability tools, auto-create and pin incident channels in Slack, auto-capture timeline from chat, and support API-driven playbook automation.
- Plan a dual-run period alongside legacy tools with a published cutover date; define rollback criteria.
- Integrate incident record with the 180 services' monitoring and dashboards so responders see everything in one place.
- Define data retention and audit trail to satisfy SOC 2 evidence requirements: who did what, when, under whose authority.
- Budget for licenses, migration effort, and two-week hardening period post-cutover.
6. Alert consolidation and event pipeline (depends on: 5)
Replace six alert sources with a single ingestion point. Deduplicate and route alerts with minimal manual judgment, removing a major source of detection delay.
- Consolidate alert endpoints from six tools into a single event pipeline; this often sits in front of the incident platform (S5).
- Implement deduplication and correlation so a single outage triggering alerts from five monitoring tools produces one page, not five.
- Map every alert to a severity level from S3 (SEV1, SEV2, SEV3, SEV4) at ingestion.
- Log every alert for postmortem analysis and trending.
- Ensure the platform's mobile app works reliably; on-call responders need to engage from any device.
7. Detection strategy: SLOs, signals, and customer-journey monitoring (depends on: 3, 6)
Customers detected 40% of incidents first—a detection gap that must be closed. Build symptom-based alerting that detects outages before customers do.
- Define SLIs and SLOs for the top 20 customer journeys: payment initiation, settlement, ledger read/write, API availability, webhook delivery, measured per region.
- Require symptom-based alerting on SLOs, not cause-based infrastructure metrics (e.g., alert on "payment success rate <99%" not "database CPU >80%").
- Add synthetic transaction monitoring from outside AWS in both regions and a third location, one-minute cadence, for all money-moving paths.
- Add ledger-critical signals: PostgreSQL replication lag, connection saturation, write latency, transaction ID exhaustion, checkpoint pressure.
- Create a detection contract per service: owner identified, at least one symptom alert defined, expected detect time documented.
- Open a customer-reported incident path: Support and account managers can declare an incident directly, counted as a detection source in metrics.
- Fund a separate resilience roadmap to reduce shared-database blast radius, because detection improvements do not protect against ledger corruption.
8. Alert quality standards and noise-reduction program (depends on: 3, 6, 7)
3,400 monthly alerts with 85% noise is the reason engineers resent the pager. Cutting noise is the price of admission for on-call buy-in.
- Publish alert standards: every page must be symptom-based, have an immediate runbook action, be owned by a team, and map to a severity level. No runbook, no page.
- Separate paging from ticketing: only customer-impacting or imminent-impact signals trigger pages. Everything else becomes a dashboard entry or a ticket.
- Set a noise budget per team and service: no service may exceed two pages per on-call shift per month. Breaching triggers a mandatory alert quality review.
- Define the default action for a noisy alert: fix the root cause, tune the threshold, or delete it within ten working days. Deletion is a legitimate successful outcome.
- Implement automatic suppression rules: silence alerts if service auto-recovered within 30 seconds; suppress known maintenance windows; group flapping alerts (>5 in 2 min) into one page; rate-limit noisy services (max 1 alert per 5 min until condition clears).
- Require expiry dates on all silencing rules so suppression cannot become permanent blindness.
- Run a 90-day noise sprint with a visible burn-down of the top 100 noisiest rules, each owned by a named manager.
9. Escalation policies and incident lifecycle (depends on: 3, 4, 5, 6)
Define the mechanical path from alert to incident declaration to resolution. Escalation must be automatic and blameless.
- Define incident lifecycle states with clear entry/exit criteria: Detected → Triaged → Declared → Mitigated → Resolved → Postmortem → Closed.
- Set acknowledgement targets: page acknowledged in 5 minutes; triage decision (is this real?) in 15 minutes; severity declaration (is this customer-impacting?) in 30 minutes.
- Build escalation ladders: if responder does not acknowledge in 5 min, page escalates to service owner, then team manager, then IC on-call. Escalation is automatic, not manual.
- Implement escalation for SEV1: IC paged via phone call + SMS + Slack + mobile; if not acknowledged in 2 min, Deputy IC paged simultaneously; Communications Lead pinged at same time.
- Define change freeze during SEV1 and SEV2: no deployments except to fix the incident. Freeze lifts only when mitigation is confirmed.
- Enforce one incident, one record: the incident record is the sole source of truth. Auto-capture timeline from Slack and bridge; never write timeline from memory later.
- Test all escalation paths weekly via synthetic page to on-call; adjust timings based on first month of operations.
10. On-call architecture and 24x7 coverage model (depends on: 3, 4, 9)
The answer to "carrying a pager for another team's code" is that every team carries its own, and the platform carries shared risk. Design a sustainable model.
- Adopt a federated model: every service has one owning team; that team's on-call carries its service's pager. No team is paged for code it does not own.
- State the consequence clearly: 16 of 28 teams currently have no on-call. They must either build one or formally transfer service ownership to a team that will.
- Create a separate, centrally staffed Incident Commander on-call roster drawn from trained senior engineers across all teams, covering 24x7 with no single point of failure.
- Merge small or low-traffic teams into shared rotations so no rotation has fewer than six engineers (below which coverage is unsustainable).
- Define primary and secondary per slot: secondary engages only on no-acknowledge or explicit request from IC.
- Align coverage across two AWS regions and New York business hours: one global IC rotation; service on-call aligned to service users' time zones.
- Define unresponsive-team escalation: 15 min without acknowledgement escalates to team manager; 30 min escalates to IC, who may direct any available engineer.
- Publish a coverage matrix of all 28 teams showing services, rotation size, and gaps, reviewed monthly.
11. On-call compensation, wellbeing, and sustainability policy (depends on: 10)
Unpaid on-call is the most cited reason for resistance. Settle compensation before rollout, not during it. Make it sustainable.
- Introduce paid on-call: a weekly stipend ($500–1,000) while on-call, regardless of incident volume, benchmarked to New York market rates.
- Pay event-based compensation: 1.5× hourly rate for time spent mitigating incidents during off-hours (minimum one-hour block per callout).
- Provide compensatory rest: no engineer works a normal business day after a night incident requiring >2 hours mitigation. Rest day is documented, not granted as a favor.
- Cap intrusion: define maximum off-hours pages per week (e.g., no more than three per shift). Mandatory review and escalation if exceeded.
- Offer a voluntary opt-out path for engineers with genuine constraints, balanced by explicit obligation that someone else is paid to cover.
- Include on-call expectation and compensation in job descriptions and hiring conversations so commitment is understood before joining.
- Publish the policy with an effective date before any team is asked to join a rotation.
- Review the policy every six months against actual page volumes, attrition rates, and survey feedback.
12. Playbooks and communication templates by severity (depends on: 3, 4)
Playbooks remove ambiguity and decision fatigue during incidents. Templates ensure consistent, compliant messaging.
- Create a one-page (or one-screen) playbook for each severity level: who gets paged (roles, order); first questions (is it real, how big, who knows); what IC declares first (status page text, account manager notification, regulatory trigger); escalation timeline.
- **SEV1 playbook**: Immediate IC + Comms + CTO notification; customer status every 5 minutes; sample message templates.
- **SEV2 playbook**: IC + Comms + tech lead notification; status every 15 minutes; decision tree for escalation to executive team.
- **SEV3 playbook**: On-call SME + Comms if customer-visible; status every 30 minutes or as resolved.
- **SEV4 playbook**: On-call SME only; update customers only if promised SLA is at risk.
- Include decision trees: is this SEV1 or SEV2? Is it our code or dependency? Escalate or containment?
- Prepare customer-communication templates pre-approved by legal and compliance: sample language for detection, impact, workaround, mitigation phases.
- Publish playbooks on wiki and embed links in incident management platform.
13. Internal and customer communications workflows (depends on: 4, 9, 12)
Specify who informs whom, in what order, via what channel. Prevent gaps like "nobody knew who was in charge for an hour."
- **Internal cadence**: First update to #incidents Slack channel within 3 minutes of declaration (even if "investigating"). Updates every 5 minutes (SEV1), 15 minutes (SEV2), 30 minutes (SEV3) or when material change occurs.
- IC calls CTO/VP Eng and incident channel lead within 1 minute of declaration (SEV1); incident declared in Slack with severity, IC name, and service affected.
- SME on-call for the failing service joins incident bridge automatically; escalation call includes them within 5 minutes.
- Designate a single Customer Communications Lead per incident (pre-identified on-call roster) who owns external messaging exclusively. Shields IC from customer contact.
- **Customer communication**: Status page updated within 3 minutes (SEV1) or 10 minutes (SEV2) even if root cause unknown; post "Investigating" with next-update ETA. Account managers of affected top-tier customers called within 5 minutes (SEV1) with templated language.
- Create a phone tree or escalation list accessible to responders; set expectation: "If you don't hear from IC in 2 minutes, call them."
- Use a single incident Slack channel per incident (auto-created by incident tool); log all communications for postmortem review.
- Define regulatory notification path: compliance must approve before sending, but do not wait for root cause; flag incidents triggering payment-processing regulations to legal immediately.
14. Status page, customer notifications, and account-manager playbook (depends on: 5, 12, 13)
Policy without tooling collapses at 3 AM. Make publishing a five-minute action.
- Upgrade or replace status page so components map to customer journeys ("payments", "settlements", "ledger") not internal services. Allow customers to subscribe per component.
- Integrate incident tool (S5) with status page so incident record drives updates and public timeline auto-populates.
- Provide one-click templates pre-filled with severity, impact language, and next-update time; reduce typing and errors.
- Create account-manager playbook: contact tree for top 50 customers, what they may say (facts only), what they must not say (speculation, blame, false ETAs), escalation path if customer escalates.
- Define SLA credit process end to end: impact detection → credit calculation (based on duration × severity) → approval → customer notification → finance treatment. Automate where possible.
- Host status page outside production failure domain on separate infrastructure so it survives total platform outage.
- Test status-page reliability during game days (S17), including simulated status-page outage and total region loss.
15. Postmortem policy, blameless process, and facilitation (depends on: 3, 4)
Only 11 of 64 action items closed means postmortems are currently a writing exercise. Rebuild around learning and tracking.
- Make postmortems mandatory: all SEV1 and SEV2, all SEV3 with customer impact or repeat pattern, any near-miss the IC flags.
- Set deadlines: draft within 5 business days, blameless review within 10 days, internal publication within 15 days.
- Adopt a single standardized template: impact and duration, timeline (detection through resolution), root cause (not "human error" but system condition that enabled error), contributing factors, what went well, what went badly, action items.
- Train a pool of blameless facilitators; require a trained facilitator (not the IC) for every SEV1 and SEV2 review.
- Prohibit counterfactual and blame language in postmortems; require contributing factors addressing tooling, process, organization, and human factors.
- Limit action items to small number of concrete, verifiable, owned changes with dates.
- Publish all postmortems internally by default, with security review only for genuinely sensitive material (e.g., unpatched vulnerability details).
- For SEV1 incidents affecting regulated customers, produce a variant customer-facing root cause report.
16. Action item tracking, reliability backlog, and completion governance (depends on: 15)
A postmortem without durable action tracking is a complaint. Solve the 11-of-64 problem.
- Create a single reliability backlog in the engineering tracker (Jira, Linear, etc.) with mandatory label, owner, due date, and link to originating incident.
- Define closure criteria: evidence required (merged code change, tested alert, verified drill) not self-reported status.
- Protect capacity: reserve a fixed percentage of each team's sprint (e.g., 10%) for reliability work; track unspent capacity and report to VP Engineering.
- Run a weekly ageing review of open actions; escalate anything overdue by >30 days to team lead and VP Engineering.
- Require IC or postmortem facilitator to sign off on action completion.
- Report completion rate and median action age in monthly incident review (target: >90% closed within 60 days).
- If the same service repeats an incident in the same area, trigger a design review rather than another action item; break the cycle.
17. Metrics, dashboards, and review cadence (depends on: 2, 3)
Measure to prove the system works. Publish dashboards so everyone sees the scoreboard.
- Define outcome metrics: time to detect (by source, target <5 min internally detected); time to mitigate SEV1/SEV2 (target <60 min); customer-detected incidents per month (target <2); SLA credits paid (target <$100K/year by month 12).
- Define process metrics: declaration latency, page acknowledgement rate, IC roster coverage (no single point of failure), first-update timeliness (% within SLA), update-cadence adherence.
- Define health metrics: alert volume and noise ratio per team, off-hours pages per engineer per month, postmortem timeliness, action closure rate and median action age.
- Build live dashboards visible to every engineer (not just managers); auto-populate from incident tool, update daily.
- Institute review cadence: weekly operational review (30 min, prior week incidents), monthly reliability review (trends, top causes, action status), quarterly executive review (CEO's office, SLA cost, systemic changes).
- Baseline every metric against S2 evidence pack; set 90-day and 12-month targets.
- End every review with decisions and owners, not just numbers.
18. Training, certification, and exercise program (depends on: 4, 12, 13, 14, 15)
A process that exists only on a wiki fails on the first real page. Build skills before deployment.
- Build curriculum: how to be on-call, how to declare an incident, how to run incidents as IC, how to communicate, how to write blameless postmortems.
- Create role-specific tracks: (1) All engineers: 30-min video on severity levels, communication expectations, postmortem format, action items (mandatory); (2) On-call responders: 2-hour workshop on playbooks (S12), alert tool (S5), escalation paths (S9), when to call manager, case studies, mobile app walkthrough; (3) Incident commanders: 4-hour IC bootcamp on leadership under pressure, decision-making, communicating with executives, status-page discipline, postmortem facilitation, practiced drills; (4) Communications leads: 2-hour training on templates, update timings, how to talk to customers, regulatory rules.
- Require certification before joining IC on-call roster: written assessment plus live simulated incident (pass/fail).
- Record videos so async teams can learn on their schedule; create runbooks and quick-reference cards (print + digital); pair new on-call engineers with experienced responder for first week.
- Run monthly tabletop exercises on realistic scenarios from the prior 12 months: region loss, ledger corruption, cascading failures.
- Run quarterly game days with intentional failure injection (database failover, status-page outage, alert tool downtime); include all on-call roles.
19. Pilot with volunteer teams (depends on: 5, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 18)
Do not roll out untested to 28 teams. Run the entire process end to end with a small cohort first.
- Recruit 3–4 volunteer teams covering a mix of criticality: one payment-path team, one ledger-adjacent team, one platform team, one low-traffic team.
- Run complete process in pilot: new severity scale (S3), consolidated tooling (S5, S6), roles (S4), escalation (S9), communications (S13, S14), postmortems (S15), action tracking (S16), paid on-call (S11), training (S18), metrics (S17).
- Instrument pilot against S17 metrics; compare results with S2 baseline.
- Hold weekly retrospectives with pilot teams; iterate on written policies, tooling, training based on feedback.
- Fix top issues found before wider rollout; document what changed and why.
- Produce pilot report with before/after numbers (MTTD, MTTR, alert noise, action completion rate) to carry into rollout conversations.
- Set explicit pilot exit criteria: rotation coverage achieved, zero unacknowledged pages over 2 weeks, all postmortems delivered on time, >80% of action items tracked.
20. Phased rollout to all 28 teams (depends on: 16, 17, 19)
Rollout is a staged migration with readiness gates, not an email announcement.
- Sequence the 28 teams into four waves of ~7 teams each, ordered by customer-impact criticality and readiness; space waves three weeks apart.
- Define per-team readiness checklist: services mapped and owned, alerts cleaned to standard (S8), runbooks written, rotation staffed, training complete, manager briefed.
- Hold gate review with program owner before each team joins; move unready teams to next wave with dated remediation plan.
- Assign named champion per wave; run internal communications cadence explaining why using pilot numbers from S19.
- Handle resistance directly: publish the "own-your-code, own-your-pager" rule and paid on-call mechanics before each wave, not after.
- Retire legacy tools, informal escalation lists, and ad-hoc status-page process at end of each wave on published cutover date.
- Harvest feedback formally at each wave and push accepted process changes through change control.
21. SOC 2 control mapping and evidence framework (depends on: 1, 3, 13, 15)
The audit will test incident response. Mapping controls now means evidence accumulates from the first real incident under the new process.
- Map the process to Trust Services Criteria for incident identification, response, evaluation, and communication of security incidents.
- Write control statements in auditor language; name a single owner per control.
- Specify the evidence artifact for each control: incident record, severity classification, paging log, communications log, postmortem, action tracker entry, training record.
- Define evidence retention and storage location (not on laptops, not on chat history that expires); plan for audit access.
- Identify which controls are blocked until certain rollout waves complete; keep a gap register with owners and review fortnightly with steering group.
- Run early walkthrough with an experienced compliance partner or pre-audit readiness team to test control design before the audit window.
22. SOC 2 dry run, gap remediation, and audit support (depends on: 20, 21)
Convert a good process into a provable one, a few months before auditors arrive. Prove the system works at scale.
- Schedule a dry run six weeks before audit window, sampling real incidents from completed waves against each control's evidence requirements.
- Remediate every gap found; prioritize anything risking a qualified opinion.
- Brief every engineer who may be interviewed so they can describe the process as practiced.
- Prepare auditor package: process documentation, sample incident records, training records, on-call schedules, action tracking register, alert quality metrics.
- Designate a single audit liaison and small evidence-request team so requests do not scatter across teams.
- Rehearse IC and Communications Lead roles under interview conditions; auditors probe realism under pressure.
- Ensure all postmortems, incidents, and evidence are retained, searchable, and accessible to auditors for the required audit period.
23. Standing governance, process ownership, and continuous improvement (depends on: 20, 22)
The classic post-audit failure is the process freezing and decaying. Lock in continuous improvement as a permanent structure.
- Establish a standing Incident Management Council chaired by the program owner, meeting monthly with engineering, support, compliance, and product representation.
- Give program owner documented mandate to change standards; require formal change control for any change to severity, roles, communications timings, or compensation.
- Re-validate severity taxonomy quarterly against real declarations; re-baseline metrics annually.
- Feed incident themes into architecture review and hiring so the program improves the system, not just the response.
- Report quarterly to executive team on metric set (MTTD, MTTR, SLA credits, customer-detected %) and top five systemic causes of incidents.
- Fund a resilience roadmap separate from incident response: reducing shared-database blast radius, multi-region failover, deploy safety.
- Refresh training and tabletop program annually and after any SEV1.
- Keep public backlog of improvement ideas (Jira, shared doc) so any engineer can suggest; publish monthly "incident newsletter" to all engineers with wins and learnings.
Previous Proposal 2 (ID: d4b0c90c-c556-49ae-945b-5c59cc4fbd11, Agent: deepseek-flash_refine_2, LLM: deepseek/deepseek-flash):
Estimated Complexity: high
Success Metrics: - Median time to detect falls from 22 minutes to under 8 minutes by month 6, and under 5 minutes for money-moving journeys by month 9.
- Share of customer-impacting incidents first detected by customers falls from 40% to under 10% by month 9.
- Median time to mitigate for SEV-1 falls from 3 h 10 min to under 60 minutes by month 9.
- 100% of SEV-1 and SEV-2 incidents have a named Triage Owner within 5 minutes and a named Incident Commander within 15 minutes, from month 4.
- Zero incidents in which command authority is unclear for more than 15 minutes, measured monthly from month 4.
- Monthly page volume falls from 3,400 to under 600, with a false-positive rate under 15%, by month 5.
- No service exceeds two pages per on-call shift for three consecutive months, measured by month 6.
- All 28 teams have a documented rotation, a service ownership map and at least one trained on-call engineer by month 6.
- Incident Commander roster holds at least 12 certified ICs covering 24x7 with no uncovered week, from month 4.
- 100% of on-call shifts are paid under a published policy from month 2.
- 100% of SEV-0, SEV-1 and SEV-2 postmortems are published internally within 15 business days, from month 5.
- Postmortem action items closed within 60 days rise from 17% to over 90%, with median age under 30 days, by month 6.
- At least one incident review or near-miss report is filed per team per quarter, from month 6.
- Status page first update is published within 30 minutes on at least 95% of SEV-1 incidents, from month 4.
- Zero missed regulatory notification windows on any incident requiring notification.
- The customer-impact ledger is complete for 100% of customer-impacting incidents from month 5, with SLA credits reconciled monthly.
- Annual SLA credits paid fall at least 50% from the $1.3M baseline within 12 months.
- Customer-impacting incidents fall at least 40% year over year from the 31-incident baseline.
- On-call satisfaction is at or above 7 out of 10, with zero voluntary attrition attributed to on-call, measured quarterly from month 6.
- Review cadence is sustained: weekly operational review in at least 90% of weeks, 12 of 12 monthly reliability reviews, 4 of 4 quarterly executive reviews.
- SOC 2 Type II is passed with zero findings related to incident response.
- 100% of new engineers complete the incident-response onboarding module within 30 days of joining.
- 100% of the 180 services have a named owner, a detection contract and at least one symptom-based alert held to the paging contract by month 7.
Steps (21):
1. Charter, mandate and the evidence clock
This step turns the CEO's email into a funded programme with one accountable owner and explicit authority, and it starts the SOC 2 clock on day one.
- Appoint a single accountable process owner — a Director of Incident Management reporting to the CTO, with a dotted line to the COO for customer and SLA matters.
- Publish a one-page charter: scope (every customer-impacting, money-moving, security and data-integrity event across all 28 teams and both regions), decision rights during an active incident, and the power to freeze deploys and override team preferences.
- Fix the funding envelope up front: tooling licences, training and drill time, and on-call compensation, with an indicative annual figure and the expected return in avoided SLA credits.
- **Start the evidence clock immediately.** A SOC 2 Type II report covers a period of operating effectiveness, so evidence cannot be backfilled; an audit in eight months means the operating period begins in the first weeks of this programme.
- Stand up a steering group of CTO, VP Engineering, Head of Support, Head of Compliance and two engineering managers, meeting fortnightly.
- Make incident-process participation a documented performance expectation for every engineering manager, not an optional extra.
- Agree the timeline explicitly: a working process in four months, evidence accumulating from month two, audit-ready by month seven.
2. Baseline evidence and problem statement (depends on: 1)
You cannot demonstrate improvement without a defensible starting point, and you cannot win the argument about noise without numbers.
- Build a 12-month incident register and re-classify all 31 customer-impacting incidents: date, severity, class, detection source, time to detect, time to mitigate, customers affected, services involved and SLA credits paid.
- Run an alert census per tool, per team and per service: total volume, page-to-action ratio, off-hours interruptions per engineer, and the 50 noisiest rules with a named owner.
- Build a silent-failure register: incidents in which no internal alert fired at all. This is the number that explains the 40% customer-detected rate.
- Survey every on-call engineer and manager on burden, fairness, escalation clarity and pay expectations, targeting above 70% response.
- Interview Support, Account Management and Sales about how customers learn of incidents, what they complain about, and what they are promised.
- Reconstruct the two command-ambiguity incidents minute by minute, from alert to mitigation, to find exactly where ownership lapsed.
- Publish the pack internally as the problem statement and retain it as management-review evidence for the audit.
3. Control mapping and evidence architecture (depends on: 1)
Most programmes leave compliance to the end. This one maps controls in the first month, because the mapping determines what the process must capture from day one.
- Map the process to the relevant Trust Services Criteria for incident identification, evaluation, response, containment, recovery and communication of events (CC7.1 to CC7.5, CC2.2, CC2.3, CC4.1, CC3.x).
- Write each control as a plain-language statement with one named owner and its evidence artifact: incident record, severity classification, paging log, communications log, postmortem, action tracker entry, training record.
- Define the **golden incident file**: one auto-assembled, single-click export per incident containing timeline, roles, severity, communications, postmortem, actions and closure evidence.
- Set retention, storage location and immutability so no control depends on a laptop, a private Slack channel or chat history that expires.
- Keep a gap register with owners and dates, reviewed fortnightly by the steering group.
- Run an early design walkthrough with the auditor's readiness team inside the first 90 days, to test the design before building on it.
4. Severity times class taxonomy (depends on: 2)
Severity alone is not enough. Severity decides how loud the response is; class decides which playbook runs and what must never be traded away.
- Define severity by impact in customer and money terms: customers affected, dollars at risk, regions lost, ledger integrity, settlement delay, data loss and SLA exposure.
- Use SEV-0 for security, privacy and regulatory events; SEV-1 for total or material loss of a payment path; SEV-2 for degradation or single-region loss; SEV-3 for limited impact with a workaround; SEV-4 for internal-only issues and near-misses.
- Define response classes: Availability, Performance, Data integrity and Ledger, Security and Privacy, Third-party dependency, Settlement and Financial, and Process failure.
- **Class can raise a response, never lower it.** A SEV-2 data-integrity incident gets SEV-1 posture, because integrity failures are not recoverable by moving faster.
- Specify automatic triggers: loss of one AWS region, ledger write failures, replication lag above threshold, a missed settlement cut-off, payment success rate below threshold for five minutes.
- State who may declare: any engineer, Support agent or account manager. State who may downgrade: the Incident Commander alone.
- Map every level to its SLA credit exposure and to the customer-visible status page state.
- Include worked examples from the last 12 months so each of the 28 teams recognises its own incidents, and re-validate the taxonomy quarterly against real declarations.
5. Roles, command and the never-without-an-owner rule (depends on: 4)
The two incidents where nobody was in charge for over an hour did not fail at declaration. They failed in the gap before it, when an alert had fired and no one owned it.
- Create one-page role cards: Triage Owner, Incident Commander, Deputy IC, Operations Lead, Internal Communications Lead, Customer Communications Lead, Scribe, Subject-Matter Responders and Executive Sponsor for SEV-1 only.
- Introduce the **Triage Owner** rule: from the moment a page is acknowledged, that person owns the incident until an IC takes over or the incident is stood down. There is never an unowned minute between first page and close.
- State the core rule plainly: the IC owns the incident, not the fix, and does not debug. An IC who starts troubleshooting has abandoned command.
- Give the IC explicit decision rights: declare and escalate severity, freeze changes, halt deploys, approve customer messaging, and pull in any engineer in the company.
- Define minimum viable staffing per severity: SEV-1 staffs every role; SEV-2 staffs IC, scribe, comms and responders; SEV-3 staffs an IC and a scribe only.
- Set handover discipline: four-hour maximum IC shifts on SEV-1, a written handover template, and a deputy named within 15 minutes of declaration.
- Set behaviour standards for responders: one incident channel, one bridge, no side channels, and asks phrased with a named owner and a time.
- Publish the role cards on the internal wiki and link them from every paging notification.
6. Declaration, lifecycle and escalation policy (depends on: 5)
This step defines the mechanical path from an alert to a declared incident and back to normal service, and it removes judgment calls from the worst moments.
- Define lifecycle states with entry and exit criteria: Detected, Triaged, Declared, Mitigated, Resolved, Postmortem, Closed — plus a Watch state with a hard 30-minute timer, after which the incident is either declared or stood down.
- **Make declaring free.** A declaration that turns out to be a false alarm is closed as a false declaration, with no blame and no follow-up, and it is tracked as a metric so the cost of caution stays visible.
- Apply the ambiguity rule: if two responders disagree about whether this is an incident, it is an incident.
- Set acknowledgement and declaration targets: page acknowledged within five minutes, triage decision within fifteen, declaration within thirty for anything customer-facing.
- Build automatic escalation ladders with timers at every layer: responder, service owner, team manager, IC on-call, VP Engineering. Escalation never requires a human decision and is never criticised.
- Define the unresponsive-team path: fifteen minutes without acknowledgement escalates to the team manager, thirty minutes to the IC, who may then direct any available engineer.
- Define change freeze and rollback authority during SEV-1 and SEV-2, and the single named condition that lifts the freeze.
- Enforce one incident, one record, with the timeline auto-captured from the channel and bridge rather than written from memory afterwards.
- Add the two-simultaneous-incidents rule: a second IC is designated, and no IC ever runs two incidents at once.
7. Detection strategy: journeys, synthetic signals and customer-report intake (depends on: 4)
Customers detected 40% of incidents first. That number is the reason this step exists, and it is fixed by measuring customer journeys rather than infrastructure.
- Define SLIs and SLOs for the top 20 customer journeys — payment initiation, authorisation, settlement, ledger read and write, API availability, webhook delivery and payout — measured per region.
- Require symptom-based alerting on those SLOs instead of cause-based alerting on infrastructure metrics.
- Add synthetic transaction monitoring from outside AWS in both regions plus a third location, on a one-minute cadence for money-moving paths.
- Add ledger and shared-PostgreSQL specific signals: replication lag, connection saturation, write latency, lock waits, transaction ID exhaustion, checkpoint pressure and bloat.
- Open a customer-report intake: Support and account managers can raise an incident directly, every customer report creates an incident record, and the customer-report path is counted as a detection source in all reporting.
- Apply the detection-gap rule: whenever a customer reports an incident before internal monitoring, a detection-gap ticket is opened automatically and owned by the responsible team.
- Publish a detection contract per service: every one of the 180 services needs an owner, at least one symptom alert, and a documented expected detect time.
- Measure current detect time per journey, set targets, and run a detection drill per team: break something in staging and see whether it pages before a human notices.
8. Three on-call rotations across 28 teams (depends on: 5)
The objection is that engineers will not carry a pager for another team's code. The answer is not to argue with it, but to build three rotations so the objection becomes structurally impossible.
- Run a **Service On-Call** rotation per team, covering only that team's own services. No engineer is ever paged for code their team does not own.
- Run a **Platform Duty** rotation for genuinely shared infrastructure: the shared PostgreSQL cluster, Kubernetes, networking, CI/CD and observability. This is nobody's product code, so it gets its own paid rotation, staffed from platform teams plus volunteers from other teams.
- Run a central **Incident Commander** roster of 12 to 16 certified senior engineers drawn from across all 28 teams, covering 24x7 on one-week shifts with a primary and a secondary.
- State the consequence honestly: 16 of 28 teams have no on-call today. Each must either build a rotation or formally transfer ownership of its services to a team that will, with the transfer dated and recorded.
- Merge small or low-traffic teams into shared rotations so no rotation has fewer than six engineers, the level below which coverage stops being sustainable.
- Cap load in the scheduling tool: no engineer is on-call more than two weeks per quarter, enforced by configuration rather than negotiation.
- Publish a coverage matrix for all 28 teams showing services, owners, rotation size and gaps, reviewed monthly.
- Write on-call participation into engineering job levels and hiring criteria, so the commitment is set before someone joins rather than after.
9. Compensation, rest and the price of opting out (depends on: 8)
Unpaid on-call is the most cited reason for resistance. The policy must be settled before rollout, not negotiated during it.
- Move immediately to paid on-call: a per-shift stipend benchmarked to the New York market, agreed with HR and Finance, and published with an effective date before any team is asked to join a new rotation.
- Pay event-based compensation for out-of-hours callouts, with a minimum call-out block and a 1.5x rate for time actually spent mitigating.
- Provide documented compensatory rest: no engineer works a normal day after a night incident, and the rest day is policy rather than a favour granted by a manager.
- Cap intrusion with a maximum number of off-hours pages per shift, and trigger a mandatory review whenever the cap is exceeded.
- Allow opt-out but **put a price on it**: an engineer may step out of a rotation, and their team must buy coverage from the paid pool at a published internal rate. This turns a cultural argument into a visible budget decision.
- Publish an explicit amnesty: incident records, near-miss reports and false declarations are never used in performance reviews. Only failure to report is a performance issue.
- Protect against gaming: pay attaches to shifts and callouts, never to incident counts.
- Review the policy every six months against real page volumes, attrition and survey results.
10. Paging contract and alert quality (depends on: 2, 7)
3,400 alerts a month at 85% noise is why engineers resent the pager. Fixing that is the price of admission for everything else in this plan.
- Publish the paging contract: every page must be symptom-based, actionable, owned, mapped to a severity and a class, and linked to a runbook. **No runbook, no page**, enforced by a CI check on the alert definition.
- Separate paging from ticketing: only customer-impacting or imminent-impact signals may wake a human; everything else becomes a ticket or a dashboard entry.
- Set a page budget per team and per service: a maximum number of pages per on-call shift. Breaching it auto-opens a remediation ticket with the engineering manager as owner.
- Put new alerts on two-week probation: a new rule runs as a ticket only and becomes a pager only after it has proved actionable, so teams stop being woken by untested rules.
- Define the default action for a noisy alert: fix, tune or delete within ten working days, with deletion treated as a legitimate and celebrated outcome and the deleted-alert count published.
- Correlate and deduplicate at the ingest pipeline so one root cause produces one page instead of forty.
- Require an expiry date on every silencing rule and temporary threshold, so suppression cannot quietly become permanent blindness.
- Run a 90-day noise sprint with a public burn-down of the top 100 noisiest rules, each owned by a named manager, and report page-to-action ratio per team monthly.
11. Incident tooling consolidation (depends on: 5, 10)
Six alerting tools and no single incident record is a structural cause of the 22-minute detection and the three-hour mitigation.
- Select one incident platform for paging, schedules, escalation policies, incident records, postmortem workflow and status-page integration, and choose it within four weeks.
- Consolidate the six alert sources into a single event pipeline feeding that platform, with deduplication, correlation and severity and class mapping applied at ingest.
- Integrate observability, runbooks, chat and voice bridges so responders see dashboards and runbooks inside the incident record and the timeline is captured automatically.
- Implement the golden incident file export defined in S3, so audit evidence is a single click rather than a reconstruction.
- Run a dual-run period alongside the legacy tools with a defined rollback, then switch the legacy tools off on a published date.
- Host the status page outside the production failure domain so it survives a total platform outage, and test that during a game day.
- Make the platform usable from a phone: an Incident Commander must be able to run a SEV-1 from a mobile device at 3am.
12. Communications: internal, customer and regulator (depends on: 4, 5)
Today the status page is written by whoever is around. This step replaces improvisation with a clock, a named owner and a pre-cleared template.
- Set the internal cadence: first update within 15 minutes of declaration, then every 30 minutes for SEV-1 and hourly for SEV-2, whether or not there is progress.
- Never let the status page be how an employee learns of an incident: internal communication leads, external follows.
- Set the customer cadence: status page updated within 30 minutes of a SEV-1 and 60 minutes of a SEV-2, with a no-new-information update still mandatory.
- Define the channel hierarchy: status page for all 2,100 customers, direct email to affected customers on SEV-1, and named account-manager calls for the top 50 accounts.
- Prepare templates per severity and class in advance, pre-approved by Legal and Compliance, each with the next-update time built in.
- Forbid speculation: customer communications never guess at cause, never assign blame, and never commit to a root cause before the postmortem.
- Define the closing communication: resolution notice, how SLA credits will be handled, and the committed date for the written report.
- Build a **regulator clock matrix**: for each event type, which regulator, which window, who signs off, and the shortest clock that drives the first action. Cover NYDFS Part 500, money-transmitter and banking notification, security breach notification, card-network rules, and public-company disclosure where applicable.
- Route every regulatory notification through Compliance, never Engineering, and pre-clear the templates.
- Publish a customer-facing root-cause report for SEV-1 incidents, especially for regulated and top-tier accounts.
- Assign a named Customer Communications Lead plus a trained deputy on every SEV-1.
13. Customer-impact ledger and SLA credit automation (depends on: 12)
The $1.3M in credits is a symptom of having no single record of customer impact. This step creates one, and makes it do four jobs at once.
- Maintain one durable customer-impact ledger per incident: which customers, which journeys, from when to when, and the estimated credit.
- Use that single record for customer communications, SLA credit computation, regulatory reporting and the annual review, so nobody reconciles two versions of the same outage.
- Automate the credit path end to end — computation, approval, customer notification and finance treatment — so credits stop being a manual scramble three weeks later.
- Reconcile accrued against paid credits monthly and report the result in the executive review.
- Track credit avoidance against programme cost, so the funding case is a number rather than an argument.
14. Postmortem policy with three artifact levels (depends on: 5)
Postmortems currently happen for some incidents, in various formats. This step makes them mandatory where they matter and light where they do not.
- Make postmortems mandatory for every SEV-0, SEV-1 and SEV-2, every SEV-3 with customer impact or a repeat pattern, every near-miss the IC flags, and every incident in which the process itself failed.
- Use three levels so the ritual matches the weight: a lightweight async incident review for SEV-3 and SEV-4, a standard facilitated postmortem for SEV-2, and a full review with an executive sponsor for SEV-1.
- Set deadlines: draft within five business days, blameless review within ten, internal publication within fifteen.
- Adopt one template: impact, timeline, detection, response, contributing factors across tooling, process, organisation and human factors, what went well, what went badly, and actions.
- Train a pool of blameless facilitators and require a trained facilitator for every SEV-1 review.
- Prohibit counterfactual and blame language in the template, and specifically ban the phrase human error as a root cause — the question is always what made the error possible.
- Publish all postmortems internally by default, with security review only for genuinely sensitive material, and produce a customer-facing root-cause variant for SEV-1.
- Treat postmortems as the learning product of the process, not as a compliance artifact.
15. Action items: capped, verifiable, with a repeat-incident rule (depends on: 14)
Eleven of 64 action items closed is not a tracking problem. It is a generation problem: the process produces more actions than the organisation can absorb.
- **Cap each postmortem at three action items.** Anything beyond three goes into a ranked reliability backlog, not into the postmortem.
- Require each action to have a named human owner, a due date, and a definition of done that is an artifact: a merged change, a new alert that fired in a drill, or a test that fails without the fix.
- Prohibit self-reported closure. Closure requires the artifact, signed off by the process owner or the Incident Commander.
- Keep one reliability backlog in the engineering tracker with a mandatory label and a link back to the originating incident.
- Reserve a fixed percentage of every team's sprint capacity for reliability work, with unspent capacity visible to the vice presidents.
- Apply the repeat-incident rule: a second incident with the same contributing factor triggers a design review owned by an architect, not another action item.
- Run a weekly ageing review of open actions and escalate anything more than 30 days overdue to the VP of Engineering.
- Target more than 90% of actions closed within 60 days and a median age under 30 days, reported monthly by team.
16. Training, certification and drills (depends on: 5, 6, 12, 14)
A process that exists only on a wiki page fails on the first real page. Skills have to be built and tested before they are needed.
- Build a short practical curriculum: how to be on call, how to triage and declare, how to run an incident as IC, how to communicate, and how to write a postmortem.
- Require certification before joining the IC roster: a written assessment plus a live simulated incident.
- Certify at least two ICs per team group, so the central roster has depth across all 28 teams and no holiday week is uncovered.
- Run monthly tabletops drawn from the last 12 months of real incidents, including region loss, ledger write failure and a missed settlement cut-off.
- Run quarterly game days with deliberately injected failure, including shared-PostgreSQL failover, status-page outage and alerting-pipeline outage.
- **Drill the process's own failure modes**, not just technical ones: IC unreachable, comms lead on PTO, two simultaneous SEV-1s, a paging storm, and a false alarm that burns an hour.
- Audit the incident process for single points of failure: who is the only person who can do each critical task, and what happens in their holiday week.
- Add a mandatory onboarding module for every engineer joining or transferring in, with completion records kept in an audit-ready form.
- Track training completion by team and publish it in the monthly reliability review.
17. Metrics and review cadence (depends on: 2, 4, 15)
Establish what good looks like, and measure it in a way that makes people report more incidents rather than fewer.
- Define outcome metrics: time to detect by source, time to mitigate by severity and class, percentage of incidents detected by customers, incidents by severity and class, SLA credits paid, and customer-impact minutes.
- Define process metrics, because the process itself has SLOs: IC assigned within five minutes, page acknowledgement rate, first-update timeliness, update-cadence adherence, postmortem timeliness, action closure rate, action median age and IC roster coverage.
- Define health metrics: alert volume and noise ratio per team, off-hours pages per engineer, on-call satisfaction, training completion, and percentage of services with a detection contract.
- **Never publish incident count as a team metric.** It rewards hiding incidents. Publish reporting metrics instead — near-misses filed, detection gaps closed, false declarations made — alongside the outcome metrics.
- Publish live dashboards visible to every engineer, refreshed daily, with each metric baselined against the S2 evidence pack and given a 90-day and a 12-month target.
- Run a 30-minute weekly operational review incident by incident, a monthly reliability review of trends and actions, and a quarterly executive review with the CEO's office covering customer impact, credits and the top five systemic causes.
- Hold a quarterly review of the process itself: what in the process wasted time, what confused responders, and what should be deleted.
- Require every review to end with decisions and named owners, never with numbers alone.
18. Pilot with three to four teams, using real incidents (depends on: 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17)
Do not roll out to 28 teams untested, and do not rely on drills alone when real incidents are available and better.
- Recruit three to four volunteer teams spanning criticality: a payment-path team, a ledger-adjacent team, the platform team that owns the shared PostgreSQL cluster, and a low-traffic team.
- Run the entire process end to end in the pilot: the new severity and class scale, roles, escalation, communications, postmortems, action tracking and paid on-call.
- Treat real incidents during the pilot as the primary training material. Hold a retro within 48 hours of each pilot incident, run by the process owner, while the friction is fresh.
- Instrument the pilot against the S2 baseline and publish before-and-after numbers.
- Hold weekly retrospectives with the pilot teams and change the written policies, the tooling and the training in response, documenting what changed and why.
- Set explicit exit criteria: rotation coverage achieved, zero unacknowledged pages over a defined period, postmortems delivered on time, and actions tracked to closure.
- Produce a pilot report that every later rollout conversation starts from.
19. Phased rollout sequenced by cost of failure (depends on: 15, 18)
Rollout is a staged migration with readiness gates, not an email announcement. The sequencing matters more than the schedule.
- Sequence the 28 teams into four waves of roughly seven, ordered by **incident density and customer-journey ownership**, not by size or alphabet: the teams whose incidents cost the most SLA credits go first, because their improvement is the visible proof.
- Leave three weeks between waves and define a per-team readiness checklist: services mapped and owned, alerts cleaned to the paging contract, runbooks written, rotation staffed to at least six engineers, training complete, manager briefed, and compensation in effect.
- Hold a gate review with the process owner before each team joins, and move an unready team to the next wave with a dated remediation plan rather than granting an exception.
- Give each wave a named champion and run an internal communications cadence that explains the why using the pilot's numbers.
- Handle resistance directly and early: publish the own-your-code-own-your-pager rule and the paid on-call mechanics before each wave, never after.
- Hold a retro within 48 hours of each wave's first incident under the new process, and feed accepted changes back through change control.
- Retire legacy tools, escalation lists and the informal status-page process at the end of each wave on a published cutover date, so two processes never run in parallel for long.
- Sequence the waves so no team faces both a rollout and the audit dry run in the same week.
20. Audit dry run and evidence review (depends on: 3, 19)
This step converts a good process into a provable one, a few months before the auditors arrive.
- Run an internal dry run six weeks before the audit window, sampling real incidents from the pilot and the early waves against each control's evidence requirements.
- Remediate every gap found, prioritising anything that risks a qualified opinion, and re-test the remediated control against the same sample.
- Brief every engineer who may be interviewed so they can describe the process as they actually practise it, not as it is written.
- Prepare the auditor package: process documentation, sample incident records, paging logs, communication logs, training records, on-call schedules and the action tracking register.
- Designate one audit liaison and a small evidence-request team, so requests do not land on 28 teams simultaneously.
- Rehearse the walkthrough with an incident commander and a communications lead, because auditors probe realism under pressure, not documents.
- Keep the audit liaison and the process owner as the same person where possible, so accountability for the control is also accountability for the evidence.
21. Standing governance, resilience roadmap and continuous improvement (depends on: 19, 20)
The classic post-audit failure is that the process freezes and then decays. This step keeps it alive and separates incident response from the engineering that prevents incidents.
- Establish a standing Incident Management Council chaired by the process owner, meeting monthly with engineering, support, compliance and product representation.
- Give the process owner a documented mandate to change standards, and require formal change control for any change to severity, classes, roles, communication timings or compensation.
- Re-validate the severity and class taxonomy quarterly against real declarations, and re-baseline the whole metric set annually.
- Fund a **resilience roadmap separate from incident response**: reduce the shared PostgreSQL blast radius, improve multi-region failover readiness, and tighten deploy safety. Better incident management cannot save a single shared ledger during a corruption event.
- Report quarterly to the executive team on the metric set and the top five systemic causes of incidents, and feed those themes into architecture review, release engineering and hiring.
- Keep a public backlog of process improvement ideas so any engineer can propose one, and change the process from it every quarter.
- Share wins publicly each quarter — faster detection, fewer customer-detected incidents, credits avoided — to sustain momentum after the audit.
- Refresh the training and tabletop programme annually and immediately after any SEV-1.
Previous Proposal 3 (ID: 985bec06-6c51-44c4-9fe4-d27c2f645d5c, Agent: qwen3.8-flash_refine_3, LLM: alibaba/qwen3.8-flash):
Estimated Complexity: high
Success Metrics: - Median Time to Detect (MTTD) < 10 minutes.
- Median Time to Mitigate (MTTR) < 45 minutes for SEV-1.
- Customer-detected incidents < 5% of total incidents.
- Monthly alert volume < 400 actionable alerts (90% reduction in noise).
- SLA credit payouts < $100k annually.
- Postmortem action item completion rate > 90% within 30 days.
- 100% of SEV-1 incidents have a designated IC and Scribe.
- On-call engineer satisfaction score > 4.0/5.0.
- Zero critical findings in SOC 2 Type II audit regarding incident response.
Steps (20):
1. Program Charter & Executive Mandate
Secure explicit authority to mandate the incident management program, transforming it from an engineering initiative into a business-critical operational standard.
* Appoint a Director of Incident Management (DIM) reporting to the CTO, with a dotted line to the Head of Compliance for audit readiness.
* Publish a charter defining scope: all customer-impacting incidents, payment path failures, and data integrity issues.
* Establish a steering committee (CTO, VP Eng, Head of Support, Legal) to resolve cross-team conflicts and approve tooling budgets.
* Define the 'Own Your Code, Own Your Pager' principle as a non-negotiable engineering value, directly addressing the 28-team resistance.
2. Baseline Audit & Evidence Framework (depends on: 1)
Create a defensible starting point for metrics and a structured evidence repository for the SOC 2 Type II audit.
* Consolidate the last 12 months of incident data into a single register, flagging the 31 customer-impacting incidents.
* Map current alert sources to specific services to quantify the 3,400 monthly alerts and identify the top 50 noisiest rules.
* Conduct a survey of on-call engineers to measure current sentiment and identify specific pain points in the unpaid rotations.
* Designate a secure, immutable storage location for all future incident artifacts (timelines, logs, postmortems) to satisfy audit requirements.
3. Severity Taxonomy & Trigger Matrix (depends on: 1)
Define unambiguous severity levels tied to specific triggers, communication SLAs, and executive notifications to eliminate classification debates.
* **SEV-1 (Critical):** System-wide outage, data loss risk, or >5% revenue-impacting transactions. Triggers immediate IC, VP Eng, and legal notification.
* **SEV-2 (Major):** Significant degradation or region-wide latency. Triggers IC and Team Lead.
* **SEV-3 (Minor):** Localized issue with workaround. Triggers on-call engineer.
* **SEV-4 (Internal):** Low-priority noise. Triggers ticket only.
* Map each severity to specific SLA credit exposures to align technical response with financial risk.
4. Incident Roles & Command Structure (depends on: 3)
Standardize the Incident Command System (ICS) to prevent the 'ambiguity of ownership' that caused previous hour-long delays.
* **Incident Commander (IC):** Single point of decision authority; does not debug. Required for SEV-1/2.
* **Scribe:** Maintains the real-time timeline log; mandatory for SEV-1.
* **Comms Lead:** Owns status page and internal broadcasts; shields IC from external noise.
* **SMEs:** Technical responders focused solely on diagnosis/mitigation under IC direction.
* Create role cards with clear 'Do' and 'Don't' responsibilities for quick reference during chaos.
5. On-Call Architecture & Compensation Policy (depends on: 4)
Establish a sustainable, paid 24x7 on-call model that respects ownership boundaries and compensates for disruption.
* Adopt a federated model: each of the 28 teams maintains its own SME rotation for owned services.
* Create a central IC rotation staffed by senior engineers across teams to ensure command coverage.
* Implement a paid policy: base stipend per shift, enhanced pay for call-outs, and mandatory recovery time after SEV-1 night shifts.
* Define the 'Unowned Service' rule: services without a dedicated team rotation must be transferred to an owning team or decommissioned.
6. Tooling Consolidation & Integration (depends on: 2, 5)
Replace the six fragmented alerting tools with a single Incident Management Platform to ensure a unified source of truth.
* Select a platform (e.g., PagerDuty + Statuspage + Slack/Teams) capable of handling 180 services.
* Integrate Kubernetes and PostgreSQL monitoring to feed alerts into the central platform.
* Automate the creation of incident channels and timelines upon alert acknowledgment.
* Ensure the status page is decoupled from the production environment to remain available during outages.
7. Alert Quality & Noise Reduction Program (depends on: 6)
Implement strict alert hygiene rules to reduce the 3,400 monthly alerts and restore trust in the pager.
* **Rule:** All paging alerts must be symptom-based (customer impact) and have a linked runbook.
* **Rule:** Implement deduplication and rate-limiting at the ingestion layer.
* **Sprint:** Conduct a 90-day noise sprint to fix, tune, or delete the top 100 noisiest rules.
* **Metric:** Enforce a 'Page-to-Action' ratio target; if an alert does not result in human action, it must be demoted to ticket-only.
8. Detection Strategy & SLOs (depends on: 3, 6)
Shift from infrastructure metrics to Service Level Objectives (SLOs) to detect customer-impacting issues faster than clients can.
* Define SLIs for critical payment paths: Transaction Success Rate, API Latency, Settlement Lag.
* Implement synthetic transaction monitoring from external vantage points in both AWS regions.
* Establish error budgets; breach of error budget triggers automatic review and potential feature freezes.
* Integrate customer-reported signals (support tickets, social media mentions) into the detection workflow as low-priority alerts.
9. Communication Protocols & Templates (depends on: 4, 6)
Define who talks, when, and what they say to stop ad-hoc status updates and protect brand reputation.
* **Status Page:** SEV-1 initial update within 15 mins; subsequent updates every 30 mins. Use legal-approved templates.
* **Internal:** IC broadcasts to #exec-leadership for SEV-1 every hour.
* **Regulatory:** Define triggers for NY financial regulator notifications; assign Legal/Compliance responsibility for filing.
* **Client Success:** Provide Account Managers with a 'Known Issues' playbook and real-time access to status page APIs for high-value accounts.
10. Postmortem Framework (Blameless) (depends on: 9)
Standardize the learning process to ensure consistent quality and adherence to the 'Blameless' culture.
* **Mandatory:** All SEV-1 and SEV-2 incidents require a postmortem within 5 business days.
* **Format:** Fixed template including Impact, Timeline, Root Cause, Contributing Factors, and Action Items.
* **Review:** Facilitated by a neutral party (not the IC) to enforce blameless language and prevent finger-pointing.
* **Publication:** All postmortems published internally on the Wiki with full searchability.
11. Action Item Tracking & Governance (depends on: 10)
Close the loop on incident resolution by enforcing rigorous tracking of corrective actions to fix the 17% completion rate.
* Automatically create Jira tickets for every action item identified in the postmortem.
* **Enforcement:** SEV-1 incident cannot be marked 'Closed' until high-priority actions are completed or deferred with VP approval.
* **Review:** Weekly review of overdue actions in the Engineering Leadership standup.
* **Metric:** Track 'Mean Time to Remediation' for action items as a key health indicator.
12. SOC 2 Control Mapping (depends on: 2, 10)
Proactively map the new incident processes to SOC 2 Trust Services Criteria to ensure audit readiness.
* Map S4 (Roles), S9 (Comms), and S10 (Postmortems) to Security and Availability criteria.
* Define 'Evidence of Operation' for each control (e.g., automated timeline logs, signed-off postmortems).
* Identify gaps between current state and audit requirements; assign remediation tasks to the DIM.
* Establish a quarterly internal compliance review to test control effectiveness before the Type II audit.
13. Training & Certification Curriculum (depends on: 4, 6)
Equip all engineers with the skills to operate within the new framework, reducing anxiety and improving response quality.
* **Universal Training:** 1-hour module on severity levels and tools for all engineers.
* **IC Certification:** Mandatory workshop and simulation for engineers joining the central IC rotation.
* **Runbook Review:** Each team must update and validate runbooks for their top 3 critical alerts.
* **Onboarding:** Include incident response basics in the engineering onboarding checklist.
14. Pilot Implementation (Wave 1) (depends on: 7, 8, 9, 13)
Deploy the new process to a controlled subset of high-traffic teams to validate assumptions before broad rollout.
* Select 3 teams: Payments Core, Ledger/API, and one Infrastructure team.
* Run the full cycle for 6 weeks: Alerts, IC handover, Comms, Postmortems.
* Collect feedback daily and adjust playbooks and tooling configurations based on real-world friction.
* Measure success using the baseline metrics from S2 (MTTD, Noise, Satisfaction).
15. Full Rollout Strategy (Waves 2-4) (depends on: 14)
Gradually migrate the remaining 25 teams using lessons learned from the pilot to manage change resistance.
* **Wave 2 (Month 3):** Deploy to 8 remaining critical customer-facing teams.
* **Wave 3 (Month 4):** Deploy to internal platform and data teams.
* **Wave 4 (Month 5):** Deploy to remaining low-traffic teams and legacy services.
* Maintain a 'Help Desk' for new on-call engineers to troubleshoot process questions during their first shift.
16. Simulations & Game Days (depends on: 15)
Test the resilience of the process and the tools under controlled failure conditions.
* **Tabletop Exercises:** Quarterly scenario walks-throughs for leadership and ICs (e.g., 'What if the primary region fails?').
* **Chaos Engineering:** Inject failures in non-production or canary environments to test alert accuracy and runbook validity.
* **Communication Drills:** Simulate SEV-1 to test the speed of status page updates and internal notification paths.
* Document findings in postmortems and create action items for identified weaknesses.
17. Metrics Dashboard & Executive Review (depends on: 2, 15)
Establish a continuous feedback loop to monitor the health of the incident management system.
* Build a real-time dashboard tracking MTTD, MTTR, Alert Volume, and SLA Credits.
* **Weekly:** Operational review of new incidents and action items with the DIM and Team Leads.
* **Monthly:** Executive review with CTO/VP Eng focusing on trends, noise ratios, and on-call fatigue.
* Publish 'Reliability Reports' to all engineers to celebrate wins and share learnings.
18. Culture & Change Management (depends on: 5, 15)
Address the cultural shift required to move from 'hero mode' to 'process mode' and mitigate pushback.
* Highlight success stories where the new process reduced toil or prevented customer churn.
* Enforce the 'Blameless' mantra by publicly correcting any leadership communication that implies blame.
* Recognize and reward effective ICs and engineers who improve runbooks or alert quality.
* Address 'Pager Fatigue' proactively by monitoring call-out volume and rotating engineers off if thresholds are breached.
19. SOC 2 Dry Run & Evidence Prep (depends on: 12, 17)
Conduct a mock audit six months out to identify gaps in evidence retention or process execution.
* Sample 10 recent incidents and verify all required artifacts (timeline, logs, postmortem, actions) are present and accessible.
* Interview on-call engineers to ensure they can describe the process and their roles without hesitation.
* Remediate any 'Control Failures' identified during the dry run.
* Prepare the 'Audit Readiness' package for the external auditors.
20. Continuous Improvement Loop (depends on: 17, 19)
Institutionalize the evolution of the incident process to prevent stagnation.
* Review the Severity Taxonomy quarterly to ensure it reflects current system behavior and customer expectations.
* Update Tooling integrations annually based on vendor roadmaps and new AWS/K8s features.
* Revise Compensation Policy annually based on market data and internal fairness reviews.
* Set annual 'Reliability Goals' for the engineering organization to drive long-term focus.
Please, considering the previous proposals as ideas that could be considered, focus on the main objective and generate an IMPROVED proposal or a completely DIFFERENT perspective if you deem it appropriate. Only if you consider any of them is amazing and impossible to improve, answer with the same proposal.
The plan has these parts:
- "steps": the list of steps, each with "step_id" (a short label unique within this proposal, such as S1, S2, S3), "title" (a short, descriptive title), "description" (what the step does and how, in Markdown: a short opening paragraph and then, if there are several concrete things to say, a bullet list) and "dependencies" (the step_ids of the other steps of THIS proposal that must be completed first; empty if none).
- "estimated_complexity": "low", "medium" or "high".
- "success_metrics": clear and measurable success metrics, one per line as a bullet list.
Think deeply about the problem, considering all aspects and implications. Develop a comprehensive understanding of what needs to be done to achieve the objective, then produce the plan.
{"steps": [{"step_id": "S1", "title": "Executive charter, governance structure, and evidence clock", "description": "Turn the CEO email into a funded, authorized program with clear ownership and documented evidence collection for SOC 2, starting today.\n\n- Appoint a Director of Incident Management reporting to CTO, with dotted line to COO (customer impact) and Head of Compliance (audit readiness).\n- Publish a one-page charter: scope (all customer-impacting, payment-path, data-integrity incidents across 28 teams and 2 regions), decision rights (IC may freeze changes, override team preferences during incidents), and authority to mandate process participation.\n- Secure annual budget for tooling, training, on-call compensation ($500–800K estimated), and resilience work. Connect funding to avoided SLA credits ($1.3M baseline).\n- Establish standing Incident Management Steering Group: CTO, VP Engineering, Head of Support, Head of Compliance, one engineering manager per region. Meet monthly.\n- **Start the SOC 2 evidence clock on day 1.** An audit in eight months means operating-period evidence begins now; design the process to capture evidence continuously, not retroactively.\n- Make incident-response participation a documented performance expectation for all engineering managers and team leads.\n- Publish timeline: working process in month 2, all 28 teams in month 5, audit-ready in month 7.", "dependencies": []}, {"step_id": "S2", "title": "Baseline measurement, incident register, and evidence pack", "description": "Establish defensible baseline metrics and identify structural gaps that explain the 40% customer-detected rate and 22-minute detection time.\n\n- Build a 12-month incident register with all 31 customer-impacting incidents: date, detection source, detection time, mitigation time, customer count, services involved, SLA credits paid, root cause class.\n- Audit the current alert estate: total volume per tool, volume per team, volume per service, page-to-action ratio, top 50 noisiest rules, off-hours interruptions per engineer.\n- Construct a **silent-failure register**: incidents with no internal alert fired at all. This explains the 40% customer-detected rate.\n- Reconstruct the two command-ambiguity incidents minute by minute: exactly when did ownership become unclear, how long, what was the decision bottleneck.\n- Survey on-call engineers (target >70% response): burden, fairness, pay expectations, escalation clarity, willingness to stay.\n- Interview Support and Account Management: how do customers discover incidents, what do they complain about, how do they contact you.\n- Publish the problem statement internally; retain all artifacts for SOC 2 audit evidence. This is the baseline against which all improvements are measured.", "dependencies": ["S1"]}, {"step_id": "S3", "title": "Control mapping and evidence architecture", "description": "Design the process to generate SOC 2-compliant evidence automatically, from the first real incident, so the audit clock ticks in your favour.\n\n- Map the new process to Trust Services Criteria CC7.1–7.5 (incident identification, response, evaluation, containment, communication), CC2.2–2.3 (authorization), CC4.1 (change management), CC3.x (information availability).\n- For each control, write a one-paragraph plain-language statement, name a single owner, and specify the evidence artifact (incident record, timeline, communications log, postmortem, action tracker, training record).\n- Define the **golden incident file**: one single-click export per incident containing severity, timeline, roles assigned, decisions made, communications sent, postmortem, and action items. This is the audit unit.\n- Specify data retention, immutability, access control, and storage location (not laptops, not chat history that expires). Ensure evidence is searchable and organized by incident date.\n- Keep a gap register with owners and dates; review fortnightly in the steering group. Identify which controls are blocked by incomplete rollout and when they unblock.\n- Run an early design walkthrough with an experienced SOC 2 readiness partner inside month 1 to stress-test control design before building on it.", "dependencies": ["S1", "S2"]}, {"step_id": "S4", "title": "Severity and response class taxonomy", "description": "Define four severity levels and four response classes so every decision—paging, communications, postmortem, compensation—keys off a defensible rule, not a judgment call.\n\n- **Severity by impact scope**: SEV-1 (total payment-path loss, data corruption, or >5% transaction failure for >5 min); SEV-2 (significant degradation or single region loss); SEV-3 (limited impact with workaround available); SEV-4 (internal issue or cosmetic); SEV-0 (reserved for security/regulatory/privacy events).\n- **Response class** (orthogonal to severity): Availability, Performance, Data Integrity & Ledger, Security & Privacy. **Key rule: class can raise severity, never lower it.** A SEV-3 data-integrity incident gets SEV-1 response posture because integrity is not recoverable by moving faster.\n- Automatic triggers: loss of one AWS region → SEV-1 or SEV-2 (class-dependent); ledger write failures → SEV-1; replication lag >10s → escalation review; payment success rate <99% for >5 min → SEV-1/2; missed settlement window → SEV-1; total external API unavailability → SEV-1.\n- Who may declare: any engineer, Support, account manager (based on observed customer impact). Who may downgrade: IC only, after investigation.\n- Map each level to SLA credit exposure and to customer-facing status-page state.\n- Include worked examples from the last 12 months so all 28 teams recognize their own incidents in the taxonomy. Re-validate quarterly against real declarations.", "dependencies": ["S2"]}, {"step_id": "S5", "title": "Incident lifecycle, Triage Owner rule, and escalation policy", "description": "Eliminate the \"nobody was in charge for over an hour\" problem by assigning ownership the moment a page is acknowledged.\n\n- Define lifecycle states with clear entry/exit criteria: Detected (alert fired) → Triaged (is this real and customer-impacting?) → Declared (severity assigned) → Mitigated (core issue resolved) → Resolved (all verifications done) → Postmortem (review scheduled) → Closed (action items tracked or dismissed).\n- **Introduce the Triage Owner rule**: the person who acknowledges the page owns the incident until an IC is assigned or the incident is stood down. There is never an unowned gap between first page and declaration. Triage Owner's sole job: decide within 15 minutes whether this requires an IC or a direct stand-down.\n- Set aggressive acknowledgement and declaration targets: page acknowledged in 5 min; triage decision (is this real?) in 15 min; severity declaration in 30 min for any customer-facing incident.\n- Implement automatic escalation ladders with no human judgment required: if responder does not acknowledge in 5 min, escalate to service owner; if no ack in 10 min, escalate to team manager; if no ack in 15 min, escalate to IC on-call. Escalation is never criticized.\n- Define unresponsive-team path: if a service's on-call is unreachable for 30 min, IC may direct any available engineer from any team to engage.\n- For SEV-1 and SEV-2: change freeze until IC declares mitigation confirmed; IC unfreezes changes explicitly.\n- Enforce one incident, one record. Timeline auto-captured from Slack channel and bridge; never written from memory later.", "dependencies": ["S4"]}, {"step_id": "S6", "title": "Incident roles, command structure, and decision rights", "description": "Define clear roles with one-page responsibility cards published and linked from every paging notification.\n\n- **Incident Commander**: owns incident outcome, not the fix. Declares severity, decides escalation, approves all customer communications, freezes changes, calls responders, hands off in shifts. Non-technical ICs are acceptable; technical depth is not required.\n- **Deputy IC**: assigned within 15 min of declaration; shadows IC; takes over if IC unavailable or after 4-hour shift on SEV-1. Maximum IC shift: 4 hours on SEV-1, 6 hours on SEV-2.\n- **Triage Owner** (new role): owns incident from first page acknowledgement until IC takes over or stand-down decision is made. Required for all incidents.\n- **Communications Lead**: owns internal Slack updates and status-page messaging; shields IC from customer contact and interruptions.\n- **Scribe**: records real-time timeline with decisions, actions, and key timestamps; not responsible for fixing.\n- **Subject-Matter Responders**: engineers with service context; take IC direction; report only to IC; no side channels or parallel debugging.\n- **Operations Lead** (SEV-1 only): coordinates multiple responders, manages incident bridge, maintains escalation list.\n- Minimum viable staffing: SEV-1 requires all roles; SEV-2 requires IC, Deputy, Comms, Scribe, SMEs; SEV-3 requires Triage Owner and IC.\n- Create laminated role cards for every on-call shift location (office, home, printed in pockets).", "dependencies": ["S5"]}, {"step_id": "S7", "title": "Three on-call rotations: Service, Platform, and Incident Commander", "description": "Directly address the \"carrying a pager for another team's code\" objection by making it structurally impossible.\n\n- **Service On-Call rotation** (federated): each of the 28 teams maintains a rotation for their own services only. No engineer is paged for code their team does not own. The answer to \"why am I carrying a pager?\" is now simply: \"for your team's code.\"\n- **Platform Duty rotation** (centrally staffed): shared PostgreSQL cluster, Kubernetes, networking, CI/CD, observability, and incident management tooling. Nobody's product code, so it gets its own dedicated rotation. Staffed from platform teams plus volunteers from other teams; paid at premium rate.\n- **Incident Commander roster** (24x7): 12–16 certified senior engineers from across all 28 teams, on one-week primary shifts with secondary backup. Covers every hour with no single point of failure and no uncovered holiday week.\n- **Consequences and gates**: 16 of 28 teams have no on-call today. Each must either (a) build a Service On-Call rotation of at least 6 engineers, or (b) formally transfer service ownership to a team that will, with transfer documented and dated. No exceptions, no waivers. Unowned services are decommissioned or transferred by end of month 2.\n- Merge small or low-traffic teams into shared rotations where service ownership is unclear (e.g., shared analytics, testing infrastructure).\n- Enforce scheduling limits in the tooling: no engineer on-call more than 2 weeks per quarter, automatically enforced by configuration, not negotiation.\n- Publish a coverage matrix for all 28 teams showing services, owners, rotation size, gaps, and monthly status.", "dependencies": ["S5", "S6"]}, {"step_id": "S8", "title": "On-call compensation, rest policy, and sustainability", "description": "Settle compensation before rollout, not during negotiations. Make on-call sustainable and valued.\n\n- **Paid on-call**: effective immediately upon joining a rotation. Weekly stipend while on shift (benchmark to New York market: $600–1,000 per week per engineer), regardless of incident volume.\n- **Event-based compensation**: 1.5× hourly rate for time spent mitigating out-of-hours incidents, minimum one-hour block per callout. Tracked by incident record (auto-capture from timeline).\n- **Compensatory rest**: no engineer works a normal 8-hour business day after a night incident requiring >2 hours mitigation. Rest day is documented policy, not a favour granted by manager.\n- **Intrusion cap**: maximum 3 unscheduled pages per week per engineer. Exceed the cap in a week and trigger an immediate review; exceed in a month and escalate to VP Engineering. Breaches are structural signal that alert quality or service stability has a problem.\n- **Voluntary opt-out**: an engineer may exit a rotation; their team must hire or buy replacement coverage from paid pool at published internal rate ($X per shift). This converts culture debate into visible budget decision.\n- **Amnesty policy**: incident records, near-miss reports, and false declarations are never used in performance reviews or compensation discussion. Only failure to report is a performance issue.\n- **Policy publication**: publish compensation structure and effective date before any team is asked to join a rotation, and include on-call expectations in job descriptions and hiring conversations.\n- **Semi-annual review**: reassess compensation and caps every six months against actual page volumes, attrition rates, and survey feedback.", "dependencies": ["S7"]}, {"step_id": "S9", "title": "Detection strategy: SLOs, synthetic monitoring, and customer-report intake", "description": "Close the 40% customer-detected gap by monitoring customer journeys instead of infrastructure metrics.\n\n- **SLO-based alerting**: Define SLIs and SLOs for the top 20 customer journeys (payment initiation, authorization, settlement, ledger read/write, API availability, webhook delivery, payout). Measure per region. Alert on SLO breach, not on infrastructure metric (e.g., alert on \"payment success rate <99%\" not \"database CPU >80%\").\n- **Synthetic transaction monitoring**: deploy synthetic transactions from outside AWS in both regions plus a third geographic location, one-minute cadence, for all money-moving paths. These are your first alarm bell.\n- **Ledger-critical signals**: PostgreSQL replication lag (target: <1s, alert >5s), connection saturation, write latency (p95), lock-wait time, transaction ID exhaustion proximity, checkpoint pressure, table bloat. These are separate alerts on shared-database health.\n- **Customer-report intake** (new detection channel): Support and Account Managers can raise an incident directly in the platform. Every customer report creates an incident record automatically, and the \"customer report\" detection source is counted in all metrics. This is a legitimate detection method, not a failure.\n- **Detection-gap rule**: whenever a customer reports an incident before internal monitoring fires, auto-create a ticket in the owning service's backlog with root cause: \"Monitoring gap on [journey].\"\n- **Detection contract per service**: every one of the 180 services needs a named owner, at least one symptom-based alert mapped to a SLO, and a documented expected detect time (target: <5 min for payment path, <10 min for others). Published on wiki and reviewed monthly.\n- **Detection drills**: run a quarterly drill per team: simulate a broken service in staging and verify it triggers a page before a human notices.\n- **Resilience roadmap separation**: detection improvements do not protect against ledger corruption or multi-region failure. Fund a separate resilience roadmap to reduce shared-database blast radius and improve failover safety.", "dependencies": ["S4", "S7"]}, {"step_id": "S10", "title": "Alert quality standards and noise-reduction program", "description": "Cut the 3,400 monthly alerts (85% noise) to <600 with 95% signal. This is the price of admission for on-call buy-in.\n\n- **Paging contract**: every page must satisfy all of (1) symptom-based (customer impact, not infrastructure cause), (2) actionable (linked runbook with immediate next step), (3) owned (named team responsible), (4) severity-mapped (SEV-1/2/3/4), (5) SLO-linked where applicable. **No runbook, no page.** Enforce with CI check on alert definition.\n- **Separation rule**: only customer-impacting or imminent-impact signals trigger pages. Everything else becomes a dashboard entry, a ticket, or a log line. Noisy infrastructure metrics go to dashboards, not pagers.\n- **Page budget per service**: no service may exceed 2 pages per on-call shift per month. Exceeding budget auto-opens a remediation ticket in the owning team's backlog (with alert-quality review assigned to tech lead).\n- **Automatic suppression rules**: (1) silence alerts if service auto-recovered within 30s, (2) suppress known maintenance windows, (3) group flapping alerts (>5 in 2 min) into one page, (4) rate-limit noisy services (max 1 page per 5 min until condition clears). All suppression rules must have an expiry date; no permanent silence without a ticket.\n- **Probation for new alerts**: new alert rules run as tickets only and alert to a Slack channel; after two weeks of proving actionability (every alert resulted in human action), they graduate to pager.\n- **Noise sprint**: run a focused 90-day program with a public burn-down of the top 100 noisiest rules. Assign each to a named manager. Default action: fix root cause, tune threshold, or delete within 10 working days. Deletion is a legitimate successful outcome (celebrate it).\n- **Correlation and deduplication**: consolidate alert sources at ingest pipeline so one outage triggering 40 alerts produces one page, not 40.\n- **Alert ownership**: every alert must have an owning team and a maintenance contact. Update monthly.", "dependencies": ["S9"]}, {"step_id": "S11", "title": "Incident tooling consolidation and integration", "description": "Replace six alert tools and ad-hoc incident records with a single source of truth that unifies paging, escalation, timeline, and audit evidence.\n\n- **Tool selection**: choose an incident-management platform (e.g., PagerDuty, Incident.io, Opsgenie) that integrates paging schedules, escalation policies, incident records, postmortem workflow, and status-page APIs. Decision gate: month 1.\n- **Event pipeline consolidation**: route all alerts from the six legacy tools into a single event pipeline that feeds the incident platform. Apply deduplication, correlation, severity/class mapping, and rate-limiting at ingest.\n- **Observability integration**: connect the incident platform to your Kubernetes dashboards, PostgreSQL monitoring, distributed tracing, and logs so responders see context in one pane. Link runbooks directly into incident records.\n- **Slack and bridge integration**: auto-create incident Slack channels, auto-invite roles, auto-capture timeline from channel transcript and voice-bridge recording. Timeline is not written from memory; it is auto-captured.\n- **Golden incident file**: implement the export defined in S3. One click produces a complete, immutable, audit-ready PDF: severity, timeline, roles, decisions, communications, postmortem, action items, and closure evidence.\n- **Dual-run period**: run both legacy and new platform in parallel for two weeks. Define rollback criteria (e.g., ", "dependencies": ["S5", "S10"]}, {"step_id": "S12", "title": "Escalation automation and incident lifecycle enforcement", "description": "Eliminate judgment calls from the worst moments. Escalation is automatic, mechanical, and blameless.\n\n- **Automatic escalation ladders**: page responder → if no ack in 5 min, page service owner → if no ack in 10 min, page team manager → if no ack in 15 min, page IC on-call + call them immediately (phone + SMS + Slack). No human decides to escalate; timers fire escalations.\n- **Severity-based escalation tempo**: SEV-1 uses faster timers (2 min for IC on-call), SEV-2 uses moderate timers (5–10 min), SEV-3 uses slower timers (15–30 min). Configured in tooling, reviewed quarterly.\n- **Dual IC rule**: if a second SEV-1 incident is detected while the first is active, immediately page and assign a separate IC. ICs never run two incidents in parallel.\n- **Change freeze and rollback authority**: SEV-1 and SEV-2 trigger automatic deploy freeze. Only the IC (with CTO/VP Eng notification) may unfreeze. Freeze lifts only when IC explicitly declares mitigation confirmed and verifies no new incident symptoms for 5 min.\n- **Unresponsive team escalation**: if service's on-call does not acknowledge in 30 min, IC may direct any engineer from any team (volunteers first, then rotated) to engage. This is documented and reported in monthly review (escalation = signal of rotation problem).\n- **One incident, one record**: all decisions logged in the incident platform. Auto-capture from Slack, bridge, status-page updates. Timeline is the source of truth; postmortem is written from timeline, never constructed after the fact.\n- **Ambiguity rule**: if two responders disagree about whether an incident should be declared, it is declared. False declarations (stand-downs within 30 min of declaration) are tracked as metrics and closed without blame.\n- **Watch state**: an unconfirmed incident can live in \"Watch\" state for max 30 min; after that, either declare it or stand it down explicitly.", "dependencies": ["S5", "S11"]}, {"step_id": "S13", "title": "Internal, customer, and regulatory communications workflows", "description": "Define who informs whom, in what order, via what channel, with explicit timings and pre-approved templates.\n\n- **Internal cadence**: first update to #incidents Slack channel within 3 min of declaration (even if \"Investigating\"). Then updates every 5 min (SEV-1), 15 min (SEV-2), or 30 min (SEV-3), or immediately on material change (e.g., mitigation achieved, scope widened). **Comms Lead owns the update; IC must not be interrupted.**\n- **Executive notification**: IC calls CTO and VP Eng within 1 min of SEV-1 declaration (not email, not Slack, call). Incident declared in Slack with severity label, IC name, and affected service. Escalation channel lead auto-pinged.\n- **Customer communication channels**: status page (all 2,100 customers), direct email to affected customers (top-tier accounts and customers affected by SEV-1), account-manager calls (top 50 accounts on SEV-1).\n- **Status page timings**: update within 3 min of SEV-1 declaration, 10 min of SEV-2, 30 min of SEV-3 (even if root cause unknown; use \"Investigating\" with next-update ETA). Updates every 5–30 min depending on severity. Always include next-update time.\n- **Pre-approved templates**: draft customer-facing language for each severity and class in advance with Legal and Compliance. Templates specify impact language (\"some of your transactions are delayed\" not \"our database failed\"), workarounds if available, and next-update commitment. Never speculate on cause in customer communication.\n- **Regulatory notification path**: identify incidents requiring regulator notification (NYDFS Part 500, money-transmitter rules, payment-card-network rules, securities disclosure). Build a clock matrix: event type → regulator → notification window → signer. Compliance owns all regulatory notifications (never Engineering). Pre-clear templates. Flag incidents to Compliance immediately upon declaration.\n- **Account manager playbook**: contact tree for top 50 accounts, templated talking points (facts only, never speculation), escalation path if customer escalates, what to offer (service credit, technical deep-dive call).\n- **Closing communication**: resolution notice, SLA credit impact, commitment date for written root-cause report, customer action required (none, or security update, etc.).", "dependencies": ["S6", "S12"]}, {"step_id": "S14", "title": "Status page infrastructure and customer-impact ledger", "description": "Make the status page reliable, customer-centric, and audit-ready. Track customer impact in a single durable record.\n\n- **Status page decoupling**: host status page outside production failure domain (separate cloud, separate infrastructure, separate database). Integrate incident platform with status page so incident record drives all public updates. Status page survives total platform outage.\n- **Component-to-journey mapping**: status page components map to customer journeys (\"Payments\", \"Settlements\", \"Payouts\", \"Ledger API\") not to internal services. Allow customers to subscribe to components; notify by email or webhook.\n- **One-click update templates**: pre-fill status-page template with severity, impact language, next-update time, and estimated resolution. Comms Lead types minimal new info (\"Root cause identified\" or \"Workaround available\"), and updates auto-post.\n- **Customer-impact ledger** (one record per incident): which customer accounts affected, which journey(s) impacted, exact start and end time of impact, estimated SLA-credit exposure. Use this single record for customer communications, credit computation, regulatory reporting, and annual review. No reconciliation of two versions of the same outage.\n- **SLA credit automation**: compute credit based on duration × severity × customer tier → auto-generate customer notification → auto-post to finance system. Reconcile accrued vs. paid credits monthly and report in executive review.\n- **Testing during game days**: simulate status-page outage and verify alerts continue to fire; test total region loss and confirm status page remains updated; drill runbook for manually updating status page if platform is down.", "dependencies": ["S13"]}, {"step_id": "S15", "title": "Postmortem policy: mandatory, blameless, three-level framework", "description": "Turn postmortems from a writing exercise (11 of 64 action items closed) into the learning engine of the system.\n\n- **Mandatory postmortems**: all SEV-0, SEV-1, and SEV-2 incidents; all SEV-3 with customer impact or repeat pattern; any near-miss IC flags; any incident where the process itself failed (IC unreachable, Comms Lead unavailable, false declaration, missed update SLA).\n- **Three-level framework** (proportionate to weight): (1) lightweight async review for SEV-4 and low-impact SEV-3 (10 min template in shared doc, owner + IC review), (2) standard facilitated postmortem for SEV-2 and impactful SEV-3 (full template, facilitated by trained neutral party, published within 10 days), (3) full executive postmortem for every SEV-1 and every security incident (executive sponsor assigned, full investigation, published within 15 days, customer-facing variant prepared).\n- **Fixed timeline**: draft postmortem within 5 business days, blameless review within 10 days, internal publication within 15 days.\n- **Single template**: impact (who, how many, how long, financial exposure), timeline (detection through resolution), root cause (not \"human error\" but system condition that enabled error; what was the gap?), contributing factors (tooling, process, organization, knowledge, monitoring), what went well, what went badly, action items (≤3, rest go to reliability backlog).\n- **Blameless facilitation**: train a pool of blameless postmortem facilitators (target: 10+ engineers). Require a trained, neutral facilitator for every SEV-1 and SEV-2 review. Prohibit counterfactual language (\"if the engineer had\"), blame language, and the phrase \"human error\" as a root cause.\n- **Publication rule**: publish all postmortems internally by default; security review only for genuinely sensitive material (e.g., unpatched vulnerability details or customer PII in logs). Create a customer-facing root-cause report for every SEV-1, especially for regulated customers, with legal and compliance sign-off.\n- **Searchability**: store postmortems in a searchable wiki or issue tracker with tags (service, class, root cause category) so teams can learn from similar incidents without repeating them.", "dependencies": ["S6", "S13"]}, {"step_id": "S16", "title": "Action item tracking, reliability backlog, and repeat-incident design rule", "description": "Close the loop on incident learning by enforcing verifiable, tracked action items and breaking cycles of repeat incidents.\n\n- **Action item capping**: each postmortem generates a maximum of 3 action items. Anything beyond 3 goes into a ranked reliability backlog, not into the postmortem, to prevent overwhelming teams.\n- **Action item requirements**: each item must have (1) a named human owner (not a team), (2) a due date (≤60 days, target ≤30 days), (3) a definition of done (merged code change, tested alert, audit evidence, architectural decision, new runbook, training completed) not self-reported status.\n- **Single reliability backlog**: create one backlog in your engineering tracker (Jira, Linear, etc.) with mandatory label (e.g., `incident-action`), link to originating incident, and link to postmortem. Track progress weekly.\n- **Closure sign-off**: Incident Commander or postmortem facilitator must sign off on closure, verifying artifact exists (code merged, alert tested in drill, runbook verified).\n- **Repeat-incident rule**: if the same service or component has a second incident with the same contributing factor, **do not create another action item**. Instead, escalate immediately to an architect or tech lead and trigger a design review (not a task, a review). This breaks the cycle of repeated patches; the system needs a structure change.\n- **Capacity protection**: reserve a fixed percentage of each team's sprint capacity (10–15%) for reliability work. Track unspent capacity and report to VP Engineering monthly; if a team is not spending it, work with them to identify and fix blockers.\n- **Ageing and escalation**: run a weekly review of open actions; escalate anything >30 days overdue to team lead and VP Engineering. Monthly report: completion rate (target >90% within 60 days) and median action age (target <30 days).", "dependencies": ["S15"]}, {"step_id": "S17", "title": "Training, certification, and exercise program", "description": "Build skills before deploying the process. Run ongoing drills so the system is tested, not guessed at.\n\n- **Curriculum**: (1) All engineers (30-min async video): severity taxonomy, communication expectations, postmortem format, when to declare an incident, where to find runbooks. (2) On-call responders (2-hr workshop): alert tool walkthrough, playbooks by severity, escalation paths and timers, when to call manager, mobile app walkthrough, case studies from the last 12 months. (3) Incident Commanders (4-hr bootcamp + test): leadership under pressure, decision-making (severity, escalation, rollback), communicating with executives, status-page discipline, postmortem facilitation, handling ambiguity, live simulated incident (pass/fail certification). (4) Communications Leads (2-hr training): templates per severity and class, customer-communication rules (no speculation, no blame), update timings, how to shield IC, regulatory triggers.\n- **IC certification**: written assessment (75% pass required) plus live simulated incident (role-play with facilitator, graded on severity declaration, escalation decisions, communication, handover). Certification valid for 12 months; recertify via annual refresher or another live sim.\n- **Depth across teams**: certify at least 2 ICs per team or team group so central roster is not siloed in one group; no holiday week is uncovered.\n- **Async content**: record all training videos so async teams can learn on their schedule. Create quick-reference cards (laminated, pocket-sized) for roles and playbooks; distribute to on-call locations (office, home).\n- **Monthly tabletop exercises**: drawn from real incidents from the last 12 months (region loss, ledger write failure, missed settlement window, cascading failures). Facilitator describes scenario; 3–4 responders play out response (Triage Owner, IC, Comms) as if real. Run 30 min; retro for 15 min afterward.\n- **Quarterly game days**: deliberately inject failures into production (database failover, status-page outage, alerting-pipeline outage, dual SEV-1 incidents). All on-call roles engage. Run 2–3 hours; measure response times, decision quality, and communication. Document findings and create action items for identified gaps.\n- **Drill the process's own failure modes**: IC unreachable (on-call unavailable, phone broken), Comms Lead on PTO, two simultaneous SEV-1s, paging storm (100+ alerts), false alarm that consumes an hour. Test escalation paths, deputy takeover, and recovery.\n- **New-engineer onboarding**: add incident-response module to all engineering onboarding (completion tracked, audit-ready). All engineers must complete within 30 days of joining or transferring in.", "dependencies": ["S6", "S13", "S15", "S16"]}, {"step_id": "S18", "title": "Metrics, dashboards, and review cadence", "description": "Measure to prove the system works. Publish live dashboards so every engineer sees the scoreboard and the system is transparent.\n\n- **Outcome metrics**: Median Time to Detect by source (target: <5 min internally detected, <10 min customer journeys); Median Time to Mitigate for SEV-1/2 (target: <60 min SEV-1); customer-detected incidents as % of total (target: <5%); incidents by severity (should be mostly SEV-3/4, few SEV-1); SLA credits paid (target: <$100K/year by month 12); annual credit avoidance vs. program cost.\n- **Process metrics**: IC assigned within 5 min (target: >95% of incidents); page acknowledgement rate (target: >98% within 5 min); first-update timeliness (target: >95% within SLA); postmortem timeliness (target: 100% of mandatory postmortems published on time); IC roster coverage (zero uncovered hours, monitored weekly).\n- **Health metrics**: alert volume and signal-to-noise ratio per team (trending toward target); off-hours pages per engineer per month (trend, cap enforcement); on-call satisfaction survey (target: >7/10); training completion by team (target: 100% within 30 days); % of services with active detection contract (target: 100%).\n- **Never publish incident count as a team metric.** Reward hiding. Instead publish detection metrics (near-misses reported per team, detection gaps closed, false declarations made).\n- **Live dashboards**: build dashboards visible to all engineers (not just managers) showing outcome, process, and health metrics. Auto-populate from incident platform and alert tool. Update daily. Link from Slack and internal wiki.\n- **Baseline all metrics against S2 evidence pack.** Set 90-day and 12-month targets for each metric. Publish targets and progress monthly.\n- **Review cadence**: (1) weekly 30-min operational review (incident by incident from prior week: what went well, what hurt, actions); (2) monthly 60-min reliability review (trends, top causes, action-item aging, alert quality per team); (3) quarterly 60-min executive review (CEO's office: customer impact, SLA credits, top five systemic causes, program ROI).\n- **Quarterly process review**: what in the process wasted responder time, what confused people, what should be deleted. Solicit feedback from ICs, Comms Leads, and responders. Document changes and reasoning.", "dependencies": ["S2", "S12", "S16", "S17"]}, {"step_id": "S19", "title": "Pilot program with 3–4 volunteer teams", "description": "Do not roll out untested to 28 teams. Run the entire process end-to-end with a small cohort using real incidents as the primary training material.\n\n- **Team selection**: recruit 3–4 volunteers spanning criticality: one payment-path team, one ledger-adjacent team, one shared infrastructure team (platform or Kubernetes), one low-traffic team. Volunteers see early adoption and influence.\n- **Full process in pilot**: new severity and class taxonomy (S4), consolidated tooling (S11), roles and Triage Owner (S5–6), three rotations (S7), escalation automation (S12), communications (S13–14), postmortems (S15), action tracking (S16), paid on-call (S8), training (S17), metrics (S18). This is not a partial test; it is the complete system.\n- **Real incidents are the training**: hold a retro within 48 hours of each pilot incident (while memory is fresh). Process Owner facilitates. Discuss: what worked, what hurt, how is the runbook, is the alert tuned, did Comms template work, did roles work, was timeline auto-captured correctly. Document feedback and iterate.\n- **Instrument against baseline**: measure all S18 metrics in pilot; compare results with S2 baseline. If MTTD dropped from 22 min to 8 min, that is the number you carry into the next conversation.\n- **Weekly retros with pilot teams**: structured 30-min sessions reviewing policies, tooling, training, templates, and process changes. Change the written policy or tooling in response; document what changed and why so later teams see the evolution, not just the finished state.\n- **Explicit exit criteria**: (1) rotation coverage sustained (≥6 engineers per rotation, zero unacknowledged pages over 3 weeks), (2) postmortems delivered on time (100% of mandatory postmortems published within 15 days), (3) action tracking established (100% of action items in backlog with owner and date), (4) metrics live (dashboards updated daily, first weekly review completed).\n- **Pilot report**: document before/after numbers (MTTD, MTTR, alert noise, action completion rate, on-call satisfaction) and key process learnings. This report is the foundation for every conversation in the rollout.", "dependencies": ["S6", "S8", "S11", "S12", "S13", "S14", "S15", "S16", "S17", "S18"]}, {"step_id": "S20", "title": "Phased rollout sequenced by cost of failure", "description": "Rollout is a staged migration with readiness gates and sequencing that prioritizes visible impact, not ease.\n\n- **Wave sequencing**: divide 28 teams into 4 waves of ~7 teams each, **ordered by incident density and customer-journey ownership** (highest-cost-of-failure teams first). Teams with the most SLA credits at stake go first; their improvement is the proof.\n- **Wave spacing**: three weeks between waves. This gives each wave time to stabilize and find problems before the next cohort joins.\n- **Readiness checklist per team**: (1) all services mapped and owned (no unowned services), (2) alerts cleaned to paging contract (runbook linked, severity mapped), (3) playbooks updated and tested in staging, (4) rotation staffed to ≥6 engineers, (5) team completes training module, (6) manager briefed on policy, (7) on-call compensation in effect.\n- **Gate review before each wave**: Process Owner holds gate review with target teams. Move unready teams to next wave with a dated remediation plan. No exceptions, no waivers; readiness is non-negotiable.\n- **Wave champion**: assign a named engineer per wave to champion the rollout, answer questions, escalate issues to Process Owner. Champions are not representatives; they are advocates and feedback collectors.\n- **Communication cadence**: weekly all-hands or newsletter for 4 weeks before each wave. Explain why (owned-code-owned-pager rule, paid on-call, SLA credit savings). Use pilot numbers. Answer FAQs. Announce champion and escalation path.\n- **First incident under new process**: hold a retro within 48 hours. Feed accepted process changes back through change control.\n- **Retire legacy tools and processes**: at end of each wave, retire legacy alert tools, informal escalation lists, ad-hoc status-page process. No parallel processes running for >3 weeks; this prevents confusion and half-learning.\n- **Sequence to avoid audit collision**: ensure no team is rolling out in the same week as the audit dry run (S21).", "dependencies": ["S16", "S18", "S19"]}, {"step_id": "S21", "title": "SOC 2 dry run and evidence review", "description": "Convert a good working process into a provable one. Test control evidence a few months before auditors arrive, when you can still fix gaps.\n\n- **Dry run timing**: run 6 weeks before audit window (around month 7 of this program).\n- **Scope**: sample 10–15 real incidents from pilot and early rollout waves. For each incident, verify evidence artifact exists and is complete: incident record, timeline (auto-captured), severity and class declaration, roles assigned and logged, communications log (Slack + status page), postmortem (if mandatory), action items in tracker with due dates, action completion evidence (code, alert test, training record).\n- **Control walkthrough**: walk through each control statement from S3 with a checklist. Is the evidence artifact present? Is it immutable? Is it searchable? Is retention adequate? Is access logged?\n- **Gap remediation**: for every gap found, estimate time to fix and prioritize by audit risk. Anything risking a qualified opinion (e.g., missing postmortem, no timeline evidence) must be fixed before the audit. Test the remediation against a new incident or a resample.\n- **Interview readiness**: brief 10–15 engineers who may be interviewed by auditors (ICs, Comms Leads, Process Owner, team managers). Ask them to describe the process as they actually practice it, not as written. Listen for confusion or gaps in understanding. Correct them.\n- **Auditor package preparation**: assemble process documentation, sample incident records (5–10 complete golden files), training records, on-call schedules, alert quality metrics, action tracker register, and status-page archive. Organize by control. Create a table of contents and index.\n- **Single audit liaison**: designate Process Owner or a small dedicated compliance person as sole point of contact for audit requests. Prevents requests scattering across 28 teams.\n- **Rehearsal**: conduct mock interview with an IC and a Comms Lead. Auditors ask tough questions under pressure (\"How do you know the timeline is accurate?\", \"What happens when both ICs are unavailable?\", \"Show me how you proved the alert was actionable.\"). Practice answering.", "dependencies": ["S3", "S20"]}, {"step_id": "S22", "title": "Standing governance and continuous improvement", "description": "Lock in durable improvement. The classic post-audit failure is the process freezing and then decaying. This step prevents that.\n\n- **Standing Incident Management Council**: chaired by Process Owner, monthly meetings, attendees: CTO, VP Engineering, Head of Support, Head of Compliance, one engineering manager per region, one IC, one team manager from recent wave. Agenda: metrics review, policy changes, gaps from recent incidents, escalation for contentious issues.\n- **Change control mandate**: give Process Owner documented authority to change severity taxonomy, response classes, roles, communications timings, and compensation policy. Any change requires: written justification, steering group approval (monthly cadence), and documented effective date before implementation. This prevents silent drift and ensures changes are deliberate.\n- **Quarterly validation**: re-validate severity and response class taxonomy against all incidents from the prior quarter. Ask: did our taxonomy correctly predict response posture? Did we misclassify? Update taxonomy if patterns emerge.\n- **Annual metric re-baselining**: every 12 months, re-run measurements from S2 (alert census, incident register) to reset targets. System should improve; targets should tighten.\n- **Resilience roadmap separation**: fund a distinct architectural or platform roadmap for incident prevention (reduce shared-database blast radius, multi-region failover, deploy safety, observability investments). Better incident response does not protect a single-ledger corruption or unplanned failover. These are separate problems.\n- **Quarterly executive report**: CTO and VP Eng report to CEO/CFO on metric set (detection time, mitigation time, SLA credits, customer-detected %), top five systemic causes of incidents, program cost vs. credit avoidance, and strategic architecture changes in flight.\n- **Public backlog of improvement ideas**: teams and engineers propose process improvements via Slack or wiki. Process Owner reviews quarterly and implements accepted ideas (e.g., \"add a dashboard for detection gaps\", \"update postmortem template\"). Publish what changed and why.\n- **Celebration and learning**: share wins publicly each quarter (\"We reduced MTTD from 22 min to 5 min\", \"Customer-detected incidents down 80%\", \"$800K SLA credits avoided\"). Refresh training and tabletop program annually and immediately after any SEV-1 to keep the system sharp and responsive to new scenarios.", "dependencies": ["S20", "S21"]}], "estimated_complexity": "high", "success_metrics": "- Median Time to Detect from 22 minutes to <5 minutes for internal detection and <10 minutes for customer journeys by month 6; <2% of customer-impacting incidents first detected by customers (vs. 40% currently) by month 9.\n- Median Time to Mitigate for SEV-1 from 3h 10min to <45 minutes by month 9; SEV-2 to <90 minutes by month 9.\n- Monthly alert volume from 3,400 to <400 with >95% signal-to-noise ratio by month 5; all alerts satisfy paging contract (symptom-based, owned, runbook-linked, severity-mapped).\n- Zero incidents with command ambiguity lasting >15 minutes from month 2 onwards; 100% of SEV-1 and SEV-2 incidents have named IC logged in timeline within 5 minutes; 100% of Triage Owners assigned within acknowledgement.\n- 100% of the 28 teams have a documented Service On-Call rotation or a signed service-ownership transfer by month 4; Platform Duty rotation and IC roster both staffed and live by month 3.\n- On-call compensation policy published and effective by month 1; on-call satisfaction score ≥7/10 by month 6; zero on-call-attributed voluntary attrition by month 6.\n- SLA credits paid from $1.3M annually to <$100K by month 12; credit avoidance (prevented credits) tracked and reported monthly.\n- 100% of mandatory postmortems (SEV-0, SEV-1, SEV-2, and repeat incidents) published internally within 15 business days by month 4.\n- Postmortem action item completion rate from 17% (11 of 64) to >90% within 60 days by month 6; median action age <30 days; zero repeat incidents caused by the same contributing factor without a design review.\n- 100% of the 180 services have a named owner, a detection contract, and at least one symptom-based alert by month 6.\n- Status-page first update published within 30 minutes for ≥95% of SEV-1 incidents by month 4.\n- 100% of new engineers complete incident-response onboarding within 30 days of joining; IC certification includes written exam and live simulation; ≥2 certified ICs per team group; zero uncovered hours in 24x7 IC roster.\n- Weekly operational review held in ≥90% of weeks; 12 of 12 monthly reliability reviews; 4 of 4 quarterly executive reviews; all reviews end with documented decisions and owners.\n- SOC 2 Type II audit passes all incident-response controls (CC7.1–7.5, CC2.2–2.3, CC4.1, CC3.x) with zero findings by month 8.\n- All pilot and rollout incidents captured with complete golden incident files (timeline, roles, communications, postmortem, actions, closure evidence) by month 3 onwards; audit dry run identifies zero critical gaps by month 7.\n- All 28 teams transitioned to new process by month 5; all legacy alert tools decommissioned; single source of truth for incidents established and sustained."}P2 kept its round-1 architecture intact and added the three things it was missing: an immediately usable interim process, a funded prevention track, and an explicit culture step. Compliance sequencing, metrics and communications are essentially unchanged and were already strong.
- New S5 day-one operating rules: ten rules needing no procurement, a manual duty-IC rotation drawn from the 12 teams that already have on-call, a register entry within 24 hours, and a daily 15-minute stand-up for month one — the only plan that produces value before tooling lands.
- New S23 turns "resilience roadmap" from a bullet into a funded track with testable items: ledger read-only tripwire, connection-pool isolation per domain, measured per-region RTO, scheduled PITR restores with published timings, rollback on SLO burn.
- S2 adds a cost-of-downtime model (dollars per minute per journey) that later drives wave sequencing and the funding case.
- S9 adds a New York labour, overtime and tax review with Legal and Finance before announcing paid on-call — the only plan that notices the legal exposure.
- S14 broadens the credit ledger into a customer-trust workstream: CTO-signed note to the top 50 accounts, quarterly public reliability report, account-manager script.
- New S24 makes incentives explicit: restate amnesty at every wave, correct blame language in the executive review itself, pull engineers off rotation when the intrusion cap breaks.
- 24 steps with S5, S14, S23 and S24 overlapping existing steps (S14 duplicates parts of S13; S24 duplicates parts of S17 and S22) — the plan is getting long at the edges.
- S23 depends only on S1 and S17, so a major engineering track sits outside the rollout gates with no stated budget split against the $500–700K envelope.
- Success metrics are almost unchanged from round 1 despite the new steps; nothing measures the day-one process (S5) or the resilience bets (S23).
- Proposal 3 : A dedicated culture and change-management step covering pager fatigue, recognition and blame correction.
- Proposal 1 : Weekly synthetic pages to test every escalation path.
- Proposal 1 : Status-page components mapped to customer journeys rather than internal services, with a named owner.
- Proposal 1 : Fund a resilience roadmap separate from incident response to reduce shared-database blast radius.
- Proposal 1 : Status-page update within 3 minutes of SEV-1 declaration and internal updates every 5 minutes.
- Proposal 1:
- Proposal 1: A roster of 40+ certified incident commanders as a success metric.
- Proposal 3 : Blocking incident closure on completion of high-priority action items.
+ Day-one operating rules and the minimum viable process+ Paging contract and the noise-reduction programme+ Customer trust workstream and the SLA credit ledger+ Postmortems: mandatory set, three levels, blameless by design+ Standing governance and process ownership+ Resilience track for the shared ledger and multi-region blast radiusPaging contract and alert qualityCustomer-impact ledger and SLA credit automationPostmortem policy with three artifact levels
The plan produced
1. Mandate, one owner, and the evidence clock
This step turns the CEO's email into a funded programme with a single accountable owner, and it starts the SOC 2 clock on day one.
- Appoint a Director of Incident Management reporting to the CTO, with a dotted line to the COO for customer and SLA matters and to the Head of Compliance for audit readiness.
- Publish a one-page charter: scope covers every customer-impacting, money-moving, security and data-integrity event across all 28 teams and both regions.
- Grant explicit authority to declare an incident, set severity, freeze deploys, page any engineer in the company, and approve customer messaging.
- Fund the envelope up front: tooling, training and drill time, on-call compensation, and a small programme team, roughly $500–700K a year against $1.3M in credits paid.
- State the return plainly to the steering group: credits avoided, churn avoided, and audit findings avoided.
Start the evidence clock now. A SOC 2 Type II report covers a period of operating effectiveness, so evidence cannot be backfilled; an audit in eight months means the operating period begins in the first weeks of this programme.
- Stand up a fortnightly steering group of the CTO, VP Engineering, Head of Support, Head of Compliance and two engineering managers.
- Open a programme risk register with the top risks, owners and review dates, and revisit it at every steering group.
- Make participation in the incident process a documented performance expectation for every engineering manager, not an optional extra.
2. Baseline evidence pack and cost-of-downtime model (after 1)
You cannot demonstrate improvement without a defensible starting point, and you cannot win the argument about noise without numbers.
- Build a 12-month incident register and re-classify all 31 customer-impacting incidents by detection source, duration, customers affected and credits paid.
- Build the silent-failure register: incidents where no internal alert fired at all, which is the number that explains the 40% customer-detected rate.
- Run an alert census per tool, per team and per service: volume, page-to-action ratio, off-hours interruptions, and the 50 noisiest rules with a named owner each.
- Survey every on-call engineer and manager on burden, fairness, escalation clarity and pay expectations, targeting above 70% response.
- Interview Support, Account Management and Sales about how customers learn of incidents and what they are promised.
- Reconstruct the two command-ambiguity incidents minute by minute to find exactly where ownership lapsed.
- Build a cost-of-downtime model: dollars per minute of impact per customer journey, used later to sequence teams and justify funding.
- Publish the pack internally as the problem statement, and retain every artifact as management-review evidence for the audit.
3. SOC 2 control mapping and evidence architecture (after 1)
Most programmes leave compliance to the end; this one maps controls in month one, because the mapping decides what the process must capture from day one.
- Map the process to the Trust Services Criteria for incident identification, evaluation, response, containment, recovery and communication (CC7.1–CC7.5, CC2.2, CC2.3, CC4.1, CC3.x).
- Write each control in plain language, with one named owner and its evidence artifact.
- Define the golden incident file: one auto-assembled, single-click export per incident containing timeline, roles, severity, communications, postmortem, actions and closure proof.
- Set retention, storage location and immutability so no control depends on a laptop, a private channel, or chat history that expires.
- Keep a gap register with owners and dates, reviewed fortnightly by the steering group.
- Meet the auditor's readiness team inside the first 90 days to test the control design before anything is built on top of it.
4. Severity and class taxonomy with the trigger matrix (after 2) from P3 step 3
Severity decides how loud the response is; class decides which playbook runs and what must never be traded away.
- Define severity in customer and money terms: customers affected, dollars at risk, regions lost, ledger integrity, settlement delay, data loss and SLA exposure.
- Use SEV0 for security, privacy or regulatory events; SEV1 for total or material loss of a payment path; SEV2 for degradation or single-region loss; SEV3 for limited impact with a workaround; SEV4 for internal-only issues and near-misses.
- Define response classes: Availability, Performance, Data integrity and Ledger, Security and Privacy, Third-party dependency, Settlement and Financial, Process failure.
Class can raise the response but never lower it. A SEV2 data-integrity incident gets SEV1 posture, because integrity failures are not recoverable by moving faster.
- Specify automatic triggers: loss of one AWS region, ledger write failures, replication lag above threshold, a missed settlement cut-off, payment success rate below threshold for five minutes.
- State who may declare — any engineer, Support agent or account manager — and who may downgrade: the Incident Commander alone.
- Map every level to its SLA credit exposure and to the customer-visible status page state.
- Include worked examples from the last 12 months so each of the 28 teams recognises its own incidents, and revalidate the taxonomy quarterly against real declarations.
5. Day-one operating rules and the minimum viable process (after 1, 4) new
The full process will take months; the first useful version must be live in two weeks using the tools that already exist.
- Publish ten day-one rules that need no procurement: a named owner within five minutes, one channel per incident, one register entry per incident, one person speaking to customers.
- Apply the ambiguity rule: if two responders disagree about whether this is an incident, it is an incident.
- Make declaring free: a false alarm is closed as a false declaration, tracked as a metric, and never criticised.
- Require a register entry within 24 hours for every customer-impacting incident, even a minimal one.
- Ban silent incidents: if we know, the customer hears it from us rather than from their own reconciliation.
- Run the first 30 days on manual command, with a rotating duty Incident Commander drawn from the 12 teams that already have on-call.
- Hold a 15-minute daily incident stand-up during month one to catch friction while it is still fresh.
6. Roles, command structure and the no-unowned-minute rule (after 4)
The two hour-long command failures did not happen at declaration; they happened in the gap before it, when an alert had fired and nobody owned it.
- Publish one-page role cards: Triage Owner, Incident Commander, Deputy IC, Operations Lead, Internal Comms Lead, Customer Comms Lead, Scribe, Subject-Matter Responders, and Executive Sponsor for SEV1 only.
- Introduce the Triage Owner rule: whoever acknowledges the page owns the incident until an IC takes over or the incident is stood down.
The IC owns the incident, not the fix, and does not debug. An IC who starts troubleshooting has abandoned command.
- Give the IC explicit decision rights: declare and escalate severity, freeze changes, halt deploys, approve customer messaging, and pull in any engineer.
- Define minimum viable staffing per severity: SEV1 fills every role; SEV2 staffs IC, scribe, comms and responders; SEV3 staffs an IC and a scribe.
- Set handover discipline: four-hour maximum IC shifts on SEV1 with a written handover, and a deputy named within 15 minutes of declaration.
- Set responder behaviour: one channel, one bridge, no side channels, and every request phrased with a named owner and a time.
- Link the role cards from every paging notification so they are one tap away at 3 AM.
7. Lifecycle, declaration and escalation policy (after 4, 6)
This step defines the mechanical path from an alert to a declared incident and back to normal service, removing judgment calls from the worst moments.
- Define states with entry and exit criteria: Detected, Triaged, Declared, Mitigated, Resolved, Postmortem, Closed, plus a Watch state with a hard 30-minute timer.
- Set targets: page acknowledged within five minutes, triage decision within fifteen, declaration within thirty for anything customer-facing.
- Build automatic escalation ladders with timers at every layer: responder, service owner, team manager, IC on-call, VP Engineering.
- Define the unresponsive-team path: fifteen minutes escalates to the team manager, thirty minutes to the IC, who may then direct any available engineer.
- Define change freeze and rollback authority during SEV1 and SEV2, with the single named condition that lifts the freeze.
- Enforce one incident, one record, with the timeline captured automatically from the channel and bridge rather than written from memory afterwards.
- Add the two-simultaneous-incidents rule: a second IC is designated, and no IC ever runs two incidents at once.
- Test every escalation path weekly with synthetic pages, and adjust the timings after the first month of real operation.
8. On-call architecture across 28 teams (after 4, 6)
The objection is that engineers will not carry a pager for another team's code; the answer is to build rotations that make the objection structurally impossible.
- Run a Service On-Call rotation per team, covering only that team's own services.
- Run a Platform Duty rotation for genuinely shared infrastructure: the PostgreSQL ledger cluster, Kubernetes, networking, CI/CD and observability.
- Run a central Incident Commander roster of 12–16 certified senior engineers on one-week shifts with a primary and a secondary.
State the consequence honestly. Sixteen of 28 teams have no rotation today; each must build one or formally transfer service ownership to a team that has one, with the transfer dated and recorded.
- Merge small or low-traffic teams into shared rotations so no rotation has fewer than six engineers.
- Cap load in the scheduling tool: no engineer on call more than two weeks per quarter, enforced by configuration rather than negotiation.
- Publish a coverage matrix of all 28 teams showing services, owners, rotation size and gaps, reviewed monthly.
- Write on-call participation into engineering job levels and hiring criteria, so the commitment is set before someone joins.
9. Compensation, rest and the economics of opting out (after 8)
Unpaid on-call is the most cited reason for resistance, so settle compensation before rollout, not during it.
- Move to paid on-call: a per-shift stipend benchmarked to the New York market, agreed with HR and Finance, published with an effective date before any team is asked to join a rotation.
- Pay for callouts at 1.5× the hourly rate for time actually spent mitigating, with a minimum block per interruption.
- Provide documented compensatory rest: no normal working day after a night incident, and the rest day is policy rather than a favour granted by a manager.
- Cap intrusion with a maximum number of off-hours pages per shift, and trigger a mandatory review whenever the cap is exceeded.
Allow opt-out, but put a price on it. An engineer may step out of a rotation, and their team buys coverage from the paid pool at a published internal rate, which turns a cultural argument into a visible budget decision.
- Publish amnesty: incident records, near-misses and false declarations are never used in performance reviews; only failure to report is a performance issue.
- Protect against gaming: pay attaches to shifts and callouts, never to incident counts.
- Check the New York labour, overtime and tax treatment with Legal and Finance before announcing, and review the policy every six months against real page volumes, attrition and survey results.
10. Detection strategy: journeys, synthetic signals and customer-report intake (after 4)
Customers detected 40% of incidents first, which makes detection the highest-leverage business problem in this programme.
- Define SLIs and SLOs for the top 20 customer journeys, measured per region: payment initiation, authorisation, settlement, ledger read and write, API availability, webhook delivery and payout.
- Alert on symptoms against those SLOs, not on cause-based infrastructure metrics.
- Add synthetic transaction monitoring from outside AWS in both regions plus a third location, on a one-minute cadence for money-moving paths.
- Add ledger and shared-PostgreSQL signals: replication lag, connection saturation, write latency, lock waits, transaction ID exhaustion, checkpoint pressure and bloat.
- Open a customer-report intake so Support and account managers can raise an incident directly, and count that path as a detection source in all reporting.
- Apply the detection-gap rule: whenever a customer reports first, a detection-gap ticket is opened automatically and owned by the responsible team.
- Publish a detection contract per service — owner, at least one symptom alert, documented expected detect time — for all 180 services.
- Run a detection drill per team: break something in staging and see whether it pages before a human notices.
11. Paging contract and the noise-reduction programme (after 2, 10) from P3 step 7
3,400 alerts a month at 85% noise is the reason engineers resent the pager, and fixing it is the price of admission for everything else in this plan.
- Publish the paging contract: every page must be symptom-based, actionable, owned, mapped to a severity and class, and linked to a runbook.
No runbook, no page, enforced by a CI check on the alert definition itself.
- Separate paging from ticketing: only customer-impacting or imminent-impact signals may wake a human.
- Set a page budget per team and per service, with a remediation ticket opened automatically, owned by the engineering manager, when the budget is breached.
- Put new alerts on two-week probation as ticket-only until they have proved actionable.
- Define the default action for a noisy alert: fix, tune or delete within ten working days, with deletion treated as a legitimate and celebrated outcome and the deleted count published.
- Deduplicate and correlate at the ingest pipeline so one root cause produces one page instead of forty.
- Require an expiry date on every silencing rule and temporary threshold, so suppression cannot quietly become permanent blindness.
- Run a 90-day noise sprint with a public burn-down of the top 100 noisiest rules, each owned by a named manager.
12. Incident tooling consolidation and the golden incident file (after 3, 6, 10, 11)
Six alerting tools and no single incident record are structural causes of the 22-minute detection and the three-hour mitigation.
- Choose one incident platform for paging, schedules, escalation policies, incident records, postmortem workflow and status-page integration, and choose it within four weeks.
- Consolidate the six alert sources into a single event pipeline feeding that platform, with deduplication, correlation and severity and class mapping applied at ingest.
- Integrate observability, runbooks, chat and voice bridges so responders see dashboards and runbooks inside the incident record, and the timeline is captured automatically.
- Implement the golden incident file export defined in S3, so audit evidence is one click rather than a reconstruction.
- Run a dual-run period alongside the legacy tools with a defined rollback, then switch the legacy tools off on a published cutover date.
- Host the status page outside the production failure domain so it survives a total platform outage, and prove that in a game day.
- Make the platform usable from a phone: an Incident Commander must be able to run a SEV1 from a mobile device at 3 AM.
13. Internal, customer and regulator communications (after 6, 7, 12)
Today the status page is written by whoever is around; this step replaces improvisation with a clock, a named owner and pre-cleared templates.
- Set the internal cadence: first update within 15 minutes of declaration, then every 30 minutes for SEV1 and hourly for SEV2, whether or not there is progress.
- Never let an employee learn of an incident from the status page: internal communication leads, external follows.
- Set the customer cadence: status page updated within 30 minutes of a SEV1 and 60 minutes of a SEV2, with a no-new-information update still mandatory.
- Define the channel hierarchy: status page for all 2,100 customers, direct email to affected customers on SEV1, and named account-manager calls for the top 50 accounts.
- Pre-approve templates per severity and class with Legal and Compliance, each carrying its own next-update time.
- Forbid speculation: customer messages never guess at cause, never assign blame, and never commit to a root cause before the postmortem.
- Define the closing communication: resolution notice, how SLA credits will be handled, and the committed date for the written report.
- Build a regulator clock matrix covering event type, regulator, notification window, signer and the shortest applicable clock, including NYDFS Part 500, money-transmitter and banking notifications, breach notification, card-network rules and public-company disclosure.
The regulatory clock starts at awareness, not at root cause. Route every notification through Compliance, never Engineering, and pre-clear the templates.
14. Customer trust workstream and the SLA credit ledger (after 13) new
The $1.3M in credits is a symptom of having no single record of customer impact, and the CEO's inbox is a symptom of customers learning things late.
- Maintain one durable customer-impact record per incident: which customers, which journeys, from when to when, and the estimated credit.
- Use that one record for communications, credit computation, regulatory reporting and the annual review, so nobody reconciles two versions of the same outage.
- Automate the credit path end to end — computation, approval, customer notification and finance treatment — so credits stop being a manual scramble three weeks later.
- Reconcile accrued against paid credits monthly and report the result in the executive review.
- Give the status page a named product owner and map its components to customer journeys, not to internal services.
- Send a CTO-signed reliability note to the top 50 accounts and publish a quarterly reliability report to all customers.
- Give account managers a script of the facts they may state, the speculation they may not, and a path for customer escalations.
- Track credit avoidance against programme cost, so the funding case stays a number rather than an argument.
15. Postmortems: mandatory set, three levels, blameless by design (after 6) new
Postmortems currently happen for some incidents, in various formats; this step makes them mandatory where they matter and light where they do not.
- Make postmortems mandatory for every SEV0, SEV1 and SEV2, every SEV3 with customer impact or a repeat pattern, every near-miss the IC flags, and every incident where the process itself failed.
- Use three levels so the ritual matches the weight: a lightweight async review for SEV3 and SEV4, a facilitated postmortem for SEV2, and a full review with an executive sponsor for SEV1.
- Set deadlines: draft within five business days, blameless review within ten, internal publication within fifteen.
- Adopt one template: impact, timeline, detection, response, contributing factors across tooling, process, organisation and human factors, what went well, what went badly, and actions.
- Train a pool of blameless facilitators and require a trained one for every SEV1 review, never the IC.
Ban blame language in the template and ban "human error" as a root cause. The question is always what system condition made the error possible.
- Publish all postmortems internally by default, with security review only for genuinely sensitive material.
- Produce a customer-facing root-cause report for SEV1 incidents, especially for regulated and top-tier accounts.
16. Action items: capped, verifiable, with the repeat-incident rule (after 15)
Eleven of 64 action items closed is not a tracking problem; it is a generation problem, because the process produces more actions than the organisation can absorb.
- Cap each postmortem at three action items, with anything beyond that going into a ranked reliability backlog.
- Require every action to have a named human owner, a due date, and a definition of done that is an artifact: a merged change, an alert that fired in a drill, or a test that fails without the fix.
- Prohibit self-reported closure; closure requires the artifact and sign-off by the process owner or the Incident Commander.
- Keep one reliability backlog in the engineering tracker with a mandatory label and a link back to the originating incident.
- Reserve a fixed share of each team's sprint capacity for reliability work, with unspent capacity visible to the vice presidents.
- Apply the repeat-incident rule: a second incident with the same contributing factor triggers a design review owned by an architect, not another action item.
- Run a weekly ageing review of open actions and escalate anything more than 30 days overdue to the VP of Engineering.
- Report action completion rate and median action age monthly, by team.
17. Metrics, dashboards and the review cadence (after 2, 4, 16) from P1 step 17
Define what good looks like, then measure it in a way that rewards reporting incidents rather than hiding them.
- Define outcome metrics: time to detect by source, time to mitigate by severity and class, share of incidents detected by customers, incidents by severity and class, SLA credits paid, and customer-impact minutes.
- Define process metrics, because the process itself has SLOs: IC assigned within five minutes, page acknowledgement rate, first-update timeliness, update-cadence adherence, postmortem timeliness, action closure rate, action median age, and IC roster coverage.
- Define health metrics: alert volume and noise ratio per team, off-hours pages per engineer, on-call satisfaction, training completion, and share of services with a detection contract.
Never publish incident count as a team metric. It rewards hiding incidents; publish reporting metrics instead — near-misses filed, detection gaps closed, false declarations made.
- Publish live dashboards visible to every engineer, refreshed daily, with every metric baselined against the S2 evidence pack and given a 90-day and a 12-month target.
- Run a 30-minute weekly operational review incident by incident, a monthly reliability review of trends and actions, and a quarterly executive review with the CEO's office.
- Hold a quarterly review of the process itself: what wasted time, what confused responders, and what should be deleted.
- End every review with decisions and named owners, never with numbers alone.
18. Training, certification and the drill programme (after 6, 7, 13, 15)
A process that lives only on a wiki page fails on the first real page, so skills are built and tested before they are needed.
- Build a practical curriculum: how to be on call, how to triage and declare, how to run an incident as IC, how to communicate, and how to write a postmortem.
- Require certification before joining the IC roster: a written assessment plus a live simulated incident.
- Certify at least two ICs per team group so the central roster has depth across all 28 teams and no holiday week is left uncovered.
- Train communications leads separately on templates, cadences, customer language and the regulatory rules.
- Run monthly tabletops drawn from the last 12 months of real incidents, including region loss, ledger write failure and a missed settlement cut-off.
- Run quarterly game days with deliberately injected failure, including PostgreSQL failover, status-page outage and alerting-pipeline outage.
- Drill the process's own failure modes, not just technical ones: IC unreachable, comms lead on leave, two simultaneous SEV1s, a paging storm, and a false alarm that burns an hour.
- Audit the process for single points of failure: who alone can perform each critical task, and what happens in their holiday week.
- Keep a mandatory onboarding module for every engineer joining or transferring in, with audit-ready completion records.
19. Pilot with three to four teams, using real incidents (after 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 18)
Do not roll out to 28 teams untested, and do not rely on drills alone when real incidents are better training material.
- Recruit three to four volunteer teams spanning criticality: a payment-path team, a ledger-adjacent team, the platform team that owns the shared PostgreSQL cluster, and a low-traffic team.
- Run the entire process end to end in the pilot: severity and class, roles, escalation, communications, postmortems, action tracking, and paid on-call.
- Treat real incidents during the pilot as the primary training material, and hold a retro within 48 hours of each one, run by the process owner while the friction is fresh.
- Instrument the pilot against the baseline and publish before-and-after numbers.
- Hold weekly retrospectives with the pilot teams, and change the written policies, the tooling and the training in response, documenting what changed and why.
- Set explicit exit criteria: rotation coverage achieved, zero unacknowledged pages over two weeks, postmortems delivered on time, and actions tracked to closure.
- Produce a pilot report that every later rollout conversation starts from.
20. Phased rollout to 28 teams, sequenced by cost of failure (after 16, 18, 19)
Rollout is a staged migration with readiness gates, not an email announcement, and the sequencing matters more than the schedule.
- Split the 28 teams into four waves of roughly seven, ordered by incident density and customer-journey ownership: the teams whose incidents cost the most SLA credits go first, because their improvement is the visible proof.
- Leave three weeks between waves and define a per-team readiness checklist: services mapped and owned, alerts cleaned to the paging contract, runbooks written, rotation staffed to at least six engineers, training complete, manager briefed, compensation in effect.
- Hold a gate review with the process owner before each team joins, and move an unready team to the next wave with a dated remediation plan rather than granting an exception.
- Give each wave a named champion and explain the why using the pilot's numbers.
- Handle resistance directly and early: publish the own-your-code-own-your-pager rule and the paid on-call mechanics before each wave, never after.
- Hold a retro within 48 hours of each wave's first incident under the new process, and push accepted changes through change control.
- Retire legacy tools, informal escalation lists and the ad-hoc status page process at the end of each wave on a published cutover date, so two processes never run in parallel for long.
- Sequence the waves so no team faces both a rollout and the audit dry run in the same week.
21. Audit dry run and evidence review (after 3, 20)
This step converts a good process into a provable one, about six weeks before the auditors arrive.
- Sample real incidents from the pilot and the early waves against each control's evidence requirements.
- Remediate every gap found and re-test the remediated control against the same sample, prioritising anything that risks a qualified opinion.
- Brief every engineer who may be interviewed so they can describe the process as they actually practise it, not as it is written.
- Prepare the auditor package: process documentation, sample incident records, paging logs, communication logs, training records, on-call schedules, the action register and alert quality metrics.
- Designate one audit liaison and a small evidence-request team, so requests do not land on all 28 teams at once.
- Rehearse the walkthrough with an Incident Commander and a communications lead, because auditors probe realism under pressure.
- Keep the audit liaison and the process owner as close to the same person as possible, so accountability for the control is also accountability for the evidence.
22. Standing governance and process ownership (after 20, 21) from P1 step 23
The classic post-audit failure is that the process freezes and then decays, so ownership has to outlive the programme.
- Establish a standing Incident Management Council chaired by the process owner, meeting monthly with engineering, support, compliance and product representation.
- Give the process owner a documented mandate to change standards, with formal change control for any change to severity, classes, roles, communication timings or compensation.
- Revalidate the severity and class taxonomy quarterly against real declarations, and re-baseline the entire metric set annually.
- Report quarterly to the executive team on the metric set and the top five systemic causes of incidents.
- Feed incident themes into architecture review, release engineering and hiring, so the programme improves the system and not only the response.
- Keep a public backlog of process improvement ideas so any engineer can propose a change, and act on it every quarter.
23. Resilience track for the shared ledger and multi-region blast radius (after 1, 17) new
Better incident management cannot save a single shared ledger during a corruption event, so this runs as a funded engineering track beside the process, not inside it.
- Reduce ledger blast radius: read replicas for read paths, connection-pool isolation per service domain, statement timeouts, write throttles and per-tenant limits.
- Build a ledger protection mode: an automatic tripwire that degrades to read-only or queues writes before corruption can spread.
- Rehearse multi-region failover with the ledger as the binding constraint, and set a measured recovery time objective per region.
- Test restore and point-in-time recovery on a schedule, and publish the timings.
- Tighten deploy safety: progressive rollout, automatic rollback on SLO burn, and change-freeze windows around settlement.
- Track a small number of resilience bets with named owners and quarterly milestones, reviewed in the same executive review as the incident metrics.
24. Culture, incentives and continuous improvement (after 20, 22)
Incident management is a cultural change before it is a process change, and the incentives decide whether reporting goes up or down.
- Restate the amnesty at every rollout wave: reporting, near-misses and false declarations are safe; hiding is not.
- Recognise detection improvements, deleted alerts, closed detection gaps and well-written postmortems publicly, and reward the engineers who do them.
- Correct blame language in leadership communication on the spot, including inside the executive review.
- Watch for pager fatigue with real data, and take an engineer off rotation as soon as the intrusion cap breaks.
- Run a quarterly process retro with the whole engineering organisation and publish what changed as a result.
- Refresh the training and tabletop programme annually and immediately after any SEV1.
- Share wins publicly each quarter — faster detection, fewer customer-detected incidents, credits avoided — to hold momentum after the audit.
- Median time to detect falls from 22 minutes to under 8 minutes by month 6, and under 5 minutes for money-moving journeys by month 9.
- Share of customer-impacting incidents first detected by customers falls from 40% to below 10% by month 9.
- Median time to mitigate for SEV1 falls from 3 h 10 min to under 60 minutes by month 9.
- 100% of SEV1 and SEV2 incidents have a named Triage Owner within 5 minutes and a named Incident Commander within 15 minutes, from month 4.
- Zero incidents with unclear command authority lasting more than 15 minutes, measured monthly from month 4.
- Monthly page volume falls from 3,400 to under 600, with a false-positive rate below 15%, by month 5.
- No service exceeds two pages per on-call shift for three consecutive months by month 6.
- 100% of the 180 services have a named owner, a detection contract and at least one symptom-based alert by month 7.
- All 28 teams have a documented rotation, a service ownership map and at least one trained on-call engineer by month 6.
- 100% of on-call shifts are paid under a published policy from month 2, with zero on-call-attributed voluntary attrition by month 6.
- On-call satisfaction scores 7 out of 10 or better in quarterly surveys from month 6.
- The IC roster holds at least 12 certified ICs covering 24x7 with no uncovered week from month 4.
- 100% of SEV0, SEV1 and SEV2 postmortems are published internally within 15 business days from month 5.
- Action items closed within 60 days rise from 17% to above 90%, with a median action age under 30 days, by month 6.
- Status page first update is posted within 30 minutes on at least 95% of SEV1 incidents from month 4.
- Zero missed regulatory notification windows on any incident requiring notification.
- The customer-impact record is complete for 100% of customer-impacting incidents from month 5, with SLA credits reconciled monthly.
- Annual SLA credits paid fall at least 50% from the $1.3M baseline within 12 months.
- Customer-impacting incidents fall at least 40% year over year from the 31-incident baseline.
- Review cadence is sustained: weekly operational review in at least 90% of weeks, 12 of 12 monthly reliability reviews, 4 of 4 quarterly executive reviews.
- Every page has a recorded disposition — fixed, tuned or deleted — within 10 working days, from month 4.
- SOC 2 Type II is passed with zero findings related to incident response.
- 100% of new engineers complete the incident-response onboarding module within 30 days of joining.
[SYSTEM]
You are an expert assistant in complex project planning.
Your task is to generate a detailed and structured action plan to reach the main objective in the most professional and most detailed way, no matter how much work or steps will be needed to perform.
Use your internal reasoning processes to think deeply about the problem and create the most comprehensive plan possible.
Take as much time and space as you need to think through all aspects of the problem.
After your thorough analysis, answer with the plan in the requested structure.
Every text field you write will be read by a busy person, so write for them: short sentences, one idea each, in short paragraphs. Text fields accept Markdown: separate paragraphs with a blank line, use a bullet list when you enumerate things and plain paragraphs when you explain or argue, and bold at most one key phrase per paragraph. A text of more than three sentences must be split into paragraphs; never deliver a single unbroken block of text.
[HUMAN]
Main Objective: "A B2B payments platform in New York (260 engineers in 28 teams, 2,100 customers, $4B processed a month, a 99.95% SLA with service credits) runs about 180 services on Kubernetes across two AWS regions, with a shared PostgreSQL cluster for the ledger. In the last 12 months: 31 customer-impacting incidents, median time to detect 22 minutes (customers detected 40% of them first), median time to mitigate 3 h 10 min, $1.3M paid in SLA credits, two incidents in which nobody was sure who was in charge for over an hour, and a CEO email about "outages we hear about from clients". On-call exists in 12 of the 28 teams, unpaid, fed by alerts from six different tools — 3,400 alerts a month, 85% of them noise; there is no severity scale, status-page updates are written by whoever is around, and postmortems happen for some incidents, in various formats, with action items that are rarely tracked (11 of 64 closed). A SOC 2 Type II audit in eight months will test incident response. Engineers push back against "carrying a pager for other teams' code".
Define the incident management process: severity levels and what each one triggers; roles (incident commander, communications, scribe, subject-matter responders) and how they are staffed 24x7 across 28 teams; the on-call structure, rotations, compensation and the rules for alert quality; detection and escalation paths; internal and customer communications (status page, account managers, regulators when required) with their timings; postmortems (when mandatory, format, blameless review, ownership and tracking of actions); the metrics and reviews that show whether it works; and how it is introduced across the teams without waiting for the audit."
For your consideration and refinement, here are proposals from the previous round:
Previous Proposal 1 (ID: 2c50b755-66c6-49f5-aea8-767330e89bc3, Agent: claudeHaiku4.5_refine_1, LLM: anthropic/claude-haiku-4-5):
Estimated Complexity: high
Success Metrics: - Median time to detect reduced from 22 minutes to under 5 minutes by month 6, with >90% internal detection (vs. 40% customer-detected now).
- Median time to mitigate reduced from 3h 10min to under 60 minutes for SEV1 and SEV2 by month 9.
- Customer-impacting incidents detected by customers drop from 40% to <5% of all customer-impacting incidents.
- Alert volume reduced from 3,400 per month to <600 per month; signal-to-noise ratio improves from 15:85 to >95:5.
- SLA credits paid reduced from $1.3M to <$100K annually by month 12.
- Zero incidents with command-and-control ambiguity lasting >15 minutes; all SEV1/2 incidents have named IC logged in timeline within 5 minutes.
- Postmortem action item completion rate reaches >80% (from 11 of 64, or 17%) by month 4.
- 100% of SEV1 and SEV2 postmortems published within 15 business days by month 5.
- All 28 teams integrated into incident management system with active on-call rotations by week 20; no team unresponsive to pages for >30 minutes.
- On-call satisfaction score reaches >7/10 on survey; zero on-call-attributed voluntary attrition by month 6.
- Incident commander roster: 40+ certified ICs covering 24x7 with no single point of failure by month 4.
- Status-page first update published within 30 minutes on ≥95% of SEV1 incidents by month 3.
- SOC 2 Type II audit passes incident response controls with zero findings by month 8.
- Weekly incident review cadence sustained in ≥90% of weeks; monthly reliability reviews 12 of 12; quarterly executive reviews 4 of 4.
- 100% of new engineers complete incident-response onboarding within 30 days of joining.
Steps (23):
1. Executive mandate and governance structure
Turn the CEO email into a funded, authorized program with clear ownership and decision rights. Without executive backing, every downstream decision stalls in negotiation.
- Appoint a single program owner (e.g., Director of Incident Management) reporting to the CTO and COO.
- Publish a one-page charter covering scope (all customer-impacting incidents), authority to override team preferences during incidents, and funding for tooling, training, and on-call compensation.
- Establish a standing Incident Management Steering Group with CTO, VP Engineering, VP Support, Head of Compliance, and one engineering manager per region meeting monthly.
- Secure budget envelope: tool licenses, training time, incident-response infrastructure, and on-call compensation (estimated $400–600K annually).
- Make incident-response participation a documented performance expectation for all engineering managers and team leads.
2. Baseline measurement and evidence pack (depends on: 1)
You cannot prove improvement without defensible baseline numbers, and you cannot win arguments about noise or impact without data.
- Build a 12-month incident register: date, detection source, impact scope, time to detect, time to mitigate, SLA credits paid, and services involved.
- Audit the current alert estate: count alerts per tool, per team, per service; compute page-to-action ratio; identify top 50 noisiest rules and off-hours interruptions per engineer.
- Survey on-call engineers and managers on burden, fairness, and escalation clarity (target >70% response rate).
- Interview Support and Account Management: how do customers discover incidents, what do they complain about?
- Document the two command-ambiguity incidents: exactly when unclear who was in charge and why, how long it lasted.
- Publish this pack internally as the problem statement and retain all artifacts for SOC 2 audit evidence.
3. Severity taxonomy and trigger matrix (depends on: 2)
Severity is the keystone. Every other rule—paging, communications, postmortems, compensation—keys off it. Define four levels plus a special SEV0 for security/regulatory events.
- **SEV1 (Critical)**: Complete service outage, data corruption, or >5% payment-path failure rate for >5 min. Every minute costs money. IC required; 99.99% uptime threatened.
- **SEV2 (Major)**: Significant degradation, single region loss, or 1–5% transaction failure. IC typically required; service credit exposure.
- **SEV3 (Minor)**: Limited customer impact with workaround available, or internal issues affecting operations. On-call SME + escalation if SLA at risk.
- **SEV4 (Cosmetic)**: Observations, non-impacting bugs, alerts. Alert-driven, no escalation unless pattern emerges.
- Specify automatic triggers: region loss, ledger write failures, missed settlement window, payment success rate thresholds.
- Define who may declare (any engineer, Support, account manager) and who may downgrade (IC only).
- Include worked examples from the last 12 months so teams recognize their incidents in the definitions.
- Review and re-validate quarterly against actual declarations.
4. Incident roles, command structure, and decision rights (depends on: 3)
The two incidents with >1 hour of command ambiguity prove this step is non-negotiable. Define clear roles with explicit decision authority.
- **Incident Commander**: Owns the incident timeline, not the fix. Declares severity, decides escalation, approves communications, freezes changes, calls responders. Non-technical ICs are acceptable.
- **Deputy IC**: Shadows IC; takes over if IC unavailable. Nominated within 5 minutes of incident declaration.
- **Communications Lead**: Owns internal Slack updates and status-page messaging. Shields IC from interruptions.
- **Scribe**: Records real-time timeline: decisions, who did what, key timestamps. Not responsible for fixing.
- **Subject-Matter Responders**: Engineers with service context. Take IC direction without debate. Report only to IC.
- **Operations Lead** (SEV1 only): Coordinates across multiple responders, manages incident bridge.
- Write one-page role cards with mission, decision authority, and escalation upward. Publish on wiki and link from every paging notification.
- Define minimum viable coverage per severity: SEV1 staffs all roles; SEV2 staffs IC, Comms, Scribe; SEV3 staffs IC + Scribe.
- Establish handover discipline: maximum 4-hour IC shifts on SEV1, written handover template required.
5. Incident tooling consolidation and integration (depends on: 1, 3)
Six alert tools and ad-hoc incident records are structural causes of the 22-minute detection and 3+ hour mitigation. Consolidate to a single incident platform that is the source of truth.
- Select an incident management platform (PagerDuty, Opsgenie, Incident.io, etc.) that supports paging, schedules, escalation, incident records, and postmortem workflow.
- Requirement: the platform must integrate with observability tools, auto-create and pin incident channels in Slack, auto-capture timeline from chat, and support API-driven playbook automation.
- Plan a dual-run period alongside legacy tools with a published cutover date; define rollback criteria.
- Integrate incident record with the 180 services' monitoring and dashboards so responders see everything in one place.
- Define data retention and audit trail to satisfy SOC 2 evidence requirements: who did what, when, under whose authority.
- Budget for licenses, migration effort, and two-week hardening period post-cutover.
6. Alert consolidation and event pipeline (depends on: 5)
Replace six alert sources with a single ingestion point. Deduplicate and route alerts with minimal manual judgment, removing a major source of detection delay.
- Consolidate alert endpoints from six tools into a single event pipeline; this often sits in front of the incident platform (S5).
- Implement deduplication and correlation so a single outage triggering alerts from five monitoring tools produces one page, not five.
- Map every alert to a severity level from S3 (SEV1, SEV2, SEV3, SEV4) at ingestion.
- Log every alert for postmortem analysis and trending.
- Ensure the platform's mobile app works reliably; on-call responders need to engage from any device.
7. Detection strategy: SLOs, signals, and customer-journey monitoring (depends on: 3, 6)
Customers detected 40% of incidents first—a detection gap that must be closed. Build symptom-based alerting that detects outages before customers do.
- Define SLIs and SLOs for the top 20 customer journeys: payment initiation, settlement, ledger read/write, API availability, webhook delivery, measured per region.
- Require symptom-based alerting on SLOs, not cause-based infrastructure metrics (e.g., alert on "payment success rate <99%" not "database CPU >80%").
- Add synthetic transaction monitoring from outside AWS in both regions and a third location, one-minute cadence, for all money-moving paths.
- Add ledger-critical signals: PostgreSQL replication lag, connection saturation, write latency, transaction ID exhaustion, checkpoint pressure.
- Create a detection contract per service: owner identified, at least one symptom alert defined, expected detect time documented.
- Open a customer-reported incident path: Support and account managers can declare an incident directly, counted as a detection source in metrics.
- Fund a separate resilience roadmap to reduce shared-database blast radius, because detection improvements do not protect against ledger corruption.
8. Alert quality standards and noise-reduction program (depends on: 3, 6, 7)
3,400 monthly alerts with 85% noise is the reason engineers resent the pager. Cutting noise is the price of admission for on-call buy-in.
- Publish alert standards: every page must be symptom-based, have an immediate runbook action, be owned by a team, and map to a severity level. No runbook, no page.
- Separate paging from ticketing: only customer-impacting or imminent-impact signals trigger pages. Everything else becomes a dashboard entry or a ticket.
- Set a noise budget per team and service: no service may exceed two pages per on-call shift per month. Breaching triggers a mandatory alert quality review.
- Define the default action for a noisy alert: fix the root cause, tune the threshold, or delete it within ten working days. Deletion is a legitimate successful outcome.
- Implement automatic suppression rules: silence alerts if service auto-recovered within 30 seconds; suppress known maintenance windows; group flapping alerts (>5 in 2 min) into one page; rate-limit noisy services (max 1 alert per 5 min until condition clears).
- Require expiry dates on all silencing rules so suppression cannot become permanent blindness.
- Run a 90-day noise sprint with a visible burn-down of the top 100 noisiest rules, each owned by a named manager.
9. Escalation policies and incident lifecycle (depends on: 3, 4, 5, 6)
Define the mechanical path from alert to incident declaration to resolution. Escalation must be automatic and blameless.
- Define incident lifecycle states with clear entry/exit criteria: Detected → Triaged → Declared → Mitigated → Resolved → Postmortem → Closed.
- Set acknowledgement targets: page acknowledged in 5 minutes; triage decision (is this real?) in 15 minutes; severity declaration (is this customer-impacting?) in 30 minutes.
- Build escalation ladders: if responder does not acknowledge in 5 min, page escalates to service owner, then team manager, then IC on-call. Escalation is automatic, not manual.
- Implement escalation for SEV1: IC paged via phone call + SMS + Slack + mobile; if not acknowledged in 2 min, Deputy IC paged simultaneously; Communications Lead pinged at same time.
- Define change freeze during SEV1 and SEV2: no deployments except to fix the incident. Freeze lifts only when mitigation is confirmed.
- Enforce one incident, one record: the incident record is the sole source of truth. Auto-capture timeline from Slack and bridge; never write timeline from memory later.
- Test all escalation paths weekly via synthetic page to on-call; adjust timings based on first month of operations.
10. On-call architecture and 24x7 coverage model (depends on: 3, 4, 9)
The answer to "carrying a pager for another team's code" is that every team carries its own, and the platform carries shared risk. Design a sustainable model.
- Adopt a federated model: every service has one owning team; that team's on-call carries its service's pager. No team is paged for code it does not own.
- State the consequence clearly: 16 of 28 teams currently have no on-call. They must either build one or formally transfer service ownership to a team that will.
- Create a separate, centrally staffed Incident Commander on-call roster drawn from trained senior engineers across all teams, covering 24x7 with no single point of failure.
- Merge small or low-traffic teams into shared rotations so no rotation has fewer than six engineers (below which coverage is unsustainable).
- Define primary and secondary per slot: secondary engages only on no-acknowledge or explicit request from IC.
- Align coverage across two AWS regions and New York business hours: one global IC rotation; service on-call aligned to service users' time zones.
- Define unresponsive-team escalation: 15 min without acknowledgement escalates to team manager; 30 min escalates to IC, who may direct any available engineer.
- Publish a coverage matrix of all 28 teams showing services, rotation size, and gaps, reviewed monthly.
11. On-call compensation, wellbeing, and sustainability policy (depends on: 10)
Unpaid on-call is the most cited reason for resistance. Settle compensation before rollout, not during it. Make it sustainable.
- Introduce paid on-call: a weekly stipend ($500–1,000) while on-call, regardless of incident volume, benchmarked to New York market rates.
- Pay event-based compensation: 1.5× hourly rate for time spent mitigating incidents during off-hours (minimum one-hour block per callout).
- Provide compensatory rest: no engineer works a normal business day after a night incident requiring >2 hours mitigation. Rest day is documented, not granted as a favor.
- Cap intrusion: define maximum off-hours pages per week (e.g., no more than three per shift). Mandatory review and escalation if exceeded.
- Offer a voluntary opt-out path for engineers with genuine constraints, balanced by explicit obligation that someone else is paid to cover.
- Include on-call expectation and compensation in job descriptions and hiring conversations so commitment is understood before joining.
- Publish the policy with an effective date before any team is asked to join a rotation.
- Review the policy every six months against actual page volumes, attrition rates, and survey feedback.
12. Playbooks and communication templates by severity (depends on: 3, 4)
Playbooks remove ambiguity and decision fatigue during incidents. Templates ensure consistent, compliant messaging.
- Create a one-page (or one-screen) playbook for each severity level: who gets paged (roles, order); first questions (is it real, how big, who knows); what IC declares first (status page text, account manager notification, regulatory trigger); escalation timeline.
- **SEV1 playbook**: Immediate IC + Comms + CTO notification; customer status every 5 minutes; sample message templates.
- **SEV2 playbook**: IC + Comms + tech lead notification; status every 15 minutes; decision tree for escalation to executive team.
- **SEV3 playbook**: On-call SME + Comms if customer-visible; status every 30 minutes or as resolved.
- **SEV4 playbook**: On-call SME only; update customers only if promised SLA is at risk.
- Include decision trees: is this SEV1 or SEV2? Is it our code or dependency? Escalate or containment?
- Prepare customer-communication templates pre-approved by legal and compliance: sample language for detection, impact, workaround, mitigation phases.
- Publish playbooks on wiki and embed links in incident management platform.
13. Internal and customer communications workflows (depends on: 4, 9, 12)
Specify who informs whom, in what order, via what channel. Prevent gaps like "nobody knew who was in charge for an hour."
- **Internal cadence**: First update to #incidents Slack channel within 3 minutes of declaration (even if "investigating"). Updates every 5 minutes (SEV1), 15 minutes (SEV2), 30 minutes (SEV3) or when material change occurs.
- IC calls CTO/VP Eng and incident channel lead within 1 minute of declaration (SEV1); incident declared in Slack with severity, IC name, and service affected.
- SME on-call for the failing service joins incident bridge automatically; escalation call includes them within 5 minutes.
- Designate a single Customer Communications Lead per incident (pre-identified on-call roster) who owns external messaging exclusively. Shields IC from customer contact.
- **Customer communication**: Status page updated within 3 minutes (SEV1) or 10 minutes (SEV2) even if root cause unknown; post "Investigating" with next-update ETA. Account managers of affected top-tier customers called within 5 minutes (SEV1) with templated language.
- Create a phone tree or escalation list accessible to responders; set expectation: "If you don't hear from IC in 2 minutes, call them."
- Use a single incident Slack channel per incident (auto-created by incident tool); log all communications for postmortem review.
- Define regulatory notification path: compliance must approve before sending, but do not wait for root cause; flag incidents triggering payment-processing regulations to legal immediately.
14. Status page, customer notifications, and account-manager playbook (depends on: 5, 12, 13)
Policy without tooling collapses at 3 AM. Make publishing a five-minute action.
- Upgrade or replace status page so components map to customer journeys ("payments", "settlements", "ledger") not internal services. Allow customers to subscribe per component.
- Integrate incident tool (S5) with status page so incident record drives updates and public timeline auto-populates.
- Provide one-click templates pre-filled with severity, impact language, and next-update time; reduce typing and errors.
- Create account-manager playbook: contact tree for top 50 customers, what they may say (facts only), what they must not say (speculation, blame, false ETAs), escalation path if customer escalates.
- Define SLA credit process end to end: impact detection → credit calculation (based on duration × severity) → approval → customer notification → finance treatment. Automate where possible.
- Host status page outside production failure domain on separate infrastructure so it survives total platform outage.
- Test status-page reliability during game days (S17), including simulated status-page outage and total region loss.
15. Postmortem policy, blameless process, and facilitation (depends on: 3, 4)
Only 11 of 64 action items closed means postmortems are currently a writing exercise. Rebuild around learning and tracking.
- Make postmortems mandatory: all SEV1 and SEV2, all SEV3 with customer impact or repeat pattern, any near-miss the IC flags.
- Set deadlines: draft within 5 business days, blameless review within 10 days, internal publication within 15 days.
- Adopt a single standardized template: impact and duration, timeline (detection through resolution), root cause (not "human error" but system condition that enabled error), contributing factors, what went well, what went badly, action items.
- Train a pool of blameless facilitators; require a trained facilitator (not the IC) for every SEV1 and SEV2 review.
- Prohibit counterfactual and blame language in postmortems; require contributing factors addressing tooling, process, organization, and human factors.
- Limit action items to small number of concrete, verifiable, owned changes with dates.
- Publish all postmortems internally by default, with security review only for genuinely sensitive material (e.g., unpatched vulnerability details).
- For SEV1 incidents affecting regulated customers, produce a variant customer-facing root cause report.
16. Action item tracking, reliability backlog, and completion governance (depends on: 15)
A postmortem without durable action tracking is a complaint. Solve the 11-of-64 problem.
- Create a single reliability backlog in the engineering tracker (Jira, Linear, etc.) with mandatory label, owner, due date, and link to originating incident.
- Define closure criteria: evidence required (merged code change, tested alert, verified drill) not self-reported status.
- Protect capacity: reserve a fixed percentage of each team's sprint (e.g., 10%) for reliability work; track unspent capacity and report to VP Engineering.
- Run a weekly ageing review of open actions; escalate anything overdue by >30 days to team lead and VP Engineering.
- Require IC or postmortem facilitator to sign off on action completion.
- Report completion rate and median action age in monthly incident review (target: >90% closed within 60 days).
- If the same service repeats an incident in the same area, trigger a design review rather than another action item; break the cycle.
17. Metrics, dashboards, and review cadence (depends on: 2, 3)
Measure to prove the system works. Publish dashboards so everyone sees the scoreboard.
- Define outcome metrics: time to detect (by source, target <5 min internally detected); time to mitigate SEV1/SEV2 (target <60 min); customer-detected incidents per month (target <2); SLA credits paid (target <$100K/year by month 12).
- Define process metrics: declaration latency, page acknowledgement rate, IC roster coverage (no single point of failure), first-update timeliness (% within SLA), update-cadence adherence.
- Define health metrics: alert volume and noise ratio per team, off-hours pages per engineer per month, postmortem timeliness, action closure rate and median action age.
- Build live dashboards visible to every engineer (not just managers); auto-populate from incident tool, update daily.
- Institute review cadence: weekly operational review (30 min, prior week incidents), monthly reliability review (trends, top causes, action status), quarterly executive review (CEO's office, SLA cost, systemic changes).
- Baseline every metric against S2 evidence pack; set 90-day and 12-month targets.
- End every review with decisions and owners, not just numbers.
18. Training, certification, and exercise program (depends on: 4, 12, 13, 14, 15)
A process that exists only on a wiki fails on the first real page. Build skills before deployment.
- Build curriculum: how to be on-call, how to declare an incident, how to run incidents as IC, how to communicate, how to write blameless postmortems.
- Create role-specific tracks: (1) All engineers: 30-min video on severity levels, communication expectations, postmortem format, action items (mandatory); (2) On-call responders: 2-hour workshop on playbooks (S12), alert tool (S5), escalation paths (S9), when to call manager, case studies, mobile app walkthrough; (3) Incident commanders: 4-hour IC bootcamp on leadership under pressure, decision-making, communicating with executives, status-page discipline, postmortem facilitation, practiced drills; (4) Communications leads: 2-hour training on templates, update timings, how to talk to customers, regulatory rules.
- Require certification before joining IC on-call roster: written assessment plus live simulated incident (pass/fail).
- Record videos so async teams can learn on their schedule; create runbooks and quick-reference cards (print + digital); pair new on-call engineers with experienced responder for first week.
- Run monthly tabletop exercises on realistic scenarios from the prior 12 months: region loss, ledger corruption, cascading failures.
- Run quarterly game days with intentional failure injection (database failover, status-page outage, alert tool downtime); include all on-call roles.
19. Pilot with volunteer teams (depends on: 5, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 18)
Do not roll out untested to 28 teams. Run the entire process end to end with a small cohort first.
- Recruit 3–4 volunteer teams covering a mix of criticality: one payment-path team, one ledger-adjacent team, one platform team, one low-traffic team.
- Run complete process in pilot: new severity scale (S3), consolidated tooling (S5, S6), roles (S4), escalation (S9), communications (S13, S14), postmortems (S15), action tracking (S16), paid on-call (S11), training (S18), metrics (S17).
- Instrument pilot against S17 metrics; compare results with S2 baseline.
- Hold weekly retrospectives with pilot teams; iterate on written policies, tooling, training based on feedback.
- Fix top issues found before wider rollout; document what changed and why.
- Produce pilot report with before/after numbers (MTTD, MTTR, alert noise, action completion rate) to carry into rollout conversations.
- Set explicit pilot exit criteria: rotation coverage achieved, zero unacknowledged pages over 2 weeks, all postmortems delivered on time, >80% of action items tracked.
20. Phased rollout to all 28 teams (depends on: 16, 17, 19)
Rollout is a staged migration with readiness gates, not an email announcement.
- Sequence the 28 teams into four waves of ~7 teams each, ordered by customer-impact criticality and readiness; space waves three weeks apart.
- Define per-team readiness checklist: services mapped and owned, alerts cleaned to standard (S8), runbooks written, rotation staffed, training complete, manager briefed.
- Hold gate review with program owner before each team joins; move unready teams to next wave with dated remediation plan.
- Assign named champion per wave; run internal communications cadence explaining why using pilot numbers from S19.
- Handle resistance directly: publish the "own-your-code, own-your-pager" rule and paid on-call mechanics before each wave, not after.
- Retire legacy tools, informal escalation lists, and ad-hoc status-page process at end of each wave on published cutover date.
- Harvest feedback formally at each wave and push accepted process changes through change control.
21. SOC 2 control mapping and evidence framework (depends on: 1, 3, 13, 15)
The audit will test incident response. Mapping controls now means evidence accumulates from the first real incident under the new process.
- Map the process to Trust Services Criteria for incident identification, response, evaluation, and communication of security incidents.
- Write control statements in auditor language; name a single owner per control.
- Specify the evidence artifact for each control: incident record, severity classification, paging log, communications log, postmortem, action tracker entry, training record.
- Define evidence retention and storage location (not on laptops, not on chat history that expires); plan for audit access.
- Identify which controls are blocked until certain rollout waves complete; keep a gap register with owners and review fortnightly with steering group.
- Run early walkthrough with an experienced compliance partner or pre-audit readiness team to test control design before the audit window.
22. SOC 2 dry run, gap remediation, and audit support (depends on: 20, 21)
Convert a good process into a provable one, a few months before auditors arrive. Prove the system works at scale.
- Schedule a dry run six weeks before audit window, sampling real incidents from completed waves against each control's evidence requirements.
- Remediate every gap found; prioritize anything risking a qualified opinion.
- Brief every engineer who may be interviewed so they can describe the process as practiced.
- Prepare auditor package: process documentation, sample incident records, training records, on-call schedules, action tracking register, alert quality metrics.
- Designate a single audit liaison and small evidence-request team so requests do not scatter across teams.
- Rehearse IC and Communications Lead roles under interview conditions; auditors probe realism under pressure.
- Ensure all postmortems, incidents, and evidence are retained, searchable, and accessible to auditors for the required audit period.
23. Standing governance, process ownership, and continuous improvement (depends on: 20, 22)
The classic post-audit failure is the process freezing and decaying. Lock in continuous improvement as a permanent structure.
- Establish a standing Incident Management Council chaired by the program owner, meeting monthly with engineering, support, compliance, and product representation.
- Give program owner documented mandate to change standards; require formal change control for any change to severity, roles, communications timings, or compensation.
- Re-validate severity taxonomy quarterly against real declarations; re-baseline metrics annually.
- Feed incident themes into architecture review and hiring so the program improves the system, not just the response.
- Report quarterly to executive team on metric set (MTTD, MTTR, SLA credits, customer-detected %) and top five systemic causes of incidents.
- Fund a resilience roadmap separate from incident response: reducing shared-database blast radius, multi-region failover, deploy safety.
- Refresh training and tabletop program annually and after any SEV1.
- Keep public backlog of improvement ideas (Jira, shared doc) so any engineer can suggest; publish monthly "incident newsletter" to all engineers with wins and learnings.
Previous Proposal 2 (ID: d4b0c90c-c556-49ae-945b-5c59cc4fbd11, Agent: deepseek-flash_refine_2, LLM: deepseek/deepseek-flash):
Estimated Complexity: high
Success Metrics: - Median time to detect falls from 22 minutes to under 8 minutes by month 6, and under 5 minutes for money-moving journeys by month 9.
- Share of customer-impacting incidents first detected by customers falls from 40% to under 10% by month 9.
- Median time to mitigate for SEV-1 falls from 3 h 10 min to under 60 minutes by month 9.
- 100% of SEV-1 and SEV-2 incidents have a named Triage Owner within 5 minutes and a named Incident Commander within 15 minutes, from month 4.
- Zero incidents in which command authority is unclear for more than 15 minutes, measured monthly from month 4.
- Monthly page volume falls from 3,400 to under 600, with a false-positive rate under 15%, by month 5.
- No service exceeds two pages per on-call shift for three consecutive months, measured by month 6.
- All 28 teams have a documented rotation, a service ownership map and at least one trained on-call engineer by month 6.
- Incident Commander roster holds at least 12 certified ICs covering 24x7 with no uncovered week, from month 4.
- 100% of on-call shifts are paid under a published policy from month 2.
- 100% of SEV-0, SEV-1 and SEV-2 postmortems are published internally within 15 business days, from month 5.
- Postmortem action items closed within 60 days rise from 17% to over 90%, with median age under 30 days, by month 6.
- At least one incident review or near-miss report is filed per team per quarter, from month 6.
- Status page first update is published within 30 minutes on at least 95% of SEV-1 incidents, from month 4.
- Zero missed regulatory notification windows on any incident requiring notification.
- The customer-impact ledger is complete for 100% of customer-impacting incidents from month 5, with SLA credits reconciled monthly.
- Annual SLA credits paid fall at least 50% from the $1.3M baseline within 12 months.
- Customer-impacting incidents fall at least 40% year over year from the 31-incident baseline.
- On-call satisfaction is at or above 7 out of 10, with zero voluntary attrition attributed to on-call, measured quarterly from month 6.
- Review cadence is sustained: weekly operational review in at least 90% of weeks, 12 of 12 monthly reliability reviews, 4 of 4 quarterly executive reviews.
- SOC 2 Type II is passed with zero findings related to incident response.
- 100% of new engineers complete the incident-response onboarding module within 30 days of joining.
- 100% of the 180 services have a named owner, a detection contract and at least one symptom-based alert held to the paging contract by month 7.
Steps (21):
1. Charter, mandate and the evidence clock
This step turns the CEO's email into a funded programme with one accountable owner and explicit authority, and it starts the SOC 2 clock on day one.
- Appoint a single accountable process owner — a Director of Incident Management reporting to the CTO, with a dotted line to the COO for customer and SLA matters.
- Publish a one-page charter: scope (every customer-impacting, money-moving, security and data-integrity event across all 28 teams and both regions), decision rights during an active incident, and the power to freeze deploys and override team preferences.
- Fix the funding envelope up front: tooling licences, training and drill time, and on-call compensation, with an indicative annual figure and the expected return in avoided SLA credits.
- **Start the evidence clock immediately.** A SOC 2 Type II report covers a period of operating effectiveness, so evidence cannot be backfilled; an audit in eight months means the operating period begins in the first weeks of this programme.
- Stand up a steering group of CTO, VP Engineering, Head of Support, Head of Compliance and two engineering managers, meeting fortnightly.
- Make incident-process participation a documented performance expectation for every engineering manager, not an optional extra.
- Agree the timeline explicitly: a working process in four months, evidence accumulating from month two, audit-ready by month seven.
2. Baseline evidence and problem statement (depends on: 1)
You cannot demonstrate improvement without a defensible starting point, and you cannot win the argument about noise without numbers.
- Build a 12-month incident register and re-classify all 31 customer-impacting incidents: date, severity, class, detection source, time to detect, time to mitigate, customers affected, services involved and SLA credits paid.
- Run an alert census per tool, per team and per service: total volume, page-to-action ratio, off-hours interruptions per engineer, and the 50 noisiest rules with a named owner.
- Build a silent-failure register: incidents in which no internal alert fired at all. This is the number that explains the 40% customer-detected rate.
- Survey every on-call engineer and manager on burden, fairness, escalation clarity and pay expectations, targeting above 70% response.
- Interview Support, Account Management and Sales about how customers learn of incidents, what they complain about, and what they are promised.
- Reconstruct the two command-ambiguity incidents minute by minute, from alert to mitigation, to find exactly where ownership lapsed.
- Publish the pack internally as the problem statement and retain it as management-review evidence for the audit.
3. Control mapping and evidence architecture (depends on: 1)
Most programmes leave compliance to the end. This one maps controls in the first month, because the mapping determines what the process must capture from day one.
- Map the process to the relevant Trust Services Criteria for incident identification, evaluation, response, containment, recovery and communication of events (CC7.1 to CC7.5, CC2.2, CC2.3, CC4.1, CC3.x).
- Write each control as a plain-language statement with one named owner and its evidence artifact: incident record, severity classification, paging log, communications log, postmortem, action tracker entry, training record.
- Define the **golden incident file**: one auto-assembled, single-click export per incident containing timeline, roles, severity, communications, postmortem, actions and closure evidence.
- Set retention, storage location and immutability so no control depends on a laptop, a private Slack channel or chat history that expires.
- Keep a gap register with owners and dates, reviewed fortnightly by the steering group.
- Run an early design walkthrough with the auditor's readiness team inside the first 90 days, to test the design before building on it.
4. Severity times class taxonomy (depends on: 2)
Severity alone is not enough. Severity decides how loud the response is; class decides which playbook runs and what must never be traded away.
- Define severity by impact in customer and money terms: customers affected, dollars at risk, regions lost, ledger integrity, settlement delay, data loss and SLA exposure.
- Use SEV-0 for security, privacy and regulatory events; SEV-1 for total or material loss of a payment path; SEV-2 for degradation or single-region loss; SEV-3 for limited impact with a workaround; SEV-4 for internal-only issues and near-misses.
- Define response classes: Availability, Performance, Data integrity and Ledger, Security and Privacy, Third-party dependency, Settlement and Financial, and Process failure.
- **Class can raise a response, never lower it.** A SEV-2 data-integrity incident gets SEV-1 posture, because integrity failures are not recoverable by moving faster.
- Specify automatic triggers: loss of one AWS region, ledger write failures, replication lag above threshold, a missed settlement cut-off, payment success rate below threshold for five minutes.
- State who may declare: any engineer, Support agent or account manager. State who may downgrade: the Incident Commander alone.
- Map every level to its SLA credit exposure and to the customer-visible status page state.
- Include worked examples from the last 12 months so each of the 28 teams recognises its own incidents, and re-validate the taxonomy quarterly against real declarations.
5. Roles, command and the never-without-an-owner rule (depends on: 4)
The two incidents where nobody was in charge for over an hour did not fail at declaration. They failed in the gap before it, when an alert had fired and no one owned it.
- Create one-page role cards: Triage Owner, Incident Commander, Deputy IC, Operations Lead, Internal Communications Lead, Customer Communications Lead, Scribe, Subject-Matter Responders and Executive Sponsor for SEV-1 only.
- Introduce the **Triage Owner** rule: from the moment a page is acknowledged, that person owns the incident until an IC takes over or the incident is stood down. There is never an unowned minute between first page and close.
- State the core rule plainly: the IC owns the incident, not the fix, and does not debug. An IC who starts troubleshooting has abandoned command.
- Give the IC explicit decision rights: declare and escalate severity, freeze changes, halt deploys, approve customer messaging, and pull in any engineer in the company.
- Define minimum viable staffing per severity: SEV-1 staffs every role; SEV-2 staffs IC, scribe, comms and responders; SEV-3 staffs an IC and a scribe only.
- Set handover discipline: four-hour maximum IC shifts on SEV-1, a written handover template, and a deputy named within 15 minutes of declaration.
- Set behaviour standards for responders: one incident channel, one bridge, no side channels, and asks phrased with a named owner and a time.
- Publish the role cards on the internal wiki and link them from every paging notification.
6. Declaration, lifecycle and escalation policy (depends on: 5)
This step defines the mechanical path from an alert to a declared incident and back to normal service, and it removes judgment calls from the worst moments.
- Define lifecycle states with entry and exit criteria: Detected, Triaged, Declared, Mitigated, Resolved, Postmortem, Closed — plus a Watch state with a hard 30-minute timer, after which the incident is either declared or stood down.
- **Make declaring free.** A declaration that turns out to be a false alarm is closed as a false declaration, with no blame and no follow-up, and it is tracked as a metric so the cost of caution stays visible.
- Apply the ambiguity rule: if two responders disagree about whether this is an incident, it is an incident.
- Set acknowledgement and declaration targets: page acknowledged within five minutes, triage decision within fifteen, declaration within thirty for anything customer-facing.
- Build automatic escalation ladders with timers at every layer: responder, service owner, team manager, IC on-call, VP Engineering. Escalation never requires a human decision and is never criticised.
- Define the unresponsive-team path: fifteen minutes without acknowledgement escalates to the team manager, thirty minutes to the IC, who may then direct any available engineer.
- Define change freeze and rollback authority during SEV-1 and SEV-2, and the single named condition that lifts the freeze.
- Enforce one incident, one record, with the timeline auto-captured from the channel and bridge rather than written from memory afterwards.
- Add the two-simultaneous-incidents rule: a second IC is designated, and no IC ever runs two incidents at once.
7. Detection strategy: journeys, synthetic signals and customer-report intake (depends on: 4)
Customers detected 40% of incidents first. That number is the reason this step exists, and it is fixed by measuring customer journeys rather than infrastructure.
- Define SLIs and SLOs for the top 20 customer journeys — payment initiation, authorisation, settlement, ledger read and write, API availability, webhook delivery and payout — measured per region.
- Require symptom-based alerting on those SLOs instead of cause-based alerting on infrastructure metrics.
- Add synthetic transaction monitoring from outside AWS in both regions plus a third location, on a one-minute cadence for money-moving paths.
- Add ledger and shared-PostgreSQL specific signals: replication lag, connection saturation, write latency, lock waits, transaction ID exhaustion, checkpoint pressure and bloat.
- Open a customer-report intake: Support and account managers can raise an incident directly, every customer report creates an incident record, and the customer-report path is counted as a detection source in all reporting.
- Apply the detection-gap rule: whenever a customer reports an incident before internal monitoring, a detection-gap ticket is opened automatically and owned by the responsible team.
- Publish a detection contract per service: every one of the 180 services needs an owner, at least one symptom alert, and a documented expected detect time.
- Measure current detect time per journey, set targets, and run a detection drill per team: break something in staging and see whether it pages before a human notices.
8. Three on-call rotations across 28 teams (depends on: 5)
The objection is that engineers will not carry a pager for another team's code. The answer is not to argue with it, but to build three rotations so the objection becomes structurally impossible.
- Run a **Service On-Call** rotation per team, covering only that team's own services. No engineer is ever paged for code their team does not own.
- Run a **Platform Duty** rotation for genuinely shared infrastructure: the shared PostgreSQL cluster, Kubernetes, networking, CI/CD and observability. This is nobody's product code, so it gets its own paid rotation, staffed from platform teams plus volunteers from other teams.
- Run a central **Incident Commander** roster of 12 to 16 certified senior engineers drawn from across all 28 teams, covering 24x7 on one-week shifts with a primary and a secondary.
- State the consequence honestly: 16 of 28 teams have no on-call today. Each must either build a rotation or formally transfer ownership of its services to a team that will, with the transfer dated and recorded.
- Merge small or low-traffic teams into shared rotations so no rotation has fewer than six engineers, the level below which coverage stops being sustainable.
- Cap load in the scheduling tool: no engineer is on-call more than two weeks per quarter, enforced by configuration rather than negotiation.
- Publish a coverage matrix for all 28 teams showing services, owners, rotation size and gaps, reviewed monthly.
- Write on-call participation into engineering job levels and hiring criteria, so the commitment is set before someone joins rather than after.
9. Compensation, rest and the price of opting out (depends on: 8)
Unpaid on-call is the most cited reason for resistance. The policy must be settled before rollout, not negotiated during it.
- Move immediately to paid on-call: a per-shift stipend benchmarked to the New York market, agreed with HR and Finance, and published with an effective date before any team is asked to join a new rotation.
- Pay event-based compensation for out-of-hours callouts, with a minimum call-out block and a 1.5x rate for time actually spent mitigating.
- Provide documented compensatory rest: no engineer works a normal day after a night incident, and the rest day is policy rather than a favour granted by a manager.
- Cap intrusion with a maximum number of off-hours pages per shift, and trigger a mandatory review whenever the cap is exceeded.
- Allow opt-out but **put a price on it**: an engineer may step out of a rotation, and their team must buy coverage from the paid pool at a published internal rate. This turns a cultural argument into a visible budget decision.
- Publish an explicit amnesty: incident records, near-miss reports and false declarations are never used in performance reviews. Only failure to report is a performance issue.
- Protect against gaming: pay attaches to shifts and callouts, never to incident counts.
- Review the policy every six months against real page volumes, attrition and survey results.
10. Paging contract and alert quality (depends on: 2, 7)
3,400 alerts a month at 85% noise is why engineers resent the pager. Fixing that is the price of admission for everything else in this plan.
- Publish the paging contract: every page must be symptom-based, actionable, owned, mapped to a severity and a class, and linked to a runbook. **No runbook, no page**, enforced by a CI check on the alert definition.
- Separate paging from ticketing: only customer-impacting or imminent-impact signals may wake a human; everything else becomes a ticket or a dashboard entry.
- Set a page budget per team and per service: a maximum number of pages per on-call shift. Breaching it auto-opens a remediation ticket with the engineering manager as owner.
- Put new alerts on two-week probation: a new rule runs as a ticket only and becomes a pager only after it has proved actionable, so teams stop being woken by untested rules.
- Define the default action for a noisy alert: fix, tune or delete within ten working days, with deletion treated as a legitimate and celebrated outcome and the deleted-alert count published.
- Correlate and deduplicate at the ingest pipeline so one root cause produces one page instead of forty.
- Require an expiry date on every silencing rule and temporary threshold, so suppression cannot quietly become permanent blindness.
- Run a 90-day noise sprint with a public burn-down of the top 100 noisiest rules, each owned by a named manager, and report page-to-action ratio per team monthly.
11. Incident tooling consolidation (depends on: 5, 10)
Six alerting tools and no single incident record is a structural cause of the 22-minute detection and the three-hour mitigation.
- Select one incident platform for paging, schedules, escalation policies, incident records, postmortem workflow and status-page integration, and choose it within four weeks.
- Consolidate the six alert sources into a single event pipeline feeding that platform, with deduplication, correlation and severity and class mapping applied at ingest.
- Integrate observability, runbooks, chat and voice bridges so responders see dashboards and runbooks inside the incident record and the timeline is captured automatically.
- Implement the golden incident file export defined in S3, so audit evidence is a single click rather than a reconstruction.
- Run a dual-run period alongside the legacy tools with a defined rollback, then switch the legacy tools off on a published date.
- Host the status page outside the production failure domain so it survives a total platform outage, and test that during a game day.
- Make the platform usable from a phone: an Incident Commander must be able to run a SEV-1 from a mobile device at 3am.
12. Communications: internal, customer and regulator (depends on: 4, 5)
Today the status page is written by whoever is around. This step replaces improvisation with a clock, a named owner and a pre-cleared template.
- Set the internal cadence: first update within 15 minutes of declaration, then every 30 minutes for SEV-1 and hourly for SEV-2, whether or not there is progress.
- Never let the status page be how an employee learns of an incident: internal communication leads, external follows.
- Set the customer cadence: status page updated within 30 minutes of a SEV-1 and 60 minutes of a SEV-2, with a no-new-information update still mandatory.
- Define the channel hierarchy: status page for all 2,100 customers, direct email to affected customers on SEV-1, and named account-manager calls for the top 50 accounts.
- Prepare templates per severity and class in advance, pre-approved by Legal and Compliance, each with the next-update time built in.
- Forbid speculation: customer communications never guess at cause, never assign blame, and never commit to a root cause before the postmortem.
- Define the closing communication: resolution notice, how SLA credits will be handled, and the committed date for the written report.
- Build a **regulator clock matrix**: for each event type, which regulator, which window, who signs off, and the shortest clock that drives the first action. Cover NYDFS Part 500, money-transmitter and banking notification, security breach notification, card-network rules, and public-company disclosure where applicable.
- Route every regulatory notification through Compliance, never Engineering, and pre-clear the templates.
- Publish a customer-facing root-cause report for SEV-1 incidents, especially for regulated and top-tier accounts.
- Assign a named Customer Communications Lead plus a trained deputy on every SEV-1.
13. Customer-impact ledger and SLA credit automation (depends on: 12)
The $1.3M in credits is a symptom of having no single record of customer impact. This step creates one, and makes it do four jobs at once.
- Maintain one durable customer-impact ledger per incident: which customers, which journeys, from when to when, and the estimated credit.
- Use that single record for customer communications, SLA credit computation, regulatory reporting and the annual review, so nobody reconciles two versions of the same outage.
- Automate the credit path end to end — computation, approval, customer notification and finance treatment — so credits stop being a manual scramble three weeks later.
- Reconcile accrued against paid credits monthly and report the result in the executive review.
- Track credit avoidance against programme cost, so the funding case is a number rather than an argument.
14. Postmortem policy with three artifact levels (depends on: 5)
Postmortems currently happen for some incidents, in various formats. This step makes them mandatory where they matter and light where they do not.
- Make postmortems mandatory for every SEV-0, SEV-1 and SEV-2, every SEV-3 with customer impact or a repeat pattern, every near-miss the IC flags, and every incident in which the process itself failed.
- Use three levels so the ritual matches the weight: a lightweight async incident review for SEV-3 and SEV-4, a standard facilitated postmortem for SEV-2, and a full review with an executive sponsor for SEV-1.
- Set deadlines: draft within five business days, blameless review within ten, internal publication within fifteen.
- Adopt one template: impact, timeline, detection, response, contributing factors across tooling, process, organisation and human factors, what went well, what went badly, and actions.
- Train a pool of blameless facilitators and require a trained facilitator for every SEV-1 review.
- Prohibit counterfactual and blame language in the template, and specifically ban the phrase human error as a root cause — the question is always what made the error possible.
- Publish all postmortems internally by default, with security review only for genuinely sensitive material, and produce a customer-facing root-cause variant for SEV-1.
- Treat postmortems as the learning product of the process, not as a compliance artifact.
15. Action items: capped, verifiable, with a repeat-incident rule (depends on: 14)
Eleven of 64 action items closed is not a tracking problem. It is a generation problem: the process produces more actions than the organisation can absorb.
- **Cap each postmortem at three action items.** Anything beyond three goes into a ranked reliability backlog, not into the postmortem.
- Require each action to have a named human owner, a due date, and a definition of done that is an artifact: a merged change, a new alert that fired in a drill, or a test that fails without the fix.
- Prohibit self-reported closure. Closure requires the artifact, signed off by the process owner or the Incident Commander.
- Keep one reliability backlog in the engineering tracker with a mandatory label and a link back to the originating incident.
- Reserve a fixed percentage of every team's sprint capacity for reliability work, with unspent capacity visible to the vice presidents.
- Apply the repeat-incident rule: a second incident with the same contributing factor triggers a design review owned by an architect, not another action item.
- Run a weekly ageing review of open actions and escalate anything more than 30 days overdue to the VP of Engineering.
- Target more than 90% of actions closed within 60 days and a median age under 30 days, reported monthly by team.
16. Training, certification and drills (depends on: 5, 6, 12, 14)
A process that exists only on a wiki page fails on the first real page. Skills have to be built and tested before they are needed.
- Build a short practical curriculum: how to be on call, how to triage and declare, how to run an incident as IC, how to communicate, and how to write a postmortem.
- Require certification before joining the IC roster: a written assessment plus a live simulated incident.
- Certify at least two ICs per team group, so the central roster has depth across all 28 teams and no holiday week is uncovered.
- Run monthly tabletops drawn from the last 12 months of real incidents, including region loss, ledger write failure and a missed settlement cut-off.
- Run quarterly game days with deliberately injected failure, including shared-PostgreSQL failover, status-page outage and alerting-pipeline outage.
- **Drill the process's own failure modes**, not just technical ones: IC unreachable, comms lead on PTO, two simultaneous SEV-1s, a paging storm, and a false alarm that burns an hour.
- Audit the incident process for single points of failure: who is the only person who can do each critical task, and what happens in their holiday week.
- Add a mandatory onboarding module for every engineer joining or transferring in, with completion records kept in an audit-ready form.
- Track training completion by team and publish it in the monthly reliability review.
17. Metrics and review cadence (depends on: 2, 4, 15)
Establish what good looks like, and measure it in a way that makes people report more incidents rather than fewer.
- Define outcome metrics: time to detect by source, time to mitigate by severity and class, percentage of incidents detected by customers, incidents by severity and class, SLA credits paid, and customer-impact minutes.
- Define process metrics, because the process itself has SLOs: IC assigned within five minutes, page acknowledgement rate, first-update timeliness, update-cadence adherence, postmortem timeliness, action closure rate, action median age and IC roster coverage.
- Define health metrics: alert volume and noise ratio per team, off-hours pages per engineer, on-call satisfaction, training completion, and percentage of services with a detection contract.
- **Never publish incident count as a team metric.** It rewards hiding incidents. Publish reporting metrics instead — near-misses filed, detection gaps closed, false declarations made — alongside the outcome metrics.
- Publish live dashboards visible to every engineer, refreshed daily, with each metric baselined against the S2 evidence pack and given a 90-day and a 12-month target.
- Run a 30-minute weekly operational review incident by incident, a monthly reliability review of trends and actions, and a quarterly executive review with the CEO's office covering customer impact, credits and the top five systemic causes.
- Hold a quarterly review of the process itself: what in the process wasted time, what confused responders, and what should be deleted.
- Require every review to end with decisions and named owners, never with numbers alone.
18. Pilot with three to four teams, using real incidents (depends on: 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17)
Do not roll out to 28 teams untested, and do not rely on drills alone when real incidents are available and better.
- Recruit three to four volunteer teams spanning criticality: a payment-path team, a ledger-adjacent team, the platform team that owns the shared PostgreSQL cluster, and a low-traffic team.
- Run the entire process end to end in the pilot: the new severity and class scale, roles, escalation, communications, postmortems, action tracking and paid on-call.
- Treat real incidents during the pilot as the primary training material. Hold a retro within 48 hours of each pilot incident, run by the process owner, while the friction is fresh.
- Instrument the pilot against the S2 baseline and publish before-and-after numbers.
- Hold weekly retrospectives with the pilot teams and change the written policies, the tooling and the training in response, documenting what changed and why.
- Set explicit exit criteria: rotation coverage achieved, zero unacknowledged pages over a defined period, postmortems delivered on time, and actions tracked to closure.
- Produce a pilot report that every later rollout conversation starts from.
19. Phased rollout sequenced by cost of failure (depends on: 15, 18)
Rollout is a staged migration with readiness gates, not an email announcement. The sequencing matters more than the schedule.
- Sequence the 28 teams into four waves of roughly seven, ordered by **incident density and customer-journey ownership**, not by size or alphabet: the teams whose incidents cost the most SLA credits go first, because their improvement is the visible proof.
- Leave three weeks between waves and define a per-team readiness checklist: services mapped and owned, alerts cleaned to the paging contract, runbooks written, rotation staffed to at least six engineers, training complete, manager briefed, and compensation in effect.
- Hold a gate review with the process owner before each team joins, and move an unready team to the next wave with a dated remediation plan rather than granting an exception.
- Give each wave a named champion and run an internal communications cadence that explains the why using the pilot's numbers.
- Handle resistance directly and early: publish the own-your-code-own-your-pager rule and the paid on-call mechanics before each wave, never after.
- Hold a retro within 48 hours of each wave's first incident under the new process, and feed accepted changes back through change control.
- Retire legacy tools, escalation lists and the informal status-page process at the end of each wave on a published cutover date, so two processes never run in parallel for long.
- Sequence the waves so no team faces both a rollout and the audit dry run in the same week.
20. Audit dry run and evidence review (depends on: 3, 19)
This step converts a good process into a provable one, a few months before the auditors arrive.
- Run an internal dry run six weeks before the audit window, sampling real incidents from the pilot and the early waves against each control's evidence requirements.
- Remediate every gap found, prioritising anything that risks a qualified opinion, and re-test the remediated control against the same sample.
- Brief every engineer who may be interviewed so they can describe the process as they actually practise it, not as it is written.
- Prepare the auditor package: process documentation, sample incident records, paging logs, communication logs, training records, on-call schedules and the action tracking register.
- Designate one audit liaison and a small evidence-request team, so requests do not land on 28 teams simultaneously.
- Rehearse the walkthrough with an incident commander and a communications lead, because auditors probe realism under pressure, not documents.
- Keep the audit liaison and the process owner as the same person where possible, so accountability for the control is also accountability for the evidence.
21. Standing governance, resilience roadmap and continuous improvement (depends on: 19, 20)
The classic post-audit failure is that the process freezes and then decays. This step keeps it alive and separates incident response from the engineering that prevents incidents.
- Establish a standing Incident Management Council chaired by the process owner, meeting monthly with engineering, support, compliance and product representation.
- Give the process owner a documented mandate to change standards, and require formal change control for any change to severity, classes, roles, communication timings or compensation.
- Re-validate the severity and class taxonomy quarterly against real declarations, and re-baseline the whole metric set annually.
- Fund a **resilience roadmap separate from incident response**: reduce the shared PostgreSQL blast radius, improve multi-region failover readiness, and tighten deploy safety. Better incident management cannot save a single shared ledger during a corruption event.
- Report quarterly to the executive team on the metric set and the top five systemic causes of incidents, and feed those themes into architecture review, release engineering and hiring.
- Keep a public backlog of process improvement ideas so any engineer can propose one, and change the process from it every quarter.
- Share wins publicly each quarter — faster detection, fewer customer-detected incidents, credits avoided — to sustain momentum after the audit.
- Refresh the training and tabletop programme annually and immediately after any SEV-1.
Previous Proposal 3 (ID: 985bec06-6c51-44c4-9fe4-d27c2f645d5c, Agent: qwen3.8-flash_refine_3, LLM: alibaba/qwen3.8-flash):
Estimated Complexity: high
Success Metrics: - Median Time to Detect (MTTD) < 10 minutes.
- Median Time to Mitigate (MTTR) < 45 minutes for SEV-1.
- Customer-detected incidents < 5% of total incidents.
- Monthly alert volume < 400 actionable alerts (90% reduction in noise).
- SLA credit payouts < $100k annually.
- Postmortem action item completion rate > 90% within 30 days.
- 100% of SEV-1 incidents have a designated IC and Scribe.
- On-call engineer satisfaction score > 4.0/5.0.
- Zero critical findings in SOC 2 Type II audit regarding incident response.
Steps (20):
1. Program Charter & Executive Mandate
Secure explicit authority to mandate the incident management program, transforming it from an engineering initiative into a business-critical operational standard.
* Appoint a Director of Incident Management (DIM) reporting to the CTO, with a dotted line to the Head of Compliance for audit readiness.
* Publish a charter defining scope: all customer-impacting incidents, payment path failures, and data integrity issues.
* Establish a steering committee (CTO, VP Eng, Head of Support, Legal) to resolve cross-team conflicts and approve tooling budgets.
* Define the 'Own Your Code, Own Your Pager' principle as a non-negotiable engineering value, directly addressing the 28-team resistance.
2. Baseline Audit & Evidence Framework (depends on: 1)
Create a defensible starting point for metrics and a structured evidence repository for the SOC 2 Type II audit.
* Consolidate the last 12 months of incident data into a single register, flagging the 31 customer-impacting incidents.
* Map current alert sources to specific services to quantify the 3,400 monthly alerts and identify the top 50 noisiest rules.
* Conduct a survey of on-call engineers to measure current sentiment and identify specific pain points in the unpaid rotations.
* Designate a secure, immutable storage location for all future incident artifacts (timelines, logs, postmortems) to satisfy audit requirements.
3. Severity Taxonomy & Trigger Matrix (depends on: 1)
Define unambiguous severity levels tied to specific triggers, communication SLAs, and executive notifications to eliminate classification debates.
* **SEV-1 (Critical):** System-wide outage, data loss risk, or >5% revenue-impacting transactions. Triggers immediate IC, VP Eng, and legal notification.
* **SEV-2 (Major):** Significant degradation or region-wide latency. Triggers IC and Team Lead.
* **SEV-3 (Minor):** Localized issue with workaround. Triggers on-call engineer.
* **SEV-4 (Internal):** Low-priority noise. Triggers ticket only.
* Map each severity to specific SLA credit exposures to align technical response with financial risk.
4. Incident Roles & Command Structure (depends on: 3)
Standardize the Incident Command System (ICS) to prevent the 'ambiguity of ownership' that caused previous hour-long delays.
* **Incident Commander (IC):** Single point of decision authority; does not debug. Required for SEV-1/2.
* **Scribe:** Maintains the real-time timeline log; mandatory for SEV-1.
* **Comms Lead:** Owns status page and internal broadcasts; shields IC from external noise.
* **SMEs:** Technical responders focused solely on diagnosis/mitigation under IC direction.
* Create role cards with clear 'Do' and 'Don't' responsibilities for quick reference during chaos.
5. On-Call Architecture & Compensation Policy (depends on: 4)
Establish a sustainable, paid 24x7 on-call model that respects ownership boundaries and compensates for disruption.
* Adopt a federated model: each of the 28 teams maintains its own SME rotation for owned services.
* Create a central IC rotation staffed by senior engineers across teams to ensure command coverage.
* Implement a paid policy: base stipend per shift, enhanced pay for call-outs, and mandatory recovery time after SEV-1 night shifts.
* Define the 'Unowned Service' rule: services without a dedicated team rotation must be transferred to an owning team or decommissioned.
6. Tooling Consolidation & Integration (depends on: 2, 5)
Replace the six fragmented alerting tools with a single Incident Management Platform to ensure a unified source of truth.
* Select a platform (e.g., PagerDuty + Statuspage + Slack/Teams) capable of handling 180 services.
* Integrate Kubernetes and PostgreSQL monitoring to feed alerts into the central platform.
* Automate the creation of incident channels and timelines upon alert acknowledgment.
* Ensure the status page is decoupled from the production environment to remain available during outages.
7. Alert Quality & Noise Reduction Program (depends on: 6)
Implement strict alert hygiene rules to reduce the 3,400 monthly alerts and restore trust in the pager.
* **Rule:** All paging alerts must be symptom-based (customer impact) and have a linked runbook.
* **Rule:** Implement deduplication and rate-limiting at the ingestion layer.
* **Sprint:** Conduct a 90-day noise sprint to fix, tune, or delete the top 100 noisiest rules.
* **Metric:** Enforce a 'Page-to-Action' ratio target; if an alert does not result in human action, it must be demoted to ticket-only.
8. Detection Strategy & SLOs (depends on: 3, 6)
Shift from infrastructure metrics to Service Level Objectives (SLOs) to detect customer-impacting issues faster than clients can.
* Define SLIs for critical payment paths: Transaction Success Rate, API Latency, Settlement Lag.
* Implement synthetic transaction monitoring from external vantage points in both AWS regions.
* Establish error budgets; breach of error budget triggers automatic review and potential feature freezes.
* Integrate customer-reported signals (support tickets, social media mentions) into the detection workflow as low-priority alerts.
9. Communication Protocols & Templates (depends on: 4, 6)
Define who talks, when, and what they say to stop ad-hoc status updates and protect brand reputation.
* **Status Page:** SEV-1 initial update within 15 mins; subsequent updates every 30 mins. Use legal-approved templates.
* **Internal:** IC broadcasts to #exec-leadership for SEV-1 every hour.
* **Regulatory:** Define triggers for NY financial regulator notifications; assign Legal/Compliance responsibility for filing.
* **Client Success:** Provide Account Managers with a 'Known Issues' playbook and real-time access to status page APIs for high-value accounts.
10. Postmortem Framework (Blameless) (depends on: 9)
Standardize the learning process to ensure consistent quality and adherence to the 'Blameless' culture.
* **Mandatory:** All SEV-1 and SEV-2 incidents require a postmortem within 5 business days.
* **Format:** Fixed template including Impact, Timeline, Root Cause, Contributing Factors, and Action Items.
* **Review:** Facilitated by a neutral party (not the IC) to enforce blameless language and prevent finger-pointing.
* **Publication:** All postmortems published internally on the Wiki with full searchability.
11. Action Item Tracking & Governance (depends on: 10)
Close the loop on incident resolution by enforcing rigorous tracking of corrective actions to fix the 17% completion rate.
* Automatically create Jira tickets for every action item identified in the postmortem.
* **Enforcement:** SEV-1 incident cannot be marked 'Closed' until high-priority actions are completed or deferred with VP approval.
* **Review:** Weekly review of overdue actions in the Engineering Leadership standup.
* **Metric:** Track 'Mean Time to Remediation' for action items as a key health indicator.
12. SOC 2 Control Mapping (depends on: 2, 10)
Proactively map the new incident processes to SOC 2 Trust Services Criteria to ensure audit readiness.
* Map S4 (Roles), S9 (Comms), and S10 (Postmortems) to Security and Availability criteria.
* Define 'Evidence of Operation' for each control (e.g., automated timeline logs, signed-off postmortems).
* Identify gaps between current state and audit requirements; assign remediation tasks to the DIM.
* Establish a quarterly internal compliance review to test control effectiveness before the Type II audit.
13. Training & Certification Curriculum (depends on: 4, 6)
Equip all engineers with the skills to operate within the new framework, reducing anxiety and improving response quality.
* **Universal Training:** 1-hour module on severity levels and tools for all engineers.
* **IC Certification:** Mandatory workshop and simulation for engineers joining the central IC rotation.
* **Runbook Review:** Each team must update and validate runbooks for their top 3 critical alerts.
* **Onboarding:** Include incident response basics in the engineering onboarding checklist.
14. Pilot Implementation (Wave 1) (depends on: 7, 8, 9, 13)
Deploy the new process to a controlled subset of high-traffic teams to validate assumptions before broad rollout.
* Select 3 teams: Payments Core, Ledger/API, and one Infrastructure team.
* Run the full cycle for 6 weeks: Alerts, IC handover, Comms, Postmortems.
* Collect feedback daily and adjust playbooks and tooling configurations based on real-world friction.
* Measure success using the baseline metrics from S2 (MTTD, Noise, Satisfaction).
15. Full Rollout Strategy (Waves 2-4) (depends on: 14)
Gradually migrate the remaining 25 teams using lessons learned from the pilot to manage change resistance.
* **Wave 2 (Month 3):** Deploy to 8 remaining critical customer-facing teams.
* **Wave 3 (Month 4):** Deploy to internal platform and data teams.
* **Wave 4 (Month 5):** Deploy to remaining low-traffic teams and legacy services.
* Maintain a 'Help Desk' for new on-call engineers to troubleshoot process questions during their first shift.
16. Simulations & Game Days (depends on: 15)
Test the resilience of the process and the tools under controlled failure conditions.
* **Tabletop Exercises:** Quarterly scenario walks-throughs for leadership and ICs (e.g., 'What if the primary region fails?').
* **Chaos Engineering:** Inject failures in non-production or canary environments to test alert accuracy and runbook validity.
* **Communication Drills:** Simulate SEV-1 to test the speed of status page updates and internal notification paths.
* Document findings in postmortems and create action items for identified weaknesses.
17. Metrics Dashboard & Executive Review (depends on: 2, 15)
Establish a continuous feedback loop to monitor the health of the incident management system.
* Build a real-time dashboard tracking MTTD, MTTR, Alert Volume, and SLA Credits.
* **Weekly:** Operational review of new incidents and action items with the DIM and Team Leads.
* **Monthly:** Executive review with CTO/VP Eng focusing on trends, noise ratios, and on-call fatigue.
* Publish 'Reliability Reports' to all engineers to celebrate wins and share learnings.
18. Culture & Change Management (depends on: 5, 15)
Address the cultural shift required to move from 'hero mode' to 'process mode' and mitigate pushback.
* Highlight success stories where the new process reduced toil or prevented customer churn.
* Enforce the 'Blameless' mantra by publicly correcting any leadership communication that implies blame.
* Recognize and reward effective ICs and engineers who improve runbooks or alert quality.
* Address 'Pager Fatigue' proactively by monitoring call-out volume and rotating engineers off if thresholds are breached.
19. SOC 2 Dry Run & Evidence Prep (depends on: 12, 17)
Conduct a mock audit six months out to identify gaps in evidence retention or process execution.
* Sample 10 recent incidents and verify all required artifacts (timeline, logs, postmortem, actions) are present and accessible.
* Interview on-call engineers to ensure they can describe the process and their roles without hesitation.
* Remediate any 'Control Failures' identified during the dry run.
* Prepare the 'Audit Readiness' package for the external auditors.
20. Continuous Improvement Loop (depends on: 17, 19)
Institutionalize the evolution of the incident process to prevent stagnation.
* Review the Severity Taxonomy quarterly to ensure it reflects current system behavior and customer expectations.
* Update Tooling integrations annually based on vendor roadmaps and new AWS/K8s features.
* Revise Compensation Policy annually based on market data and internal fairness reviews.
* Set annual 'Reliability Goals' for the engineering organization to drive long-term focus.
Please, considering the previous proposals as ideas that could be considered, focus on the main objective and generate an IMPROVED proposal or a completely DIFFERENT perspective if you deem it appropriate. Only if you consider any of them is amazing and impossible to improve, answer with the same proposal.
The plan has these parts:
- "steps": the list of steps, each with "step_id" (a short label unique within this proposal, such as S1, S2, S3), "title" (a short, descriptive title), "description" (what the step does and how, in Markdown: a short opening paragraph and then, if there are several concrete things to say, a bullet list) and "dependencies" (the step_ids of the other steps of THIS proposal that must be completed first; empty if none).
- "estimated_complexity": "low", "medium" or "high".
- "success_metrics": clear and measurable success metrics, one per line as a bullet list.
Think deeply about the problem, considering all aspects and implications. Develop a comprehensive understanding of what needs to be done to achieve the objective, then produce the plan.
Answer with a single JSON object, and nothing else, that conforms to this JSON schema:
{"description": "The structure every proposing agent must answer with, in every round.", "properties": {"steps": {"items": {"description": "One step of a proposed plan. Dependencies are step_ids of the same proposal.", "properties": {"step_id": {"title": "Step Id", "type": "string"}, "title": {"title": "Title", "type": "string"}, "description": {"title": "Description", "type": "string"}, "dependencies": {"items": {"type": "string"}, "title": "Dependencies", "type": "array"}}, "required": ["step_id", "title", "description", "dependencies"], "title": "Step", "type": "object", "additionalProperties": false}, "title": "Steps", "type": "array"}, "estimated_complexity": {"title": "Estimated Complexity", "type": "string"}, "success_metrics": {"title": "Success Metrics", "type": "string"}}, "required": ["steps", "estimated_complexity", "success_metrics"], "title": "ProposalContent", "type": "object", "additionalProperties": false}
{"steps": [{"step_id": "S1", "title": "Mandate, one owner, and the evidence clock", "description": "This step turns the CEO's email into a funded programme with a single accountable owner, and it starts the SOC 2 clock on day one.\n\n- Appoint a Director of Incident Management reporting to the CTO, with a dotted line to the COO for customer and SLA matters and to the Head of Compliance for audit readiness.\n- Publish a one-page charter: scope covers every customer-impacting, money-moving, security and data-integrity event across all 28 teams and both regions.\n- Grant explicit authority to declare an incident, set severity, freeze deploys, page any engineer in the company, and approve customer messaging.\n- Fund the envelope up front: tooling, training and drill time, on-call compensation, and a small programme team, roughly $500–700K a year against $1.3M in credits paid.\n- State the return plainly to the steering group: credits avoided, churn avoided, and audit findings avoided.\n\n**Start the evidence clock now.** A SOC 2 Type II report covers a period of operating effectiveness, so evidence cannot be backfilled; an audit in eight months means the operating period begins in the first weeks of this programme.\n\n- Stand up a fortnightly steering group of the CTO, VP Engineering, Head of Support, Head of Compliance and two engineering managers.\n- Open a programme risk register with the top risks, owners and review dates, and revisit it at every steering group.\n- Make participation in the incident process a documented performance expectation for every engineering manager, not an optional extra.", "dependencies": []}, {"step_id": "S2", "title": "Baseline evidence pack and cost-of-downtime model", "description": "You cannot demonstrate improvement without a defensible starting point, and you cannot win the argument about noise without numbers.\n\n- Build a 12-month incident register and re-classify all 31 customer-impacting incidents by detection source, duration, customers affected and credits paid.\n- Build the silent-failure register: incidents where no internal alert fired at all, which is the number that explains the 40% customer-detected rate.\n- Run an alert census per tool, per team and per service: volume, page-to-action ratio, off-hours interruptions, and the 50 noisiest rules with a named owner each.\n- Survey every on-call engineer and manager on burden, fairness, escalation clarity and pay expectations, targeting above 70% response.\n- Interview Support, Account Management and Sales about how customers learn of incidents and what they are promised.\n- Reconstruct the two command-ambiguity incidents minute by minute to find exactly where ownership lapsed.\n- Build a cost-of-downtime model: dollars per minute of impact per customer journey, used later to sequence teams and justify funding.\n- Publish the pack internally as the problem statement, and retain every artifact as management-review evidence for the audit.", "dependencies": ["S1"]}, {"step_id": "S3", "title": "SOC 2 control mapping and evidence architecture", "description": "Most programmes leave compliance to the end; this one maps controls in month one, because the mapping decides what the process must capture from day one.\n\n- Map the process to the Trust Services Criteria for incident identification, evaluation, response, containment, recovery and communication (CC7.1–CC7.5, CC2.2, CC2.3, CC4.1, CC3.x).\n- Write each control in plain language, with one named owner and its evidence artifact.\n- Define the golden incident file: one auto-assembled, single-click export per incident containing timeline, roles, severity, communications, postmortem, actions and closure proof.\n- Set retention, storage location and immutability so no control depends on a laptop, a private channel, or chat history that expires.\n- Keep a gap register with owners and dates, reviewed fortnightly by the steering group.\n- Meet the auditor's readiness team inside the first 90 days to test the control design before anything is built on top of it.", "dependencies": ["S1"]}, {"step_id": "S4", "title": "Severity and class taxonomy with the trigger matrix", "description": "Severity decides how loud the response is; class decides which playbook runs and what must never be traded away.\n\n- Define severity in customer and money terms: customers affected, dollars at risk, regions lost, ledger integrity, settlement delay, data loss and SLA exposure.\n- Use SEV0 for security, privacy or regulatory events; SEV1 for total or material loss of a payment path; SEV2 for degradation or single-region loss; SEV3 for limited impact with a workaround; SEV4 for internal-only issues and near-misses.\n- Define response classes: Availability, Performance, Data integrity and Ledger, Security and Privacy, Third-party dependency, Settlement and Financial, Process failure.\n\n**Class can raise the response but never lower it.** A SEV2 data-integrity incident gets SEV1 posture, because integrity failures are not recoverable by moving faster.\n\n- Specify automatic triggers: loss of one AWS region, ledger write failures, replication lag above threshold, a missed settlement cut-off, payment success rate below threshold for five minutes.\n- State who may declare — any engineer, Support agent or account manager — and who may downgrade: the Incident Commander alone.\n- Map every level to its SLA credit exposure and to the customer-visible status page state.\n- Include worked examples from the last 12 months so each of the 28 teams recognises its own incidents, and revalidate the taxonomy quarterly against real declarations.", "dependencies": ["S2"]}, {"step_id": "S5", "title": "Day-one operating rules and the minimum viable process", "description": "The full process will take months; the first useful version must be live in two weeks using the tools that already exist.\n\n- Publish ten day-one rules that need no procurement: a named owner within five minutes, one channel per incident, one register entry per incident, one person speaking to customers.\n- Apply the ambiguity rule: if two responders disagree about whether this is an incident, it is an incident.\n- Make declaring free: a false alarm is closed as a false declaration, tracked as a metric, and never criticised.\n- Require a register entry within 24 hours for every customer-impacting incident, even a minimal one.\n- Ban silent incidents: if we know, the customer hears it from us rather than from their own reconciliation.\n- Run the first 30 days on manual command, with a rotating duty Incident Commander drawn from the 12 teams that already have on-call.\n- Hold a 15-minute daily incident stand-up during month one to catch friction while it is still fresh.", "dependencies": ["S1", "S4"]}, {"step_id": "S6", "title": "Roles, command structure and the no-unowned-minute rule", "description": "The two hour-long command failures did not happen at declaration; they happened in the gap before it, when an alert had fired and nobody owned it.\n\n- Publish one-page role cards: Triage Owner, Incident Commander, Deputy IC, Operations Lead, Internal Comms Lead, Customer Comms Lead, Scribe, Subject-Matter Responders, and Executive Sponsor for SEV1 only.\n- Introduce the Triage Owner rule: whoever acknowledges the page owns the incident until an IC takes over or the incident is stood down.\n\n**The IC owns the incident, not the fix, and does not debug.** An IC who starts troubleshooting has abandoned command.\n\n- Give the IC explicit decision rights: declare and escalate severity, freeze changes, halt deploys, approve customer messaging, and pull in any engineer.\n- Define minimum viable staffing per severity: SEV1 fills every role; SEV2 staffs IC, scribe, comms and responders; SEV3 staffs an IC and a scribe.\n- Set handover discipline: four-hour maximum IC shifts on SEV1 with a written handover, and a deputy named within 15 minutes of declaration.\n- Set responder behaviour: one channel, one bridge, no side channels, and every request phrased with a named owner and a time.\n- Link the role cards from every paging notification so they are one tap away at 3 AM.", "dependencies": ["S4"]}, {"step_id": "S7", "title": "Lifecycle, declaration and escalation policy", "description": "This step defines the mechanical path from an alert to a declared incident and back to normal service, removing judgment calls from the worst moments.\n\n- Define states with entry and exit criteria: Detected, Triaged, Declared, Mitigated, Resolved, Postmortem, Closed, plus a Watch state with a hard 30-minute timer.\n- Set targets: page acknowledged within five minutes, triage decision within fifteen, declaration within thirty for anything customer-facing.\n- Build automatic escalation ladders with timers at every layer: responder, service owner, team manager, IC on-call, VP Engineering.\n- Define the unresponsive-team path: fifteen minutes escalates to the team manager, thirty minutes to the IC, who may then direct any available engineer.\n- Define change freeze and rollback authority during SEV1 and SEV2, with the single named condition that lifts the freeze.\n- Enforce one incident, one record, with the timeline captured automatically from the channel and bridge rather than written from memory afterwards.\n- Add the two-simultaneous-incidents rule: a second IC is designated, and no IC ever runs two incidents at once.\n- Test every escalation path weekly with synthetic pages, and adjust the timings after the first month of real operation.", "dependencies": ["S4", "S6"]}, {"step_id": "S8", "title": "On-call architecture across 28 teams", "description": "The objection is that engineers will not carry a pager for another team's code; the answer is to build rotations that make the objection structurally impossible.\n\n- Run a Service On-Call rotation per team, covering only that team's own services.\n- Run a Platform Duty rotation for genuinely shared infrastructure: the PostgreSQL ledger cluster, Kubernetes, networking, CI/CD and observability.\n- Run a central Incident Commander roster of 12–16 certified senior engineers on one-week shifts with a primary and a secondary.\n\n**State the consequence honestly.** Sixteen of 28 teams have no rotation today; each must build one or formally transfer service ownership to a team that has one, with the transfer dated and recorded.\n\n- Merge small or low-traffic teams into shared rotations so no rotation has fewer than six engineers.\n- Cap load in the scheduling tool: no engineer on call more than two weeks per quarter, enforced by configuration rather than negotiation.\n- Publish a coverage matrix of all 28 teams showing services, owners, rotation size and gaps, reviewed monthly.\n- Write on-call participation into engineering job levels and hiring criteria, so the commitment is set before someone joins.", "dependencies": ["S4", "S6"]}, {"step_id": "S9", "title": "Compensation, rest and the economics of opting out", "description": "Unpaid on-call is the most cited reason for resistance, so settle compensation before rollout, not during it.\n\n- Move to paid on-call: a per-shift stipend benchmarked to the New York market, agreed with HR and Finance, published with an effective date before any team is asked to join a rotation.\n- Pay for callouts at 1.5× the hourly rate for time actually spent mitigating, with a minimum block per interruption.\n- Provide documented compensatory rest: no normal working day after a night incident, and the rest day is policy rather than a favour granted by a manager.\n- Cap intrusion with a maximum number of off-hours pages per shift, and trigger a mandatory review whenever the cap is exceeded.\n\n**Allow opt-out, but put a price on it.** An engineer may step out of a rotation, and their team buys coverage from the paid pool at a published internal rate, which turns a cultural argument into a visible budget decision.\n\n- Publish amnesty: incident records, near-misses and false declarations are never used in performance reviews; only failure to report is a performance issue.\n- Protect against gaming: pay attaches to shifts and callouts, never to incident counts.\n- Check the New York labour, overtime and tax treatment with Legal and Finance before announcing, and review the policy every six months against real page volumes, attrition and survey results.", "dependencies": ["S8"]}, {"step_id": "S10", "title": "Detection strategy: journeys, synthetic signals and customer-report intake", "description": "Customers detected 40% of incidents first, which makes detection the highest-leverage business problem in this programme.\n\n- Define SLIs and SLOs for the top 20 customer journeys, measured per region: payment initiation, authorisation, settlement, ledger read and write, API availability, webhook delivery and payout.\n- Alert on symptoms against those SLOs, not on cause-based infrastructure metrics.\n- Add synthetic transaction monitoring from outside AWS in both regions plus a third location, on a one-minute cadence for money-moving paths.\n- Add ledger and shared-PostgreSQL signals: replication lag, connection saturation, write latency, lock waits, transaction ID exhaustion, checkpoint pressure and bloat.\n- Open a customer-report intake so Support and account managers can raise an incident directly, and count that path as a detection source in all reporting.\n- Apply the detection-gap rule: whenever a customer reports first, a detection-gap ticket is opened automatically and owned by the responsible team.\n- Publish a detection contract per service — owner, at least one symptom alert, documented expected detect time — for all 180 services.\n- Run a detection drill per team: break something in staging and see whether it pages before a human notices.", "dependencies": ["S4"]}, {"step_id": "S11", "title": "Paging contract and the noise-reduction programme", "description": "3,400 alerts a month at 85% noise is the reason engineers resent the pager, and fixing it is the price of admission for everything else in this plan.\n\n- Publish the paging contract: every page must be symptom-based, actionable, owned, mapped to a severity and class, and linked to a runbook.\n\n**No runbook, no page**, enforced by a CI check on the alert definition itself.\n\n- Separate paging from ticketing: only customer-impacting or imminent-impact signals may wake a human.\n- Set a page budget per team and per service, with a remediation ticket opened automatically, owned by the engineering manager, when the budget is breached.\n- Put new alerts on two-week probation as ticket-only until they have proved actionable.\n- Define the default action for a noisy alert: fix, tune or delete within ten working days, with deletion treated as a legitimate and celebrated outcome and the deleted count published.\n- Deduplicate and correlate at the ingest pipeline so one root cause produces one page instead of forty.\n- Require an expiry date on every silencing rule and temporary threshold, so suppression cannot quietly become permanent blindness.\n- Run a 90-day noise sprint with a public burn-down of the top 100 noisiest rules, each owned by a named manager.", "dependencies": ["S2", "S10"]}, {"step_id": "S12", "title": "Incident tooling consolidation and the golden incident file", "description": "Six alerting tools and no single incident record are structural causes of the 22-minute detection and the three-hour mitigation.\n\n- Choose one incident platform for paging, schedules, escalation policies, incident records, postmortem workflow and status-page integration, and choose it within four weeks.\n- Consolidate the six alert sources into a single event pipeline feeding that platform, with deduplication, correlation and severity and class mapping applied at ingest.\n- Integrate observability, runbooks, chat and voice bridges so responders see dashboards and runbooks inside the incident record, and the timeline is captured automatically.\n- Implement the golden incident file export defined in S3, so audit evidence is one click rather than a reconstruction.\n- Run a dual-run period alongside the legacy tools with a defined rollback, then switch the legacy tools off on a published cutover date.\n- Host the status page outside the production failure domain so it survives a total platform outage, and prove that in a game day.\n- Make the platform usable from a phone: an Incident Commander must be able to run a SEV1 from a mobile device at 3 AM.", "dependencies": ["S3", "S6", "S10", "S11"]}, {"step_id": "S13", "title": "Internal, customer and regulator communications", "description": "Today the status page is written by whoever is around; this step replaces improvisation with a clock, a named owner and pre-cleared templates.\n\n- Set the internal cadence: first update within 15 minutes of declaration, then every 30 minutes for SEV1 and hourly for SEV2, whether or not there is progress.\n- Never let an employee learn of an incident from the status page: internal communication leads, external follows.\n- Set the customer cadence: status page updated within 30 minutes of a SEV1 and 60 minutes of a SEV2, with a no-new-information update still mandatory.\n- Define the channel hierarchy: status page for all 2,100 customers, direct email to affected customers on SEV1, and named account-manager calls for the top 50 accounts.\n- Pre-approve templates per severity and class with Legal and Compliance, each carrying its own next-update time.\n- Forbid speculation: customer messages never guess at cause, never assign blame, and never commit to a root cause before the postmortem.\n- Define the closing communication: resolution notice, how SLA credits will be handled, and the committed date for the written report.\n- Build a regulator clock matrix covering event type, regulator, notification window, signer and the shortest applicable clock, including NYDFS Part 500, money-transmitter and banking notifications, breach notification, card-network rules and public-company disclosure.\n\n**The regulatory clock starts at awareness, not at root cause.** Route every notification through Compliance, never Engineering, and pre-clear the templates.", "dependencies": ["S6", "S7", "S12"]}, {"step_id": "S14", "title": "Customer trust workstream and the SLA credit ledger", "description": "The $1.3M in credits is a symptom of having no single record of customer impact, and the CEO's inbox is a symptom of customers learning things late.\n\n- Maintain one durable customer-impact record per incident: which customers, which journeys, from when to when, and the estimated credit.\n- Use that one record for communications, credit computation, regulatory reporting and the annual review, so nobody reconciles two versions of the same outage.\n- Automate the credit path end to end — computation, approval, customer notification and finance treatment — so credits stop being a manual scramble three weeks later.\n- Reconcile accrued against paid credits monthly and report the result in the executive review.\n- Give the status page a named product owner and map its components to customer journeys, not to internal services.\n- Send a CTO-signed reliability note to the top 50 accounts and publish a quarterly reliability report to all customers.\n- Give account managers a script of the facts they may state, the speculation they may not, and a path for customer escalations.\n- Track credit avoidance against programme cost, so the funding case stays a number rather than an argument.", "dependencies": ["S13"]}, {"step_id": "S15", "title": "Postmortems: mandatory set, three levels, blameless by design", "description": "Postmortems currently happen for some incidents, in various formats; this step makes them mandatory where they matter and light where they do not.\n\n- Make postmortems mandatory for every SEV0, SEV1 and SEV2, every SEV3 with customer impact or a repeat pattern, every near-miss the IC flags, and every incident where the process itself failed.\n- Use three levels so the ritual matches the weight: a lightweight async review for SEV3 and SEV4, a facilitated postmortem for SEV2, and a full review with an executive sponsor for SEV1.\n- Set deadlines: draft within five business days, blameless review within ten, internal publication within fifteen.\n- Adopt one template: impact, timeline, detection, response, contributing factors across tooling, process, organisation and human factors, what went well, what went badly, and actions.\n- Train a pool of blameless facilitators and require a trained one for every SEV1 review, never the IC.\n\n**Ban blame language in the template and ban \"human error\" as a root cause.** The question is always what system condition made the error possible.\n\n- Publish all postmortems internally by default, with security review only for genuinely sensitive material.\n- Produce a customer-facing root-cause report for SEV1 incidents, especially for regulated and top-tier accounts.", "dependencies": ["S6"]}, {"step_id": "S16", "title": "Action items: capped, verifiable, with the repeat-incident rule", "description": "Eleven of 64 action items closed is not a tracking problem; it is a generation problem, because the process produces more actions than the organisation can absorb.\n\n- Cap each postmortem at three action items, with anything beyond that going into a ranked reliability backlog.\n- Require every action to have a named human owner, a due date, and a definition of done that is an artifact: a merged change, an alert that fired in a drill, or a test that fails without the fix.\n- Prohibit self-reported closure; closure requires the artifact and sign-off by the process owner or the Incident Commander.\n- Keep one reliability backlog in the engineering tracker with a mandatory label and a link back to the originating incident.\n- Reserve a fixed share of each team's sprint capacity for reliability work, with unspent capacity visible to the vice presidents.\n- Apply the repeat-incident rule: a second incident with the same contributing factor triggers a design review owned by an architect, not another action item.\n- Run a weekly ageing review of open actions and escalate anything more than 30 days overdue to the VP of Engineering.\n- Report action completion rate and median action age monthly, by team.", "dependencies": ["S15"]}, {"step_id": "S17", "title": "Metrics, dashboards and the review cadence", "description": "Define what good looks like, then measure it in a way that rewards reporting incidents rather than hiding them.\n\n- Define outcome metrics: time to detect by source, time to mitigate by severity and class, share of incidents detected by customers, incidents by severity and class, SLA credits paid, and customer-impact minutes.\n- Define process metrics, because the process itself has SLOs: IC assigned within five minutes, page acknowledgement rate, first-update timeliness, update-cadence adherence, postmortem timeliness, action closure rate, action median age, and IC roster coverage.\n- Define health metrics: alert volume and noise ratio per team, off-hours pages per engineer, on-call satisfaction, training completion, and share of services with a detection contract.\n\n**Never publish incident count as a team metric.** It rewards hiding incidents; publish reporting metrics instead — near-misses filed, detection gaps closed, false declarations made.\n\n- Publish live dashboards visible to every engineer, refreshed daily, with every metric baselined against the S2 evidence pack and given a 90-day and a 12-month target.\n- Run a 30-minute weekly operational review incident by incident, a monthly reliability review of trends and actions, and a quarterly executive review with the CEO's office.\n- Hold a quarterly review of the process itself: what wasted time, what confused responders, and what should be deleted.\n- End every review with decisions and named owners, never with numbers alone.", "dependencies": ["S2", "S4", "S16"]}, {"step_id": "S18", "title": "Training, certification and the drill programme", "description": "A process that lives only on a wiki page fails on the first real page, so skills are built and tested before they are needed.\n\n- Build a practical curriculum: how to be on call, how to triage and declare, how to run an incident as IC, how to communicate, and how to write a postmortem.\n- Require certification before joining the IC roster: a written assessment plus a live simulated incident.\n- Certify at least two ICs per team group so the central roster has depth across all 28 teams and no holiday week is left uncovered.\n- Train communications leads separately on templates, cadences, customer language and the regulatory rules.\n- Run monthly tabletops drawn from the last 12 months of real incidents, including region loss, ledger write failure and a missed settlement cut-off.\n- Run quarterly game days with deliberately injected failure, including PostgreSQL failover, status-page outage and alerting-pipeline outage.\n- Drill the process's own failure modes, not just technical ones: IC unreachable, comms lead on leave, two simultaneous SEV1s, a paging storm, and a false alarm that burns an hour.\n- Audit the process for single points of failure: who alone can perform each critical task, and what happens in their holiday week.\n- Keep a mandatory onboarding module for every engineer joining or transferring in, with audit-ready completion records.", "dependencies": ["S6", "S7", "S13", "S15"]}, {"step_id": "S19", "title": "Pilot with three to four teams, using real incidents", "description": "Do not roll out to 28 teams untested, and do not rely on drills alone when real incidents are better training material.\n\n- Recruit three to four volunteer teams spanning criticality: a payment-path team, a ledger-adjacent team, the platform team that owns the shared PostgreSQL cluster, and a low-traffic team.\n- Run the entire process end to end in the pilot: severity and class, roles, escalation, communications, postmortems, action tracking, and paid on-call.\n- Treat real incidents during the pilot as the primary training material, and hold a retro within 48 hours of each one, run by the process owner while the friction is fresh.\n- Instrument the pilot against the baseline and publish before-and-after numbers.\n- Hold weekly retrospectives with the pilot teams, and change the written policies, the tooling and the training in response, documenting what changed and why.\n- Set explicit exit criteria: rotation coverage achieved, zero unacknowledged pages over two weeks, postmortems delivered on time, and actions tracked to closure.\n- Produce a pilot report that every later rollout conversation starts from.", "dependencies": ["S7", "S8", "S9", "S10", "S11", "S12", "S13", "S14", "S15", "S16", "S18"]}, {"step_id": "S20", "title": "Phased rollout to 28 teams, sequenced by cost of failure", "description": "Rollout is a staged migration with readiness gates, not an email announcement, and the sequencing matters more than the schedule.\n\n- Split the 28 teams into four waves of roughly seven, ordered by incident density and customer-journey ownership: the teams whose incidents cost the most SLA credits go first, because their improvement is the visible proof.\n- Leave three weeks between waves and define a per-team readiness checklist: services mapped and owned, alerts cleaned to the paging contract, runbooks written, rotation staffed to at least six engineers, training complete, manager briefed, compensation in effect.\n- Hold a gate review with the process owner before each team joins, and move an unready team to the next wave with a dated remediation plan rather than granting an exception.\n- Give each wave a named champion and explain the why using the pilot's numbers.\n- Handle resistance directly and early: publish the own-your-code-own-your-pager rule and the paid on-call mechanics before each wave, never after.\n- Hold a retro within 48 hours of each wave's first incident under the new process, and push accepted changes through change control.\n- Retire legacy tools, informal escalation lists and the ad-hoc status page process at the end of each wave on a published cutover date, so two processes never run in parallel for long.\n- Sequence the waves so no team faces both a rollout and the audit dry run in the same week.", "dependencies": ["S16", "S18", "S19"]}, {"step_id": "S21", "title": "Audit dry run and evidence review", "description": "This step converts a good process into a provable one, about six weeks before the auditors arrive.\n\n- Sample real incidents from the pilot and the early waves against each control's evidence requirements.\n- Remediate every gap found and re-test the remediated control against the same sample, prioritising anything that risks a qualified opinion.\n- Brief every engineer who may be interviewed so they can describe the process as they actually practise it, not as it is written.\n- Prepare the auditor package: process documentation, sample incident records, paging logs, communication logs, training records, on-call schedules, the action register and alert quality metrics.\n- Designate one audit liaison and a small evidence-request team, so requests do not land on all 28 teams at once.\n- Rehearse the walkthrough with an Incident Commander and a communications lead, because auditors probe realism under pressure.\n- Keep the audit liaison and the process owner as close to the same person as possible, so accountability for the control is also accountability for the evidence.", "dependencies": ["S3", "S20"]}, {"step_id": "S22", "title": "Standing governance and process ownership", "description": "The classic post-audit failure is that the process freezes and then decays, so ownership has to outlive the programme.\n\n- Establish a standing Incident Management Council chaired by the process owner, meeting monthly with engineering, support, compliance and product representation.\n- Give the process owner a documented mandate to change standards, with formal change control for any change to severity, classes, roles, communication timings or compensation.\n- Revalidate the severity and class taxonomy quarterly against real declarations, and re-baseline the entire metric set annually.\n- Report quarterly to the executive team on the metric set and the top five systemic causes of incidents.\n- Feed incident themes into architecture review, release engineering and hiring, so the programme improves the system and not only the response.\n- Keep a public backlog of process improvement ideas so any engineer can propose a change, and act on it every quarter.", "dependencies": ["S20", "S21"]}, {"step_id": "S23", "title": "Resilience track for the shared ledger and multi-region blast radius", "description": "Better incident management cannot save a single shared ledger during a corruption event, so this runs as a funded engineering track beside the process, not inside it.\n\n- Reduce ledger blast radius: read replicas for read paths, connection-pool isolation per service domain, statement timeouts, write throttles and per-tenant limits.\n- Build a ledger protection mode: an automatic tripwire that degrades to read-only or queues writes before corruption can spread.\n- Rehearse multi-region failover with the ledger as the binding constraint, and set a measured recovery time objective per region.\n- Test restore and point-in-time recovery on a schedule, and publish the timings.\n- Tighten deploy safety: progressive rollout, automatic rollback on SLO burn, and change-freeze windows around settlement.\n- Track a small number of resilience bets with named owners and quarterly milestones, reviewed in the same executive review as the incident metrics.", "dependencies": ["S1", "S17"]}, {"step_id": "S24", "title": "Culture, incentives and continuous improvement", "description": "Incident management is a cultural change before it is a process change, and the incentives decide whether reporting goes up or down.\n\n- Restate the amnesty at every rollout wave: reporting, near-misses and false declarations are safe; hiding is not.\n- Recognise detection improvements, deleted alerts, closed detection gaps and well-written postmortems publicly, and reward the engineers who do them.\n- Correct blame language in leadership communication on the spot, including inside the executive review.\n- Watch for pager fatigue with real data, and take an engineer off rotation as soon as the intrusion cap breaks.\n- Run a quarterly process retro with the whole engineering organisation and publish what changed as a result.\n- Refresh the training and tabletop programme annually and immediately after any SEV1.\n- Share wins publicly each quarter — faster detection, fewer customer-detected incidents, credits avoided — to hold momentum after the audit.", "dependencies": ["S20", "S22"]}], "estimated_complexity": "high", "success_metrics": "- Median time to detect falls from 22 minutes to under 8 minutes by month 6, and under 5 minutes for money-moving journeys by month 9.\n- Share of customer-impacting incidents first detected by customers falls from 40% to below 10% by month 9.\n- Median time to mitigate for SEV1 falls from 3 h 10 min to under 60 minutes by month 9.\n- 100% of SEV1 and SEV2 incidents have a named Triage Owner within 5 minutes and a named Incident Commander within 15 minutes, from month 4.\n- Zero incidents with unclear command authority lasting more than 15 minutes, measured monthly from month 4.\n- Monthly page volume falls from 3,400 to under 600, with a false-positive rate below 15%, by month 5.\n- No service exceeds two pages per on-call shift for three consecutive months by month 6.\n- 100% of the 180 services have a named owner, a detection contract and at least one symptom-based alert by month 7.\n- All 28 teams have a documented rotation, a service ownership map and at least one trained on-call engineer by month 6.\n- 100% of on-call shifts are paid under a published policy from month 2, with zero on-call-attributed voluntary attrition by month 6.\n- On-call satisfaction scores 7 out of 10 or better in quarterly surveys from month 6.\n- The IC roster holds at least 12 certified ICs covering 24x7 with no uncovered week from month 4.\n- 100% of SEV0, SEV1 and SEV2 postmortems are published internally within 15 business days from month 5.\n- Action items closed within 60 days rise from 17% to above 90%, with a median action age under 30 days, by month 6.\n- Status page first update is posted within 30 minutes on at least 95% of SEV1 incidents from month 4.\n- Zero missed regulatory notification windows on any incident requiring notification.\n- The customer-impact record is complete for 100% of customer-impacting incidents from month 5, with SLA credits reconciled monthly.\n- Annual SLA credits paid fall at least 50% from the $1.3M baseline within 12 months.\n- Customer-impacting incidents fall at least 40% year over year from the 31-incident baseline.\n- Review cadence is sustained: weekly operational review in at least 90% of weeks, 12 of 12 monthly reliability reviews, 4 of 4 quarterly executive reviews.\n- Every page has a recorded disposition — fixed, tuned or deleted — within 10 working days, from month 4.\n- SOC 2 Type II is passed with zero findings related to incident response.\n- 100% of new engineers complete the incident-response onboarding module within 30 days of joining."}P3 discarded all step content and submitted 22 titles with a single sentence of rationale each. Every operational detail it had in round 1 — severity thresholds, status-page timings, compensation mechanics, wave schedule, training curriculum — is gone, and several metrics were loosened or stripped of dates.
- Step 1 now names the SOC 2 evidence clock, acknowledging that Type II evidence cannot be backfilled.
- Step 3 is retitled "Severity & Class Taxonomy", nodding to the class dimension.
- MTTD target tightened from <10 to <5 minutes.
- All step bullets removed: no SEV-1 definition, no status-page timing, no stipend structure, no wave schedule, no IC certification content — the plan is now an outline, not a process.
- Metrics loosened: SEV-1 MTTR from <45 to <60 minutes, alert volume from <400 to <600, action completion from >90% to >80%; no metric carries a date or month.
- "40+ certified Incident Commanders" is retained with no justification and conflicts with the 12–16 roster both other plans converged on.
- Dropped concrete round-1 content that was its own: the 6-week pilot with named teams (Payments Core, Ledger/API), the Wave 2–5 month schedule, the new-on-call Help Desk, error budgets triggering feature freezes.
- Ordering is now weaker: SOC 2 control mapping sits at step 20 after full rollout (step 19), despite depending only on 1 and 3 — evidence design arrives after eight months of incidents have already been handled.
- Step 5 (tooling) depends only on step 1, so platform selection precedes the severity taxonomy and alert-quality rules it must encode.
- Proposal 2 : The SOC 2 evidence clock starting on day one.
- Proposal 2 : Severity paired with a response class.
- Proposal 2 : An IC roster of 12–16 certified senior engineers.
- Proposal 2 : The Triage Owner rule and the three-action-item cap.
+ Severity & Class Taxonomy+ Incident Command Roles & Structure+ Detection Strategy & SLO Implementation+ Communication Workflows (Internal & External)Severity taxonomy and trigger matrixIncident roles, command structure, and decision rightsAlert consolidation and event pipelineDetection strategy: SLOs, signals, and customer-journey monitoringInternal and customer communications workflows
The plan produced
1. Executive Mandate & SOC 2 Evidence Clock from P2 step 1
Establish clear ownership and funding to transform the CEO's concern into a governed program with immediate compliance implications.
2. Baseline Measurement & Problem Statement (after 1) from P2 step 2
Quantify the current state to build a defensible problem statement and establish the data foundation for improvement tracking.
3. Severity & Class Taxonomy (after 2) from P2 step 4
Define unambiguous incident levels based on business impact to trigger specific response protocols.
4. Incident Command Roles & Structure (after 3)
Standardize roles and decision rights to eliminate command ambiguity during active incidents.
5. Tooling Consolidation & Integration (after 1)
Deploy a unified incident management platform to replace fragmented alerting tools and create a single source of truth.
6. Detection Strategy & SLO Implementation (after 3)
Shift monitoring focus to customer-impacting symptoms to detect issues before customers do.
7. Alert Quality Standards & Noise Reduction (after 5, 6)
Implement strict alerting rules to reduce noise and restore on-call trust.
8. On-Call Architecture & Coverage Model (after 4) from P1 step 10
Design a federated, team-owned on-call structure supported by central command coverage.
9. On-Call Compensation & Wellbeing Policy (after 8)
Introduce paid, sustainable on-call rotations to remove resistance and mitigate burnout.
10. Playbooks & Communication Templates (after 3, 4) from P1 step 12
Create standardized playbooks and templates for each severity level to reduce cognitive load.
11. Escalation & Incident Lifecycle Policy (after 4, 5) from P1 step 9
Define the mechanical progression from alert to resolution, including automatic escalation paths.
12. Communication Workflows (Internal & External) (after 4, 10)
Establish timed protocols for internal, customer, and regulatory communications.
13. Status Page & Customer Notification Automation (after 5, 12) from P1 step 14
Integrate incident management with public status pages to ensure timely and accurate updates.
14. Postmortem Policy & Blameless Framework (after 3) from P1 step 15
Mandate and standardize postmortems to focus on system improvement rather than individual error.
15. Action Item Tracking & Governance (after 14)
Ensure postmortem action items are tracked, owned, and completed to prevent recurrence.
16. Metrics, Dashboards & Review Cadence (after 2, 3) from P1 step 17
Define and monitor key performance indicators to validate the effectiveness of the incident process.
17. Training, Certification & Drills (after 4, 10, 12) from P2 step 16
Build practical skills and readiness for incident response across all roles.
18. Pilot Program with Volunteer Teams (after 5, 7, 9, 11, 13, 15, 17) from P1 step 19
Test the end-to-end process on a small scale to refine procedures before full rollout.
19. Phased Rollout to All Teams (after 18) from P1 step 20
Deploy the process to all 28 teams in manageable waves to minimize disruption and manage resistance.
20. SOC 2 Control Mapping & Evidence Framework (after 1, 3) from P1 step 21
Align the incident management process with SOC 2 requirements to ensure audit readiness.
21. SOC 2 Dry Run & Audit Support (after 19, 20)
Verify evidence sufficiency and prepare the team for the external audit.
22. Governance & Continuous Improvement (after 19, 21)
Establish standing governance to sustain the process and drive ongoing improvements.
- Median Time to Detect (MTTD) < 5 minutes.
- Median Time to Mitigate (MTTR) < 60 minutes for SEV1.
-
90% of incidents detected internally (vs. 40% previously).
- SLA credits reduced to <$100K annually.
- Monthly alert volume < 600 with >95% signal-to-noise ratio.
- Postmortem action item completion rate > 80%.
- Zero incidents with command ambiguity > 15 minutes.
- 100% of SEV1/2 postmortems published within 15 days.
- Status page updates within 30 minutes for 95% of SEV1.
- 100% of teams integrated into on-call rotations.
- On-call satisfaction score > 7/10.
- 40+ certified Incident Commanders available 24x7.
- SOC 2 Type II audit passed with zero findings.
[SYSTEM]
You are an expert assistant in complex project planning.
Your task is to generate a detailed and structured action plan to reach the main objective in the most professional and most detailed way, no matter how much work or steps will be needed to perform.
Use your internal reasoning processes to think deeply about the problem and create the most comprehensive plan possible.
Take as much time and space as you need to think through all aspects of the problem.
After your thorough analysis, answer with the plan in the requested structure.
Every text field you write will be read by a busy person, so write for them: short sentences, one idea each, in short paragraphs. Text fields accept Markdown: separate paragraphs with a blank line, use a bullet list when you enumerate things and plain paragraphs when you explain or argue, and bold at most one key phrase per paragraph. A text of more than three sentences must be split into paragraphs; never deliver a single unbroken block of text.
[HUMAN]
Main Objective: "A B2B payments platform in New York (260 engineers in 28 teams, 2,100 customers, $4B processed a month, a 99.95% SLA with service credits) runs about 180 services on Kubernetes across two AWS regions, with a shared PostgreSQL cluster for the ledger. In the last 12 months: 31 customer-impacting incidents, median time to detect 22 minutes (customers detected 40% of them first), median time to mitigate 3 h 10 min, $1.3M paid in SLA credits, two incidents in which nobody was sure who was in charge for over an hour, and a CEO email about "outages we hear about from clients". On-call exists in 12 of the 28 teams, unpaid, fed by alerts from six different tools — 3,400 alerts a month, 85% of them noise; there is no severity scale, status-page updates are written by whoever is around, and postmortems happen for some incidents, in various formats, with action items that are rarely tracked (11 of 64 closed). A SOC 2 Type II audit in eight months will test incident response. Engineers push back against "carrying a pager for other teams' code".
Define the incident management process: severity levels and what each one triggers; roles (incident commander, communications, scribe, subject-matter responders) and how they are staffed 24x7 across 28 teams; the on-call structure, rotations, compensation and the rules for alert quality; detection and escalation paths; internal and customer communications (status page, account managers, regulators when required) with their timings; postmortems (when mandatory, format, blameless review, ownership and tracking of actions); the metrics and reviews that show whether it works; and how it is introduced across the teams without waiting for the audit."
For your consideration and refinement, here are proposals from the previous round:
Previous Proposal 1 (ID: 2c50b755-66c6-49f5-aea8-767330e89bc3, Agent: claudeHaiku4.5_refine_1, LLM: anthropic/claude-haiku-4-5):
Estimated Complexity: high
Success Metrics: - Median time to detect reduced from 22 minutes to under 5 minutes by month 6, with >90% internal detection (vs. 40% customer-detected now).
- Median time to mitigate reduced from 3h 10min to under 60 minutes for SEV1 and SEV2 by month 9.
- Customer-impacting incidents detected by customers drop from 40% to <5% of all customer-impacting incidents.
- Alert volume reduced from 3,400 per month to <600 per month; signal-to-noise ratio improves from 15:85 to >95:5.
- SLA credits paid reduced from $1.3M to <$100K annually by month 12.
- Zero incidents with command-and-control ambiguity lasting >15 minutes; all SEV1/2 incidents have named IC logged in timeline within 5 minutes.
- Postmortem action item completion rate reaches >80% (from 11 of 64, or 17%) by month 4.
- 100% of SEV1 and SEV2 postmortems published within 15 business days by month 5.
- All 28 teams integrated into incident management system with active on-call rotations by week 20; no team unresponsive to pages for >30 minutes.
- On-call satisfaction score reaches >7/10 on survey; zero on-call-attributed voluntary attrition by month 6.
- Incident commander roster: 40+ certified ICs covering 24x7 with no single point of failure by month 4.
- Status-page first update published within 30 minutes on ≥95% of SEV1 incidents by month 3.
- SOC 2 Type II audit passes incident response controls with zero findings by month 8.
- Weekly incident review cadence sustained in ≥90% of weeks; monthly reliability reviews 12 of 12; quarterly executive reviews 4 of 4.
- 100% of new engineers complete incident-response onboarding within 30 days of joining.
Steps (23):
1. Executive mandate and governance structure
Turn the CEO email into a funded, authorized program with clear ownership and decision rights. Without executive backing, every downstream decision stalls in negotiation.
- Appoint a single program owner (e.g., Director of Incident Management) reporting to the CTO and COO.
- Publish a one-page charter covering scope (all customer-impacting incidents), authority to override team preferences during incidents, and funding for tooling, training, and on-call compensation.
- Establish a standing Incident Management Steering Group with CTO, VP Engineering, VP Support, Head of Compliance, and one engineering manager per region meeting monthly.
- Secure budget envelope: tool licenses, training time, incident-response infrastructure, and on-call compensation (estimated $400–600K annually).
- Make incident-response participation a documented performance expectation for all engineering managers and team leads.
2. Baseline measurement and evidence pack (depends on: 1)
You cannot prove improvement without defensible baseline numbers, and you cannot win arguments about noise or impact without data.
- Build a 12-month incident register: date, detection source, impact scope, time to detect, time to mitigate, SLA credits paid, and services involved.
- Audit the current alert estate: count alerts per tool, per team, per service; compute page-to-action ratio; identify top 50 noisiest rules and off-hours interruptions per engineer.
- Survey on-call engineers and managers on burden, fairness, and escalation clarity (target >70% response rate).
- Interview Support and Account Management: how do customers discover incidents, what do they complain about?
- Document the two command-ambiguity incidents: exactly when unclear who was in charge and why, how long it lasted.
- Publish this pack internally as the problem statement and retain all artifacts for SOC 2 audit evidence.
3. Severity taxonomy and trigger matrix (depends on: 2)
Severity is the keystone. Every other rule—paging, communications, postmortems, compensation—keys off it. Define four levels plus a special SEV0 for security/regulatory events.
- **SEV1 (Critical)**: Complete service outage, data corruption, or >5% payment-path failure rate for >5 min. Every minute costs money. IC required; 99.99% uptime threatened.
- **SEV2 (Major)**: Significant degradation, single region loss, or 1–5% transaction failure. IC typically required; service credit exposure.
- **SEV3 (Minor)**: Limited customer impact with workaround available, or internal issues affecting operations. On-call SME + escalation if SLA at risk.
- **SEV4 (Cosmetic)**: Observations, non-impacting bugs, alerts. Alert-driven, no escalation unless pattern emerges.
- Specify automatic triggers: region loss, ledger write failures, missed settlement window, payment success rate thresholds.
- Define who may declare (any engineer, Support, account manager) and who may downgrade (IC only).
- Include worked examples from the last 12 months so teams recognize their incidents in the definitions.
- Review and re-validate quarterly against actual declarations.
4. Incident roles, command structure, and decision rights (depends on: 3)
The two incidents with >1 hour of command ambiguity prove this step is non-negotiable. Define clear roles with explicit decision authority.
- **Incident Commander**: Owns the incident timeline, not the fix. Declares severity, decides escalation, approves communications, freezes changes, calls responders. Non-technical ICs are acceptable.
- **Deputy IC**: Shadows IC; takes over if IC unavailable. Nominated within 5 minutes of incident declaration.
- **Communications Lead**: Owns internal Slack updates and status-page messaging. Shields IC from interruptions.
- **Scribe**: Records real-time timeline: decisions, who did what, key timestamps. Not responsible for fixing.
- **Subject-Matter Responders**: Engineers with service context. Take IC direction without debate. Report only to IC.
- **Operations Lead** (SEV1 only): Coordinates across multiple responders, manages incident bridge.
- Write one-page role cards with mission, decision authority, and escalation upward. Publish on wiki and link from every paging notification.
- Define minimum viable coverage per severity: SEV1 staffs all roles; SEV2 staffs IC, Comms, Scribe; SEV3 staffs IC + Scribe.
- Establish handover discipline: maximum 4-hour IC shifts on SEV1, written handover template required.
5. Incident tooling consolidation and integration (depends on: 1, 3)
Six alert tools and ad-hoc incident records are structural causes of the 22-minute detection and 3+ hour mitigation. Consolidate to a single incident platform that is the source of truth.
- Select an incident management platform (PagerDuty, Opsgenie, Incident.io, etc.) that supports paging, schedules, escalation, incident records, and postmortem workflow.
- Requirement: the platform must integrate with observability tools, auto-create and pin incident channels in Slack, auto-capture timeline from chat, and support API-driven playbook automation.
- Plan a dual-run period alongside legacy tools with a published cutover date; define rollback criteria.
- Integrate incident record with the 180 services' monitoring and dashboards so responders see everything in one place.
- Define data retention and audit trail to satisfy SOC 2 evidence requirements: who did what, when, under whose authority.
- Budget for licenses, migration effort, and two-week hardening period post-cutover.
6. Alert consolidation and event pipeline (depends on: 5)
Replace six alert sources with a single ingestion point. Deduplicate and route alerts with minimal manual judgment, removing a major source of detection delay.
- Consolidate alert endpoints from six tools into a single event pipeline; this often sits in front of the incident platform (S5).
- Implement deduplication and correlation so a single outage triggering alerts from five monitoring tools produces one page, not five.
- Map every alert to a severity level from S3 (SEV1, SEV2, SEV3, SEV4) at ingestion.
- Log every alert for postmortem analysis and trending.
- Ensure the platform's mobile app works reliably; on-call responders need to engage from any device.
7. Detection strategy: SLOs, signals, and customer-journey monitoring (depends on: 3, 6)
Customers detected 40% of incidents first—a detection gap that must be closed. Build symptom-based alerting that detects outages before customers do.
- Define SLIs and SLOs for the top 20 customer journeys: payment initiation, settlement, ledger read/write, API availability, webhook delivery, measured per region.
- Require symptom-based alerting on SLOs, not cause-based infrastructure metrics (e.g., alert on "payment success rate <99%" not "database CPU >80%").
- Add synthetic transaction monitoring from outside AWS in both regions and a third location, one-minute cadence, for all money-moving paths.
- Add ledger-critical signals: PostgreSQL replication lag, connection saturation, write latency, transaction ID exhaustion, checkpoint pressure.
- Create a detection contract per service: owner identified, at least one symptom alert defined, expected detect time documented.
- Open a customer-reported incident path: Support and account managers can declare an incident directly, counted as a detection source in metrics.
- Fund a separate resilience roadmap to reduce shared-database blast radius, because detection improvements do not protect against ledger corruption.
8. Alert quality standards and noise-reduction program (depends on: 3, 6, 7)
3,400 monthly alerts with 85% noise is the reason engineers resent the pager. Cutting noise is the price of admission for on-call buy-in.
- Publish alert standards: every page must be symptom-based, have an immediate runbook action, be owned by a team, and map to a severity level. No runbook, no page.
- Separate paging from ticketing: only customer-impacting or imminent-impact signals trigger pages. Everything else becomes a dashboard entry or a ticket.
- Set a noise budget per team and service: no service may exceed two pages per on-call shift per month. Breaching triggers a mandatory alert quality review.
- Define the default action for a noisy alert: fix the root cause, tune the threshold, or delete it within ten working days. Deletion is a legitimate successful outcome.
- Implement automatic suppression rules: silence alerts if service auto-recovered within 30 seconds; suppress known maintenance windows; group flapping alerts (>5 in 2 min) into one page; rate-limit noisy services (max 1 alert per 5 min until condition clears).
- Require expiry dates on all silencing rules so suppression cannot become permanent blindness.
- Run a 90-day noise sprint with a visible burn-down of the top 100 noisiest rules, each owned by a named manager.
9. Escalation policies and incident lifecycle (depends on: 3, 4, 5, 6)
Define the mechanical path from alert to incident declaration to resolution. Escalation must be automatic and blameless.
- Define incident lifecycle states with clear entry/exit criteria: Detected → Triaged → Declared → Mitigated → Resolved → Postmortem → Closed.
- Set acknowledgement targets: page acknowledged in 5 minutes; triage decision (is this real?) in 15 minutes; severity declaration (is this customer-impacting?) in 30 minutes.
- Build escalation ladders: if responder does not acknowledge in 5 min, page escalates to service owner, then team manager, then IC on-call. Escalation is automatic, not manual.
- Implement escalation for SEV1: IC paged via phone call + SMS + Slack + mobile; if not acknowledged in 2 min, Deputy IC paged simultaneously; Communications Lead pinged at same time.
- Define change freeze during SEV1 and SEV2: no deployments except to fix the incident. Freeze lifts only when mitigation is confirmed.
- Enforce one incident, one record: the incident record is the sole source of truth. Auto-capture timeline from Slack and bridge; never write timeline from memory later.
- Test all escalation paths weekly via synthetic page to on-call; adjust timings based on first month of operations.
10. On-call architecture and 24x7 coverage model (depends on: 3, 4, 9)
The answer to "carrying a pager for another team's code" is that every team carries its own, and the platform carries shared risk. Design a sustainable model.
- Adopt a federated model: every service has one owning team; that team's on-call carries its service's pager. No team is paged for code it does not own.
- State the consequence clearly: 16 of 28 teams currently have no on-call. They must either build one or formally transfer service ownership to a team that will.
- Create a separate, centrally staffed Incident Commander on-call roster drawn from trained senior engineers across all teams, covering 24x7 with no single point of failure.
- Merge small or low-traffic teams into shared rotations so no rotation has fewer than six engineers (below which coverage is unsustainable).
- Define primary and secondary per slot: secondary engages only on no-acknowledge or explicit request from IC.
- Align coverage across two AWS regions and New York business hours: one global IC rotation; service on-call aligned to service users' time zones.
- Define unresponsive-team escalation: 15 min without acknowledgement escalates to team manager; 30 min escalates to IC, who may direct any available engineer.
- Publish a coverage matrix of all 28 teams showing services, rotation size, and gaps, reviewed monthly.
11. On-call compensation, wellbeing, and sustainability policy (depends on: 10)
Unpaid on-call is the most cited reason for resistance. Settle compensation before rollout, not during it. Make it sustainable.
- Introduce paid on-call: a weekly stipend ($500–1,000) while on-call, regardless of incident volume, benchmarked to New York market rates.
- Pay event-based compensation: 1.5× hourly rate for time spent mitigating incidents during off-hours (minimum one-hour block per callout).
- Provide compensatory rest: no engineer works a normal business day after a night incident requiring >2 hours mitigation. Rest day is documented, not granted as a favor.
- Cap intrusion: define maximum off-hours pages per week (e.g., no more than three per shift). Mandatory review and escalation if exceeded.
- Offer a voluntary opt-out path for engineers with genuine constraints, balanced by explicit obligation that someone else is paid to cover.
- Include on-call expectation and compensation in job descriptions and hiring conversations so commitment is understood before joining.
- Publish the policy with an effective date before any team is asked to join a rotation.
- Review the policy every six months against actual page volumes, attrition rates, and survey feedback.
12. Playbooks and communication templates by severity (depends on: 3, 4)
Playbooks remove ambiguity and decision fatigue during incidents. Templates ensure consistent, compliant messaging.
- Create a one-page (or one-screen) playbook for each severity level: who gets paged (roles, order); first questions (is it real, how big, who knows); what IC declares first (status page text, account manager notification, regulatory trigger); escalation timeline.
- **SEV1 playbook**: Immediate IC + Comms + CTO notification; customer status every 5 minutes; sample message templates.
- **SEV2 playbook**: IC + Comms + tech lead notification; status every 15 minutes; decision tree for escalation to executive team.
- **SEV3 playbook**: On-call SME + Comms if customer-visible; status every 30 minutes or as resolved.
- **SEV4 playbook**: On-call SME only; update customers only if promised SLA is at risk.
- Include decision trees: is this SEV1 or SEV2? Is it our code or dependency? Escalate or containment?
- Prepare customer-communication templates pre-approved by legal and compliance: sample language for detection, impact, workaround, mitigation phases.
- Publish playbooks on wiki and embed links in incident management platform.
13. Internal and customer communications workflows (depends on: 4, 9, 12)
Specify who informs whom, in what order, via what channel. Prevent gaps like "nobody knew who was in charge for an hour."
- **Internal cadence**: First update to #incidents Slack channel within 3 minutes of declaration (even if "investigating"). Updates every 5 minutes (SEV1), 15 minutes (SEV2), 30 minutes (SEV3) or when material change occurs.
- IC calls CTO/VP Eng and incident channel lead within 1 minute of declaration (SEV1); incident declared in Slack with severity, IC name, and service affected.
- SME on-call for the failing service joins incident bridge automatically; escalation call includes them within 5 minutes.
- Designate a single Customer Communications Lead per incident (pre-identified on-call roster) who owns external messaging exclusively. Shields IC from customer contact.
- **Customer communication**: Status page updated within 3 minutes (SEV1) or 10 minutes (SEV2) even if root cause unknown; post "Investigating" with next-update ETA. Account managers of affected top-tier customers called within 5 minutes (SEV1) with templated language.
- Create a phone tree or escalation list accessible to responders; set expectation: "If you don't hear from IC in 2 minutes, call them."
- Use a single incident Slack channel per incident (auto-created by incident tool); log all communications for postmortem review.
- Define regulatory notification path: compliance must approve before sending, but do not wait for root cause; flag incidents triggering payment-processing regulations to legal immediately.
14. Status page, customer notifications, and account-manager playbook (depends on: 5, 12, 13)
Policy without tooling collapses at 3 AM. Make publishing a five-minute action.
- Upgrade or replace status page so components map to customer journeys ("payments", "settlements", "ledger") not internal services. Allow customers to subscribe per component.
- Integrate incident tool (S5) with status page so incident record drives updates and public timeline auto-populates.
- Provide one-click templates pre-filled with severity, impact language, and next-update time; reduce typing and errors.
- Create account-manager playbook: contact tree for top 50 customers, what they may say (facts only), what they must not say (speculation, blame, false ETAs), escalation path if customer escalates.
- Define SLA credit process end to end: impact detection → credit calculation (based on duration × severity) → approval → customer notification → finance treatment. Automate where possible.
- Host status page outside production failure domain on separate infrastructure so it survives total platform outage.
- Test status-page reliability during game days (S17), including simulated status-page outage and total region loss.
15. Postmortem policy, blameless process, and facilitation (depends on: 3, 4)
Only 11 of 64 action items closed means postmortems are currently a writing exercise. Rebuild around learning and tracking.
- Make postmortems mandatory: all SEV1 and SEV2, all SEV3 with customer impact or repeat pattern, any near-miss the IC flags.
- Set deadlines: draft within 5 business days, blameless review within 10 days, internal publication within 15 days.
- Adopt a single standardized template: impact and duration, timeline (detection through resolution), root cause (not "human error" but system condition that enabled error), contributing factors, what went well, what went badly, action items.
- Train a pool of blameless facilitators; require a trained facilitator (not the IC) for every SEV1 and SEV2 review.
- Prohibit counterfactual and blame language in postmortems; require contributing factors addressing tooling, process, organization, and human factors.
- Limit action items to small number of concrete, verifiable, owned changes with dates.
- Publish all postmortems internally by default, with security review only for genuinely sensitive material (e.g., unpatched vulnerability details).
- For SEV1 incidents affecting regulated customers, produce a variant customer-facing root cause report.
16. Action item tracking, reliability backlog, and completion governance (depends on: 15)
A postmortem without durable action tracking is a complaint. Solve the 11-of-64 problem.
- Create a single reliability backlog in the engineering tracker (Jira, Linear, etc.) with mandatory label, owner, due date, and link to originating incident.
- Define closure criteria: evidence required (merged code change, tested alert, verified drill) not self-reported status.
- Protect capacity: reserve a fixed percentage of each team's sprint (e.g., 10%) for reliability work; track unspent capacity and report to VP Engineering.
- Run a weekly ageing review of open actions; escalate anything overdue by >30 days to team lead and VP Engineering.
- Require IC or postmortem facilitator to sign off on action completion.
- Report completion rate and median action age in monthly incident review (target: >90% closed within 60 days).
- If the same service repeats an incident in the same area, trigger a design review rather than another action item; break the cycle.
17. Metrics, dashboards, and review cadence (depends on: 2, 3)
Measure to prove the system works. Publish dashboards so everyone sees the scoreboard.
- Define outcome metrics: time to detect (by source, target <5 min internally detected); time to mitigate SEV1/SEV2 (target <60 min); customer-detected incidents per month (target <2); SLA credits paid (target <$100K/year by month 12).
- Define process metrics: declaration latency, page acknowledgement rate, IC roster coverage (no single point of failure), first-update timeliness (% within SLA), update-cadence adherence.
- Define health metrics: alert volume and noise ratio per team, off-hours pages per engineer per month, postmortem timeliness, action closure rate and median action age.
- Build live dashboards visible to every engineer (not just managers); auto-populate from incident tool, update daily.
- Institute review cadence: weekly operational review (30 min, prior week incidents), monthly reliability review (trends, top causes, action status), quarterly executive review (CEO's office, SLA cost, systemic changes).
- Baseline every metric against S2 evidence pack; set 90-day and 12-month targets.
- End every review with decisions and owners, not just numbers.
18. Training, certification, and exercise program (depends on: 4, 12, 13, 14, 15)
A process that exists only on a wiki fails on the first real page. Build skills before deployment.
- Build curriculum: how to be on-call, how to declare an incident, how to run incidents as IC, how to communicate, how to write blameless postmortems.
- Create role-specific tracks: (1) All engineers: 30-min video on severity levels, communication expectations, postmortem format, action items (mandatory); (2) On-call responders: 2-hour workshop on playbooks (S12), alert tool (S5), escalation paths (S9), when to call manager, case studies, mobile app walkthrough; (3) Incident commanders: 4-hour IC bootcamp on leadership under pressure, decision-making, communicating with executives, status-page discipline, postmortem facilitation, practiced drills; (4) Communications leads: 2-hour training on templates, update timings, how to talk to customers, regulatory rules.
- Require certification before joining IC on-call roster: written assessment plus live simulated incident (pass/fail).
- Record videos so async teams can learn on their schedule; create runbooks and quick-reference cards (print + digital); pair new on-call engineers with experienced responder for first week.
- Run monthly tabletop exercises on realistic scenarios from the prior 12 months: region loss, ledger corruption, cascading failures.
- Run quarterly game days with intentional failure injection (database failover, status-page outage, alert tool downtime); include all on-call roles.
19. Pilot with volunteer teams (depends on: 5, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 18)
Do not roll out untested to 28 teams. Run the entire process end to end with a small cohort first.
- Recruit 3–4 volunteer teams covering a mix of criticality: one payment-path team, one ledger-adjacent team, one platform team, one low-traffic team.
- Run complete process in pilot: new severity scale (S3), consolidated tooling (S5, S6), roles (S4), escalation (S9), communications (S13, S14), postmortems (S15), action tracking (S16), paid on-call (S11), training (S18), metrics (S17).
- Instrument pilot against S17 metrics; compare results with S2 baseline.
- Hold weekly retrospectives with pilot teams; iterate on written policies, tooling, training based on feedback.
- Fix top issues found before wider rollout; document what changed and why.
- Produce pilot report with before/after numbers (MTTD, MTTR, alert noise, action completion rate) to carry into rollout conversations.
- Set explicit pilot exit criteria: rotation coverage achieved, zero unacknowledged pages over 2 weeks, all postmortems delivered on time, >80% of action items tracked.
20. Phased rollout to all 28 teams (depends on: 16, 17, 19)
Rollout is a staged migration with readiness gates, not an email announcement.
- Sequence the 28 teams into four waves of ~7 teams each, ordered by customer-impact criticality and readiness; space waves three weeks apart.
- Define per-team readiness checklist: services mapped and owned, alerts cleaned to standard (S8), runbooks written, rotation staffed, training complete, manager briefed.
- Hold gate review with program owner before each team joins; move unready teams to next wave with dated remediation plan.
- Assign named champion per wave; run internal communications cadence explaining why using pilot numbers from S19.
- Handle resistance directly: publish the "own-your-code, own-your-pager" rule and paid on-call mechanics before each wave, not after.
- Retire legacy tools, informal escalation lists, and ad-hoc status-page process at end of each wave on published cutover date.
- Harvest feedback formally at each wave and push accepted process changes through change control.
21. SOC 2 control mapping and evidence framework (depends on: 1, 3, 13, 15)
The audit will test incident response. Mapping controls now means evidence accumulates from the first real incident under the new process.
- Map the process to Trust Services Criteria for incident identification, response, evaluation, and communication of security incidents.
- Write control statements in auditor language; name a single owner per control.
- Specify the evidence artifact for each control: incident record, severity classification, paging log, communications log, postmortem, action tracker entry, training record.
- Define evidence retention and storage location (not on laptops, not on chat history that expires); plan for audit access.
- Identify which controls are blocked until certain rollout waves complete; keep a gap register with owners and review fortnightly with steering group.
- Run early walkthrough with an experienced compliance partner or pre-audit readiness team to test control design before the audit window.
22. SOC 2 dry run, gap remediation, and audit support (depends on: 20, 21)
Convert a good process into a provable one, a few months before auditors arrive. Prove the system works at scale.
- Schedule a dry run six weeks before audit window, sampling real incidents from completed waves against each control's evidence requirements.
- Remediate every gap found; prioritize anything risking a qualified opinion.
- Brief every engineer who may be interviewed so they can describe the process as practiced.
- Prepare auditor package: process documentation, sample incident records, training records, on-call schedules, action tracking register, alert quality metrics.
- Designate a single audit liaison and small evidence-request team so requests do not scatter across teams.
- Rehearse IC and Communications Lead roles under interview conditions; auditors probe realism under pressure.
- Ensure all postmortems, incidents, and evidence are retained, searchable, and accessible to auditors for the required audit period.
23. Standing governance, process ownership, and continuous improvement (depends on: 20, 22)
The classic post-audit failure is the process freezing and decaying. Lock in continuous improvement as a permanent structure.
- Establish a standing Incident Management Council chaired by the program owner, meeting monthly with engineering, support, compliance, and product representation.
- Give program owner documented mandate to change standards; require formal change control for any change to severity, roles, communications timings, or compensation.
- Re-validate severity taxonomy quarterly against real declarations; re-baseline metrics annually.
- Feed incident themes into architecture review and hiring so the program improves the system, not just the response.
- Report quarterly to executive team on metric set (MTTD, MTTR, SLA credits, customer-detected %) and top five systemic causes of incidents.
- Fund a resilience roadmap separate from incident response: reducing shared-database blast radius, multi-region failover, deploy safety.
- Refresh training and tabletop program annually and after any SEV1.
- Keep public backlog of improvement ideas (Jira, shared doc) so any engineer can suggest; publish monthly "incident newsletter" to all engineers with wins and learnings.
Previous Proposal 2 (ID: d4b0c90c-c556-49ae-945b-5c59cc4fbd11, Agent: deepseek-flash_refine_2, LLM: deepseek/deepseek-flash):
Estimated Complexity: high
Success Metrics: - Median time to detect falls from 22 minutes to under 8 minutes by month 6, and under 5 minutes for money-moving journeys by month 9.
- Share of customer-impacting incidents first detected by customers falls from 40% to under 10% by month 9.
- Median time to mitigate for SEV-1 falls from 3 h 10 min to under 60 minutes by month 9.
- 100% of SEV-1 and SEV-2 incidents have a named Triage Owner within 5 minutes and a named Incident Commander within 15 minutes, from month 4.
- Zero incidents in which command authority is unclear for more than 15 minutes, measured monthly from month 4.
- Monthly page volume falls from 3,400 to under 600, with a false-positive rate under 15%, by month 5.
- No service exceeds two pages per on-call shift for three consecutive months, measured by month 6.
- All 28 teams have a documented rotation, a service ownership map and at least one trained on-call engineer by month 6.
- Incident Commander roster holds at least 12 certified ICs covering 24x7 with no uncovered week, from month 4.
- 100% of on-call shifts are paid under a published policy from month 2.
- 100% of SEV-0, SEV-1 and SEV-2 postmortems are published internally within 15 business days, from month 5.
- Postmortem action items closed within 60 days rise from 17% to over 90%, with median age under 30 days, by month 6.
- At least one incident review or near-miss report is filed per team per quarter, from month 6.
- Status page first update is published within 30 minutes on at least 95% of SEV-1 incidents, from month 4.
- Zero missed regulatory notification windows on any incident requiring notification.
- The customer-impact ledger is complete for 100% of customer-impacting incidents from month 5, with SLA credits reconciled monthly.
- Annual SLA credits paid fall at least 50% from the $1.3M baseline within 12 months.
- Customer-impacting incidents fall at least 40% year over year from the 31-incident baseline.
- On-call satisfaction is at or above 7 out of 10, with zero voluntary attrition attributed to on-call, measured quarterly from month 6.
- Review cadence is sustained: weekly operational review in at least 90% of weeks, 12 of 12 monthly reliability reviews, 4 of 4 quarterly executive reviews.
- SOC 2 Type II is passed with zero findings related to incident response.
- 100% of new engineers complete the incident-response onboarding module within 30 days of joining.
- 100% of the 180 services have a named owner, a detection contract and at least one symptom-based alert held to the paging contract by month 7.
Steps (21):
1. Charter, mandate and the evidence clock
This step turns the CEO's email into a funded programme with one accountable owner and explicit authority, and it starts the SOC 2 clock on day one.
- Appoint a single accountable process owner — a Director of Incident Management reporting to the CTO, with a dotted line to the COO for customer and SLA matters.
- Publish a one-page charter: scope (every customer-impacting, money-moving, security and data-integrity event across all 28 teams and both regions), decision rights during an active incident, and the power to freeze deploys and override team preferences.
- Fix the funding envelope up front: tooling licences, training and drill time, and on-call compensation, with an indicative annual figure and the expected return in avoided SLA credits.
- **Start the evidence clock immediately.** A SOC 2 Type II report covers a period of operating effectiveness, so evidence cannot be backfilled; an audit in eight months means the operating period begins in the first weeks of this programme.
- Stand up a steering group of CTO, VP Engineering, Head of Support, Head of Compliance and two engineering managers, meeting fortnightly.
- Make incident-process participation a documented performance expectation for every engineering manager, not an optional extra.
- Agree the timeline explicitly: a working process in four months, evidence accumulating from month two, audit-ready by month seven.
2. Baseline evidence and problem statement (depends on: 1)
You cannot demonstrate improvement without a defensible starting point, and you cannot win the argument about noise without numbers.
- Build a 12-month incident register and re-classify all 31 customer-impacting incidents: date, severity, class, detection source, time to detect, time to mitigate, customers affected, services involved and SLA credits paid.
- Run an alert census per tool, per team and per service: total volume, page-to-action ratio, off-hours interruptions per engineer, and the 50 noisiest rules with a named owner.
- Build a silent-failure register: incidents in which no internal alert fired at all. This is the number that explains the 40% customer-detected rate.
- Survey every on-call engineer and manager on burden, fairness, escalation clarity and pay expectations, targeting above 70% response.
- Interview Support, Account Management and Sales about how customers learn of incidents, what they complain about, and what they are promised.
- Reconstruct the two command-ambiguity incidents minute by minute, from alert to mitigation, to find exactly where ownership lapsed.
- Publish the pack internally as the problem statement and retain it as management-review evidence for the audit.
3. Control mapping and evidence architecture (depends on: 1)
Most programmes leave compliance to the end. This one maps controls in the first month, because the mapping determines what the process must capture from day one.
- Map the process to the relevant Trust Services Criteria for incident identification, evaluation, response, containment, recovery and communication of events (CC7.1 to CC7.5, CC2.2, CC2.3, CC4.1, CC3.x).
- Write each control as a plain-language statement with one named owner and its evidence artifact: incident record, severity classification, paging log, communications log, postmortem, action tracker entry, training record.
- Define the **golden incident file**: one auto-assembled, single-click export per incident containing timeline, roles, severity, communications, postmortem, actions and closure evidence.
- Set retention, storage location and immutability so no control depends on a laptop, a private Slack channel or chat history that expires.
- Keep a gap register with owners and dates, reviewed fortnightly by the steering group.
- Run an early design walkthrough with the auditor's readiness team inside the first 90 days, to test the design before building on it.
4. Severity times class taxonomy (depends on: 2)
Severity alone is not enough. Severity decides how loud the response is; class decides which playbook runs and what must never be traded away.
- Define severity by impact in customer and money terms: customers affected, dollars at risk, regions lost, ledger integrity, settlement delay, data loss and SLA exposure.
- Use SEV-0 for security, privacy and regulatory events; SEV-1 for total or material loss of a payment path; SEV-2 for degradation or single-region loss; SEV-3 for limited impact with a workaround; SEV-4 for internal-only issues and near-misses.
- Define response classes: Availability, Performance, Data integrity and Ledger, Security and Privacy, Third-party dependency, Settlement and Financial, and Process failure.
- **Class can raise a response, never lower it.** A SEV-2 data-integrity incident gets SEV-1 posture, because integrity failures are not recoverable by moving faster.
- Specify automatic triggers: loss of one AWS region, ledger write failures, replication lag above threshold, a missed settlement cut-off, payment success rate below threshold for five minutes.
- State who may declare: any engineer, Support agent or account manager. State who may downgrade: the Incident Commander alone.
- Map every level to its SLA credit exposure and to the customer-visible status page state.
- Include worked examples from the last 12 months so each of the 28 teams recognises its own incidents, and re-validate the taxonomy quarterly against real declarations.
5. Roles, command and the never-without-an-owner rule (depends on: 4)
The two incidents where nobody was in charge for over an hour did not fail at declaration. They failed in the gap before it, when an alert had fired and no one owned it.
- Create one-page role cards: Triage Owner, Incident Commander, Deputy IC, Operations Lead, Internal Communications Lead, Customer Communications Lead, Scribe, Subject-Matter Responders and Executive Sponsor for SEV-1 only.
- Introduce the **Triage Owner** rule: from the moment a page is acknowledged, that person owns the incident until an IC takes over or the incident is stood down. There is never an unowned minute between first page and close.
- State the core rule plainly: the IC owns the incident, not the fix, and does not debug. An IC who starts troubleshooting has abandoned command.
- Give the IC explicit decision rights: declare and escalate severity, freeze changes, halt deploys, approve customer messaging, and pull in any engineer in the company.
- Define minimum viable staffing per severity: SEV-1 staffs every role; SEV-2 staffs IC, scribe, comms and responders; SEV-3 staffs an IC and a scribe only.
- Set handover discipline: four-hour maximum IC shifts on SEV-1, a written handover template, and a deputy named within 15 minutes of declaration.
- Set behaviour standards for responders: one incident channel, one bridge, no side channels, and asks phrased with a named owner and a time.
- Publish the role cards on the internal wiki and link them from every paging notification.
6. Declaration, lifecycle and escalation policy (depends on: 5)
This step defines the mechanical path from an alert to a declared incident and back to normal service, and it removes judgment calls from the worst moments.
- Define lifecycle states with entry and exit criteria: Detected, Triaged, Declared, Mitigated, Resolved, Postmortem, Closed — plus a Watch state with a hard 30-minute timer, after which the incident is either declared or stood down.
- **Make declaring free.** A declaration that turns out to be a false alarm is closed as a false declaration, with no blame and no follow-up, and it is tracked as a metric so the cost of caution stays visible.
- Apply the ambiguity rule: if two responders disagree about whether this is an incident, it is an incident.
- Set acknowledgement and declaration targets: page acknowledged within five minutes, triage decision within fifteen, declaration within thirty for anything customer-facing.
- Build automatic escalation ladders with timers at every layer: responder, service owner, team manager, IC on-call, VP Engineering. Escalation never requires a human decision and is never criticised.
- Define the unresponsive-team path: fifteen minutes without acknowledgement escalates to the team manager, thirty minutes to the IC, who may then direct any available engineer.
- Define change freeze and rollback authority during SEV-1 and SEV-2, and the single named condition that lifts the freeze.
- Enforce one incident, one record, with the timeline auto-captured from the channel and bridge rather than written from memory afterwards.
- Add the two-simultaneous-incidents rule: a second IC is designated, and no IC ever runs two incidents at once.
7. Detection strategy: journeys, synthetic signals and customer-report intake (depends on: 4)
Customers detected 40% of incidents first. That number is the reason this step exists, and it is fixed by measuring customer journeys rather than infrastructure.
- Define SLIs and SLOs for the top 20 customer journeys — payment initiation, authorisation, settlement, ledger read and write, API availability, webhook delivery and payout — measured per region.
- Require symptom-based alerting on those SLOs instead of cause-based alerting on infrastructure metrics.
- Add synthetic transaction monitoring from outside AWS in both regions plus a third location, on a one-minute cadence for money-moving paths.
- Add ledger and shared-PostgreSQL specific signals: replication lag, connection saturation, write latency, lock waits, transaction ID exhaustion, checkpoint pressure and bloat.
- Open a customer-report intake: Support and account managers can raise an incident directly, every customer report creates an incident record, and the customer-report path is counted as a detection source in all reporting.
- Apply the detection-gap rule: whenever a customer reports an incident before internal monitoring, a detection-gap ticket is opened automatically and owned by the responsible team.
- Publish a detection contract per service: every one of the 180 services needs an owner, at least one symptom alert, and a documented expected detect time.
- Measure current detect time per journey, set targets, and run a detection drill per team: break something in staging and see whether it pages before a human notices.
8. Three on-call rotations across 28 teams (depends on: 5)
The objection is that engineers will not carry a pager for another team's code. The answer is not to argue with it, but to build three rotations so the objection becomes structurally impossible.
- Run a **Service On-Call** rotation per team, covering only that team's own services. No engineer is ever paged for code their team does not own.
- Run a **Platform Duty** rotation for genuinely shared infrastructure: the shared PostgreSQL cluster, Kubernetes, networking, CI/CD and observability. This is nobody's product code, so it gets its own paid rotation, staffed from platform teams plus volunteers from other teams.
- Run a central **Incident Commander** roster of 12 to 16 certified senior engineers drawn from across all 28 teams, covering 24x7 on one-week shifts with a primary and a secondary.
- State the consequence honestly: 16 of 28 teams have no on-call today. Each must either build a rotation or formally transfer ownership of its services to a team that will, with the transfer dated and recorded.
- Merge small or low-traffic teams into shared rotations so no rotation has fewer than six engineers, the level below which coverage stops being sustainable.
- Cap load in the scheduling tool: no engineer is on-call more than two weeks per quarter, enforced by configuration rather than negotiation.
- Publish a coverage matrix for all 28 teams showing services, owners, rotation size and gaps, reviewed monthly.
- Write on-call participation into engineering job levels and hiring criteria, so the commitment is set before someone joins rather than after.
9. Compensation, rest and the price of opting out (depends on: 8)
Unpaid on-call is the most cited reason for resistance. The policy must be settled before rollout, not negotiated during it.
- Move immediately to paid on-call: a per-shift stipend benchmarked to the New York market, agreed with HR and Finance, and published with an effective date before any team is asked to join a new rotation.
- Pay event-based compensation for out-of-hours callouts, with a minimum call-out block and a 1.5x rate for time actually spent mitigating.
- Provide documented compensatory rest: no engineer works a normal day after a night incident, and the rest day is policy rather than a favour granted by a manager.
- Cap intrusion with a maximum number of off-hours pages per shift, and trigger a mandatory review whenever the cap is exceeded.
- Allow opt-out but **put a price on it**: an engineer may step out of a rotation, and their team must buy coverage from the paid pool at a published internal rate. This turns a cultural argument into a visible budget decision.
- Publish an explicit amnesty: incident records, near-miss reports and false declarations are never used in performance reviews. Only failure to report is a performance issue.
- Protect against gaming: pay attaches to shifts and callouts, never to incident counts.
- Review the policy every six months against real page volumes, attrition and survey results.
10. Paging contract and alert quality (depends on: 2, 7)
3,400 alerts a month at 85% noise is why engineers resent the pager. Fixing that is the price of admission for everything else in this plan.
- Publish the paging contract: every page must be symptom-based, actionable, owned, mapped to a severity and a class, and linked to a runbook. **No runbook, no page**, enforced by a CI check on the alert definition.
- Separate paging from ticketing: only customer-impacting or imminent-impact signals may wake a human; everything else becomes a ticket or a dashboard entry.
- Set a page budget per team and per service: a maximum number of pages per on-call shift. Breaching it auto-opens a remediation ticket with the engineering manager as owner.
- Put new alerts on two-week probation: a new rule runs as a ticket only and becomes a pager only after it has proved actionable, so teams stop being woken by untested rules.
- Define the default action for a noisy alert: fix, tune or delete within ten working days, with deletion treated as a legitimate and celebrated outcome and the deleted-alert count published.
- Correlate and deduplicate at the ingest pipeline so one root cause produces one page instead of forty.
- Require an expiry date on every silencing rule and temporary threshold, so suppression cannot quietly become permanent blindness.
- Run a 90-day noise sprint with a public burn-down of the top 100 noisiest rules, each owned by a named manager, and report page-to-action ratio per team monthly.
11. Incident tooling consolidation (depends on: 5, 10)
Six alerting tools and no single incident record is a structural cause of the 22-minute detection and the three-hour mitigation.
- Select one incident platform for paging, schedules, escalation policies, incident records, postmortem workflow and status-page integration, and choose it within four weeks.
- Consolidate the six alert sources into a single event pipeline feeding that platform, with deduplication, correlation and severity and class mapping applied at ingest.
- Integrate observability, runbooks, chat and voice bridges so responders see dashboards and runbooks inside the incident record and the timeline is captured automatically.
- Implement the golden incident file export defined in S3, so audit evidence is a single click rather than a reconstruction.
- Run a dual-run period alongside the legacy tools with a defined rollback, then switch the legacy tools off on a published date.
- Host the status page outside the production failure domain so it survives a total platform outage, and test that during a game day.
- Make the platform usable from a phone: an Incident Commander must be able to run a SEV-1 from a mobile device at 3am.
12. Communications: internal, customer and regulator (depends on: 4, 5)
Today the status page is written by whoever is around. This step replaces improvisation with a clock, a named owner and a pre-cleared template.
- Set the internal cadence: first update within 15 minutes of declaration, then every 30 minutes for SEV-1 and hourly for SEV-2, whether or not there is progress.
- Never let the status page be how an employee learns of an incident: internal communication leads, external follows.
- Set the customer cadence: status page updated within 30 minutes of a SEV-1 and 60 minutes of a SEV-2, with a no-new-information update still mandatory.
- Define the channel hierarchy: status page for all 2,100 customers, direct email to affected customers on SEV-1, and named account-manager calls for the top 50 accounts.
- Prepare templates per severity and class in advance, pre-approved by Legal and Compliance, each with the next-update time built in.
- Forbid speculation: customer communications never guess at cause, never assign blame, and never commit to a root cause before the postmortem.
- Define the closing communication: resolution notice, how SLA credits will be handled, and the committed date for the written report.
- Build a **regulator clock matrix**: for each event type, which regulator, which window, who signs off, and the shortest clock that drives the first action. Cover NYDFS Part 500, money-transmitter and banking notification, security breach notification, card-network rules, and public-company disclosure where applicable.
- Route every regulatory notification through Compliance, never Engineering, and pre-clear the templates.
- Publish a customer-facing root-cause report for SEV-1 incidents, especially for regulated and top-tier accounts.
- Assign a named Customer Communications Lead plus a trained deputy on every SEV-1.
13. Customer-impact ledger and SLA credit automation (depends on: 12)
The $1.3M in credits is a symptom of having no single record of customer impact. This step creates one, and makes it do four jobs at once.
- Maintain one durable customer-impact ledger per incident: which customers, which journeys, from when to when, and the estimated credit.
- Use that single record for customer communications, SLA credit computation, regulatory reporting and the annual review, so nobody reconciles two versions of the same outage.
- Automate the credit path end to end — computation, approval, customer notification and finance treatment — so credits stop being a manual scramble three weeks later.
- Reconcile accrued against paid credits monthly and report the result in the executive review.
- Track credit avoidance against programme cost, so the funding case is a number rather than an argument.
14. Postmortem policy with three artifact levels (depends on: 5)
Postmortems currently happen for some incidents, in various formats. This step makes them mandatory where they matter and light where they do not.
- Make postmortems mandatory for every SEV-0, SEV-1 and SEV-2, every SEV-3 with customer impact or a repeat pattern, every near-miss the IC flags, and every incident in which the process itself failed.
- Use three levels so the ritual matches the weight: a lightweight async incident review for SEV-3 and SEV-4, a standard facilitated postmortem for SEV-2, and a full review with an executive sponsor for SEV-1.
- Set deadlines: draft within five business days, blameless review within ten, internal publication within fifteen.
- Adopt one template: impact, timeline, detection, response, contributing factors across tooling, process, organisation and human factors, what went well, what went badly, and actions.
- Train a pool of blameless facilitators and require a trained facilitator for every SEV-1 review.
- Prohibit counterfactual and blame language in the template, and specifically ban the phrase human error as a root cause — the question is always what made the error possible.
- Publish all postmortems internally by default, with security review only for genuinely sensitive material, and produce a customer-facing root-cause variant for SEV-1.
- Treat postmortems as the learning product of the process, not as a compliance artifact.
15. Action items: capped, verifiable, with a repeat-incident rule (depends on: 14)
Eleven of 64 action items closed is not a tracking problem. It is a generation problem: the process produces more actions than the organisation can absorb.
- **Cap each postmortem at three action items.** Anything beyond three goes into a ranked reliability backlog, not into the postmortem.
- Require each action to have a named human owner, a due date, and a definition of done that is an artifact: a merged change, a new alert that fired in a drill, or a test that fails without the fix.
- Prohibit self-reported closure. Closure requires the artifact, signed off by the process owner or the Incident Commander.
- Keep one reliability backlog in the engineering tracker with a mandatory label and a link back to the originating incident.
- Reserve a fixed percentage of every team's sprint capacity for reliability work, with unspent capacity visible to the vice presidents.
- Apply the repeat-incident rule: a second incident with the same contributing factor triggers a design review owned by an architect, not another action item.
- Run a weekly ageing review of open actions and escalate anything more than 30 days overdue to the VP of Engineering.
- Target more than 90% of actions closed within 60 days and a median age under 30 days, reported monthly by team.
16. Training, certification and drills (depends on: 5, 6, 12, 14)
A process that exists only on a wiki page fails on the first real page. Skills have to be built and tested before they are needed.
- Build a short practical curriculum: how to be on call, how to triage and declare, how to run an incident as IC, how to communicate, and how to write a postmortem.
- Require certification before joining the IC roster: a written assessment plus a live simulated incident.
- Certify at least two ICs per team group, so the central roster has depth across all 28 teams and no holiday week is uncovered.
- Run monthly tabletops drawn from the last 12 months of real incidents, including region loss, ledger write failure and a missed settlement cut-off.
- Run quarterly game days with deliberately injected failure, including shared-PostgreSQL failover, status-page outage and alerting-pipeline outage.
- **Drill the process's own failure modes**, not just technical ones: IC unreachable, comms lead on PTO, two simultaneous SEV-1s, a paging storm, and a false alarm that burns an hour.
- Audit the incident process for single points of failure: who is the only person who can do each critical task, and what happens in their holiday week.
- Add a mandatory onboarding module for every engineer joining or transferring in, with completion records kept in an audit-ready form.
- Track training completion by team and publish it in the monthly reliability review.
17. Metrics and review cadence (depends on: 2, 4, 15)
Establish what good looks like, and measure it in a way that makes people report more incidents rather than fewer.
- Define outcome metrics: time to detect by source, time to mitigate by severity and class, percentage of incidents detected by customers, incidents by severity and class, SLA credits paid, and customer-impact minutes.
- Define process metrics, because the process itself has SLOs: IC assigned within five minutes, page acknowledgement rate, first-update timeliness, update-cadence adherence, postmortem timeliness, action closure rate, action median age and IC roster coverage.
- Define health metrics: alert volume and noise ratio per team, off-hours pages per engineer, on-call satisfaction, training completion, and percentage of services with a detection contract.
- **Never publish incident count as a team metric.** It rewards hiding incidents. Publish reporting metrics instead — near-misses filed, detection gaps closed, false declarations made — alongside the outcome metrics.
- Publish live dashboards visible to every engineer, refreshed daily, with each metric baselined against the S2 evidence pack and given a 90-day and a 12-month target.
- Run a 30-minute weekly operational review incident by incident, a monthly reliability review of trends and actions, and a quarterly executive review with the CEO's office covering customer impact, credits and the top five systemic causes.
- Hold a quarterly review of the process itself: what in the process wasted time, what confused responders, and what should be deleted.
- Require every review to end with decisions and named owners, never with numbers alone.
18. Pilot with three to four teams, using real incidents (depends on: 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17)
Do not roll out to 28 teams untested, and do not rely on drills alone when real incidents are available and better.
- Recruit three to four volunteer teams spanning criticality: a payment-path team, a ledger-adjacent team, the platform team that owns the shared PostgreSQL cluster, and a low-traffic team.
- Run the entire process end to end in the pilot: the new severity and class scale, roles, escalation, communications, postmortems, action tracking and paid on-call.
- Treat real incidents during the pilot as the primary training material. Hold a retro within 48 hours of each pilot incident, run by the process owner, while the friction is fresh.
- Instrument the pilot against the S2 baseline and publish before-and-after numbers.
- Hold weekly retrospectives with the pilot teams and change the written policies, the tooling and the training in response, documenting what changed and why.
- Set explicit exit criteria: rotation coverage achieved, zero unacknowledged pages over a defined period, postmortems delivered on time, and actions tracked to closure.
- Produce a pilot report that every later rollout conversation starts from.
19. Phased rollout sequenced by cost of failure (depends on: 15, 18)
Rollout is a staged migration with readiness gates, not an email announcement. The sequencing matters more than the schedule.
- Sequence the 28 teams into four waves of roughly seven, ordered by **incident density and customer-journey ownership**, not by size or alphabet: the teams whose incidents cost the most SLA credits go first, because their improvement is the visible proof.
- Leave three weeks between waves and define a per-team readiness checklist: services mapped and owned, alerts cleaned to the paging contract, runbooks written, rotation staffed to at least six engineers, training complete, manager briefed, and compensation in effect.
- Hold a gate review with the process owner before each team joins, and move an unready team to the next wave with a dated remediation plan rather than granting an exception.
- Give each wave a named champion and run an internal communications cadence that explains the why using the pilot's numbers.
- Handle resistance directly and early: publish the own-your-code-own-your-pager rule and the paid on-call mechanics before each wave, never after.
- Hold a retro within 48 hours of each wave's first incident under the new process, and feed accepted changes back through change control.
- Retire legacy tools, escalation lists and the informal status-page process at the end of each wave on a published cutover date, so two processes never run in parallel for long.
- Sequence the waves so no team faces both a rollout and the audit dry run in the same week.
20. Audit dry run and evidence review (depends on: 3, 19)
This step converts a good process into a provable one, a few months before the auditors arrive.
- Run an internal dry run six weeks before the audit window, sampling real incidents from the pilot and the early waves against each control's evidence requirements.
- Remediate every gap found, prioritising anything that risks a qualified opinion, and re-test the remediated control against the same sample.
- Brief every engineer who may be interviewed so they can describe the process as they actually practise it, not as it is written.
- Prepare the auditor package: process documentation, sample incident records, paging logs, communication logs, training records, on-call schedules and the action tracking register.
- Designate one audit liaison and a small evidence-request team, so requests do not land on 28 teams simultaneously.
- Rehearse the walkthrough with an incident commander and a communications lead, because auditors probe realism under pressure, not documents.
- Keep the audit liaison and the process owner as the same person where possible, so accountability for the control is also accountability for the evidence.
21. Standing governance, resilience roadmap and continuous improvement (depends on: 19, 20)
The classic post-audit failure is that the process freezes and then decays. This step keeps it alive and separates incident response from the engineering that prevents incidents.
- Establish a standing Incident Management Council chaired by the process owner, meeting monthly with engineering, support, compliance and product representation.
- Give the process owner a documented mandate to change standards, and require formal change control for any change to severity, classes, roles, communication timings or compensation.
- Re-validate the severity and class taxonomy quarterly against real declarations, and re-baseline the whole metric set annually.
- Fund a **resilience roadmap separate from incident response**: reduce the shared PostgreSQL blast radius, improve multi-region failover readiness, and tighten deploy safety. Better incident management cannot save a single shared ledger during a corruption event.
- Report quarterly to the executive team on the metric set and the top five systemic causes of incidents, and feed those themes into architecture review, release engineering and hiring.
- Keep a public backlog of process improvement ideas so any engineer can propose one, and change the process from it every quarter.
- Share wins publicly each quarter — faster detection, fewer customer-detected incidents, credits avoided — to sustain momentum after the audit.
- Refresh the training and tabletop programme annually and immediately after any SEV-1.
Previous Proposal 3 (ID: 985bec06-6c51-44c4-9fe4-d27c2f645d5c, Agent: qwen3.8-flash_refine_3, LLM: alibaba/qwen3.8-flash):
Estimated Complexity: high
Success Metrics: - Median Time to Detect (MTTD) < 10 minutes.
- Median Time to Mitigate (MTTR) < 45 minutes for SEV-1.
- Customer-detected incidents < 5% of total incidents.
- Monthly alert volume < 400 actionable alerts (90% reduction in noise).
- SLA credit payouts < $100k annually.
- Postmortem action item completion rate > 90% within 30 days.
- 100% of SEV-1 incidents have a designated IC and Scribe.
- On-call engineer satisfaction score > 4.0/5.0.
- Zero critical findings in SOC 2 Type II audit regarding incident response.
Steps (20):
1. Program Charter & Executive Mandate
Secure explicit authority to mandate the incident management program, transforming it from an engineering initiative into a business-critical operational standard.
* Appoint a Director of Incident Management (DIM) reporting to the CTO, with a dotted line to the Head of Compliance for audit readiness.
* Publish a charter defining scope: all customer-impacting incidents, payment path failures, and data integrity issues.
* Establish a steering committee (CTO, VP Eng, Head of Support, Legal) to resolve cross-team conflicts and approve tooling budgets.
* Define the 'Own Your Code, Own Your Pager' principle as a non-negotiable engineering value, directly addressing the 28-team resistance.
2. Baseline Audit & Evidence Framework (depends on: 1)
Create a defensible starting point for metrics and a structured evidence repository for the SOC 2 Type II audit.
* Consolidate the last 12 months of incident data into a single register, flagging the 31 customer-impacting incidents.
* Map current alert sources to specific services to quantify the 3,400 monthly alerts and identify the top 50 noisiest rules.
* Conduct a survey of on-call engineers to measure current sentiment and identify specific pain points in the unpaid rotations.
* Designate a secure, immutable storage location for all future incident artifacts (timelines, logs, postmortems) to satisfy audit requirements.
3. Severity Taxonomy & Trigger Matrix (depends on: 1)
Define unambiguous severity levels tied to specific triggers, communication SLAs, and executive notifications to eliminate classification debates.
* **SEV-1 (Critical):** System-wide outage, data loss risk, or >5% revenue-impacting transactions. Triggers immediate IC, VP Eng, and legal notification.
* **SEV-2 (Major):** Significant degradation or region-wide latency. Triggers IC and Team Lead.
* **SEV-3 (Minor):** Localized issue with workaround. Triggers on-call engineer.
* **SEV-4 (Internal):** Low-priority noise. Triggers ticket only.
* Map each severity to specific SLA credit exposures to align technical response with financial risk.
4. Incident Roles & Command Structure (depends on: 3)
Standardize the Incident Command System (ICS) to prevent the 'ambiguity of ownership' that caused previous hour-long delays.
* **Incident Commander (IC):** Single point of decision authority; does not debug. Required for SEV-1/2.
* **Scribe:** Maintains the real-time timeline log; mandatory for SEV-1.
* **Comms Lead:** Owns status page and internal broadcasts; shields IC from external noise.
* **SMEs:** Technical responders focused solely on diagnosis/mitigation under IC direction.
* Create role cards with clear 'Do' and 'Don't' responsibilities for quick reference during chaos.
5. On-Call Architecture & Compensation Policy (depends on: 4)
Establish a sustainable, paid 24x7 on-call model that respects ownership boundaries and compensates for disruption.
* Adopt a federated model: each of the 28 teams maintains its own SME rotation for owned services.
* Create a central IC rotation staffed by senior engineers across teams to ensure command coverage.
* Implement a paid policy: base stipend per shift, enhanced pay for call-outs, and mandatory recovery time after SEV-1 night shifts.
* Define the 'Unowned Service' rule: services without a dedicated team rotation must be transferred to an owning team or decommissioned.
6. Tooling Consolidation & Integration (depends on: 2, 5)
Replace the six fragmented alerting tools with a single Incident Management Platform to ensure a unified source of truth.
* Select a platform (e.g., PagerDuty + Statuspage + Slack/Teams) capable of handling 180 services.
* Integrate Kubernetes and PostgreSQL monitoring to feed alerts into the central platform.
* Automate the creation of incident channels and timelines upon alert acknowledgment.
* Ensure the status page is decoupled from the production environment to remain available during outages.
7. Alert Quality & Noise Reduction Program (depends on: 6)
Implement strict alert hygiene rules to reduce the 3,400 monthly alerts and restore trust in the pager.
* **Rule:** All paging alerts must be symptom-based (customer impact) and have a linked runbook.
* **Rule:** Implement deduplication and rate-limiting at the ingestion layer.
* **Sprint:** Conduct a 90-day noise sprint to fix, tune, or delete the top 100 noisiest rules.
* **Metric:** Enforce a 'Page-to-Action' ratio target; if an alert does not result in human action, it must be demoted to ticket-only.
8. Detection Strategy & SLOs (depends on: 3, 6)
Shift from infrastructure metrics to Service Level Objectives (SLOs) to detect customer-impacting issues faster than clients can.
* Define SLIs for critical payment paths: Transaction Success Rate, API Latency, Settlement Lag.
* Implement synthetic transaction monitoring from external vantage points in both AWS regions.
* Establish error budgets; breach of error budget triggers automatic review and potential feature freezes.
* Integrate customer-reported signals (support tickets, social media mentions) into the detection workflow as low-priority alerts.
9. Communication Protocols & Templates (depends on: 4, 6)
Define who talks, when, and what they say to stop ad-hoc status updates and protect brand reputation.
* **Status Page:** SEV-1 initial update within 15 mins; subsequent updates every 30 mins. Use legal-approved templates.
* **Internal:** IC broadcasts to #exec-leadership for SEV-1 every hour.
* **Regulatory:** Define triggers for NY financial regulator notifications; assign Legal/Compliance responsibility for filing.
* **Client Success:** Provide Account Managers with a 'Known Issues' playbook and real-time access to status page APIs for high-value accounts.
10. Postmortem Framework (Blameless) (depends on: 9)
Standardize the learning process to ensure consistent quality and adherence to the 'Blameless' culture.
* **Mandatory:** All SEV-1 and SEV-2 incidents require a postmortem within 5 business days.
* **Format:** Fixed template including Impact, Timeline, Root Cause, Contributing Factors, and Action Items.
* **Review:** Facilitated by a neutral party (not the IC) to enforce blameless language and prevent finger-pointing.
* **Publication:** All postmortems published internally on the Wiki with full searchability.
11. Action Item Tracking & Governance (depends on: 10)
Close the loop on incident resolution by enforcing rigorous tracking of corrective actions to fix the 17% completion rate.
* Automatically create Jira tickets for every action item identified in the postmortem.
* **Enforcement:** SEV-1 incident cannot be marked 'Closed' until high-priority actions are completed or deferred with VP approval.
* **Review:** Weekly review of overdue actions in the Engineering Leadership standup.
* **Metric:** Track 'Mean Time to Remediation' for action items as a key health indicator.
12. SOC 2 Control Mapping (depends on: 2, 10)
Proactively map the new incident processes to SOC 2 Trust Services Criteria to ensure audit readiness.
* Map S4 (Roles), S9 (Comms), and S10 (Postmortems) to Security and Availability criteria.
* Define 'Evidence of Operation' for each control (e.g., automated timeline logs, signed-off postmortems).
* Identify gaps between current state and audit requirements; assign remediation tasks to the DIM.
* Establish a quarterly internal compliance review to test control effectiveness before the Type II audit.
13. Training & Certification Curriculum (depends on: 4, 6)
Equip all engineers with the skills to operate within the new framework, reducing anxiety and improving response quality.
* **Universal Training:** 1-hour module on severity levels and tools for all engineers.
* **IC Certification:** Mandatory workshop and simulation for engineers joining the central IC rotation.
* **Runbook Review:** Each team must update and validate runbooks for their top 3 critical alerts.
* **Onboarding:** Include incident response basics in the engineering onboarding checklist.
14. Pilot Implementation (Wave 1) (depends on: 7, 8, 9, 13)
Deploy the new process to a controlled subset of high-traffic teams to validate assumptions before broad rollout.
* Select 3 teams: Payments Core, Ledger/API, and one Infrastructure team.
* Run the full cycle for 6 weeks: Alerts, IC handover, Comms, Postmortems.
* Collect feedback daily and adjust playbooks and tooling configurations based on real-world friction.
* Measure success using the baseline metrics from S2 (MTTD, Noise, Satisfaction).
15. Full Rollout Strategy (Waves 2-4) (depends on: 14)
Gradually migrate the remaining 25 teams using lessons learned from the pilot to manage change resistance.
* **Wave 2 (Month 3):** Deploy to 8 remaining critical customer-facing teams.
* **Wave 3 (Month 4):** Deploy to internal platform and data teams.
* **Wave 4 (Month 5):** Deploy to remaining low-traffic teams and legacy services.
* Maintain a 'Help Desk' for new on-call engineers to troubleshoot process questions during their first shift.
16. Simulations & Game Days (depends on: 15)
Test the resilience of the process and the tools under controlled failure conditions.
* **Tabletop Exercises:** Quarterly scenario walks-throughs for leadership and ICs (e.g., 'What if the primary region fails?').
* **Chaos Engineering:** Inject failures in non-production or canary environments to test alert accuracy and runbook validity.
* **Communication Drills:** Simulate SEV-1 to test the speed of status page updates and internal notification paths.
* Document findings in postmortems and create action items for identified weaknesses.
17. Metrics Dashboard & Executive Review (depends on: 2, 15)
Establish a continuous feedback loop to monitor the health of the incident management system.
* Build a real-time dashboard tracking MTTD, MTTR, Alert Volume, and SLA Credits.
* **Weekly:** Operational review of new incidents and action items with the DIM and Team Leads.
* **Monthly:** Executive review with CTO/VP Eng focusing on trends, noise ratios, and on-call fatigue.
* Publish 'Reliability Reports' to all engineers to celebrate wins and share learnings.
18. Culture & Change Management (depends on: 5, 15)
Address the cultural shift required to move from 'hero mode' to 'process mode' and mitigate pushback.
* Highlight success stories where the new process reduced toil or prevented customer churn.
* Enforce the 'Blameless' mantra by publicly correcting any leadership communication that implies blame.
* Recognize and reward effective ICs and engineers who improve runbooks or alert quality.
* Address 'Pager Fatigue' proactively by monitoring call-out volume and rotating engineers off if thresholds are breached.
19. SOC 2 Dry Run & Evidence Prep (depends on: 12, 17)
Conduct a mock audit six months out to identify gaps in evidence retention or process execution.
* Sample 10 recent incidents and verify all required artifacts (timeline, logs, postmortem, actions) are present and accessible.
* Interview on-call engineers to ensure they can describe the process and their roles without hesitation.
* Remediate any 'Control Failures' identified during the dry run.
* Prepare the 'Audit Readiness' package for the external auditors.
20. Continuous Improvement Loop (depends on: 17, 19)
Institutionalize the evolution of the incident process to prevent stagnation.
* Review the Severity Taxonomy quarterly to ensure it reflects current system behavior and customer expectations.
* Update Tooling integrations annually based on vendor roadmaps and new AWS/K8s features.
* Revise Compensation Policy annually based on market data and internal fairness reviews.
* Set annual 'Reliability Goals' for the engineering organization to drive long-term focus.
Please, considering the previous proposals as ideas that could be considered, focus on the main objective and generate an IMPROVED proposal or a completely DIFFERENT perspective if you deem it appropriate. Only if you consider any of them is amazing and impossible to improve, answer with the same proposal.
The plan has these parts:
- "steps": the list of steps, each with "step_id" (a short label unique within this proposal, such as S1, S2, S3), "title" (a short, descriptive title), "description" (what the step does and how, in Markdown: a short opening paragraph and then, if there are several concrete things to say, a bullet list) and "dependencies" (the step_ids of the other steps of THIS proposal that must be completed first; empty if none).
- "estimated_complexity": "low", "medium" or "high".
- "success_metrics": clear and measurable success metrics, one per line as a bullet list.
Think deeply about the problem, considering all aspects and implications. Develop a comprehensive understanding of what needs to be done to achieve the objective, then produce the plan.
{"steps": [{"step_id": "S1", "title": "Executive Mandate & SOC 2 Evidence Clock", "description": "Establish clear ownership and funding to transform the CEO's concern into a governed program with immediate compliance implications.", "dependencies": []}, {"step_id": "S2", "title": "Baseline Measurement & Problem Statement", "description": "Quantify the current state to build a defensible problem statement and establish the data foundation for improvement tracking.", "dependencies": ["S1"]}, {"step_id": "S3", "title": "Severity & Class Taxonomy", "description": "Define unambiguous incident levels based on business impact to trigger specific response protocols.", "dependencies": ["S2"]}, {"step_id": "S4", "title": "Incident Command Roles & Structure", "description": "Standardize roles and decision rights to eliminate command ambiguity during active incidents.", "dependencies": ["S3"]}, {"step_id": "S5", "title": "Tooling Consolidation & Integration", "description": "Deploy a unified incident management platform to replace fragmented alerting tools and create a single source of truth.", "dependencies": ["S1"]}, {"step_id": "S6", "title": "Detection Strategy & SLO Implementation", "description": "Shift monitoring focus to customer-impacting symptoms to detect issues before customers do.", "dependencies": ["S3"]}, {"step_id": "S7", "title": "Alert Quality Standards & Noise Reduction", "description": "Implement strict alerting rules to reduce noise and restore on-call trust.", "dependencies": ["S5", "S6"]}, {"step_id": "S8", "title": "On-Call Architecture & Coverage Model", "description": "Design a federated, team-owned on-call structure supported by central command coverage.", "dependencies": ["S4"]}, {"step_id": "S9", "title": "On-Call Compensation & Wellbeing Policy", "description": "Introduce paid, sustainable on-call rotations to remove resistance and mitigate burnout.", "dependencies": ["S8"]}, {"step_id": "S10", "title": "Playbooks & Communication Templates", "description": "Create standardized playbooks and templates for each severity level to reduce cognitive load.", "dependencies": ["S3", "S4"]}, {"step_id": "S11", "title": "Escalation & Incident Lifecycle Policy", "description": "Define the mechanical progression from alert to resolution, including automatic escalation paths.", "dependencies": ["S4", "S5"]}, {"step_id": "S12", "title": "Communication Workflows (Internal & External)", "description": "Establish timed protocols for internal, customer, and regulatory communications.", "dependencies": ["S4", "S10"]}, {"step_id": "S13", "title": "Status Page & Customer Notification Automation", "description": "Integrate incident management with public status pages to ensure timely and accurate updates.", "dependencies": ["S5", "S12"]}, {"step_id": "S14", "title": "Postmortem Policy & Blameless Framework", "description": "Mandate and standardize postmortems to focus on system improvement rather than individual error.", "dependencies": ["S3"]}, {"step_id": "S15", "title": "Action Item Tracking & Governance", "description": "Ensure postmortem action items are tracked, owned, and completed to prevent recurrence.", "dependencies": ["S14"]}, {"step_id": "S16", "title": "Metrics, Dashboards & Review Cadence", "description": "Define and monitor key performance indicators to validate the effectiveness of the incident process.", "dependencies": ["S2", "S3"]}, {"step_id": "S17", "title": "Training, Certification & Drills", "description": "Build practical skills and readiness for incident response across all roles.", "dependencies": ["S4", "S10", "S12"]}, {"step_id": "S18", "title": "Pilot Program with Volunteer Teams", "description": "Test the end-to-end process on a small scale to refine procedures before full rollout.", "dependencies": ["S5", "S7", "S9", "S11", "S13", "S15", "S17"]}, {"step_id": "S19", "title": "Phased Rollout to All Teams", "description": "Deploy the process to all 28 teams in manageable waves to minimize disruption and manage resistance.", "dependencies": ["S18"]}, {"step_id": "S20", "title": "SOC 2 Control Mapping & Evidence Framework", "description": "Align the incident management process with SOC 2 requirements to ensure audit readiness.", "dependencies": ["S1", "S3"]}, {"step_id": "S21", "title": "SOC 2 Dry Run & Audit Support", "description": "Verify evidence sufficiency and prepare the team for the external audit.", "dependencies": ["S19", "S20"]}, {"step_id": "S22", "title": "Governance & Continuous Improvement", "description": "Establish standing governance to sustain the process and drive ongoing improvements.", "dependencies": ["S19", "S21"]}], "estimated_complexity": "high", "success_metrics": "- Median Time to Detect (MTTD) < 5 minutes.\n- Median Time to Mitigate (MTTR) < 60 minutes for SEV1.\n- >90% of incidents detected internally (vs. 40% previously).\n- SLA credits reduced to <$100K annually.\n- Monthly alert volume < 600 with >95% signal-to-noise ratio.\n- Postmortem action item completion rate > 80%.\n- Zero incidents with command ambiguity > 15 minutes.\n- 100% of SEV1/2 postmortems published within 15 days.\n- Status page updates within 30 minutes for 95% of SEV1.\n- 100% of teams integrated into on-call rotations.\n- On-call satisfaction score > 7/10.\n- 40+ certified Incident Commanders available 24x7.\n- SOC 2 Type II audit passed with zero findings."}Votes
The calls of the vote
The votes
| Voter | Model | Voted for | Justification | Usage | Raw |
|---|---|---|---|---|---|
| claudeHaiku4.5_voter_1 | claudeHaiku4.5 · anthropic/claude-haiku-4-5 |
Proposal A · claudeHaiku4.5_refine_1 | Proposal 1 delivers the most actionable and detailed implementation path, with 22 concrete, richly specified steps covering all requirements. It combines aggressive but realistic timelines (month 5 complete rollout, month 8 audit-ready), unambiguous handling of the on-call resistance problem via three explicit rotations with no exceptions, and early control mapping (step 3) that positions for SOC 2 success. While Proposal 2 is nearly equivalent and slightly stronger on resilience separation and cultural sustainability, Proposal 1's clarity, specificity, and urgency-matched pacing directly serve the objectives of a company facing an immediate audit and critical operational failures. |
21.3k in · 4.1k out · 1 min 2 s | show[SYSTEM]
You are an expert and objective evaluator of project plan proposals.
Your task is to select the BEST proposal based on criteria of completeness, clarity, and alignment with the main objective.
Use your internal reasoning processes to thoroughly analyze each proposal, considering all aspects and implications.
Take as much time and space as you need to evaluate each proposal in depth before making your decision.
Every text field you write will be read by a busy person, so write for them: short sentences, one idea each, in short paragraphs. Text fields accept Markdown: separate paragraphs with a blank line, use a bullet list when you enumerate things and plain paragraphs when you explain or argue, and bold at most one key phrase per paragraph. A text of more than three sentences must be split into paragraphs; never deliver a single unbroken block of text.
[HUMAN]
Main Objective: "A B2B payments platform in New York (260 engineers in 28 teams, 2,100 customers, $4B processed a month, a 99.95% SLA with service credits) runs about 180 services on Kubernetes across two AWS regions, with a shared PostgreSQL cluster for the ledger. In the last 12 months: 31 customer-impacting incidents, median time to detect 22 minutes (customers detected 40% of them first), median time to mitigate 3 h 10 min, $1.3M paid in SLA credits, two incidents in which nobody was sure who was in charge for over an hour, and a CEO email about "outages we hear about from clients". On-call exists in 12 of the 28 teams, unpaid, fed by alerts from six different tools — 3,400 alerts a month, 85% of them noise; there is no severity scale, status-page updates are written by whoever is around, and postmortems happen for some incidents, in various formats, with action items that are rarely tracked (11 of 64 closed). A SOC 2 Type II audit in eight months will test incident response. Engineers push back against "carrying a pager for other teams' code".
Define the incident management process: severity levels and what each one triggers; roles (incident commander, communications, scribe, subject-matter responders) and how they are staffed 24x7 across 28 teams; the on-call structure, rotations, compensation and the rules for alert quality; detection and escalation paths; internal and customer communications (status page, account managers, regulators when required) with their timings; postmortems (when mandatory, format, blameless review, ownership and tracking of actions); the metrics and reviews that show whether it works; and how it is introduced across the teams without waiting for the audit."
Proposals to Evaluate:
--- PROPOSAL 1 ---
Proposal ID: 80c18452-b9c3-4820-9dec-cf8cb9a8c7b0
Content:
Estimated Complexity: high
Success Metrics: - Median Time to Detect from 22 minutes to <5 minutes for internal detection and <10 minutes for customer journeys by month 6; <2% of customer-impacting incidents first detected by customers (vs. 40% currently) by month 9.
- Median Time to Mitigate for SEV-1 from 3h 10min to <45 minutes by month 9; SEV-2 to <90 minutes by month 9.
- Monthly alert volume from 3,400 to <400 with >95% signal-to-noise ratio by month 5; all alerts satisfy paging contract (symptom-based, owned, runbook-linked, severity-mapped).
- Zero incidents with command ambiguity lasting >15 minutes from month 2 onwards; 100% of SEV-1 and SEV-2 incidents have named IC logged in timeline within 5 minutes; 100% of Triage Owners assigned within acknowledgement.
- 100% of the 28 teams have a documented Service On-Call rotation or a signed service-ownership transfer by month 4; Platform Duty rotation and IC roster both staffed and live by month 3.
- On-call compensation policy published and effective by month 1; on-call satisfaction score ≥7/10 by month 6; zero on-call-attributed voluntary attrition by month 6.
- SLA credits paid from $1.3M annually to <$100K by month 12; credit avoidance (prevented credits) tracked and reported monthly.
- 100% of mandatory postmortems (SEV-0, SEV-1, SEV-2, and repeat incidents) published internally within 15 business days by month 4.
- Postmortem action item completion rate from 17% (11 of 64) to >90% within 60 days by month 6; median action age <30 days; zero repeat incidents caused by the same contributing factor without a design review.
- 100% of the 180 services have a named owner, a detection contract, and at least one symptom-based alert by month 6.
- Status-page first update published within 30 minutes for ≥95% of SEV-1 incidents by month 4.
- 100% of new engineers complete incident-response onboarding within 30 days of joining; IC certification includes written exam and live simulation; ≥2 certified ICs per team group; zero uncovered hours in 24x7 IC roster.
- Weekly operational review held in ≥90% of weeks; 12 of 12 monthly reliability reviews; 4 of 4 quarterly executive reviews; all reviews end with documented decisions and owners.
- SOC 2 Type II audit passes all incident-response controls (CC7.1–7.5, CC2.2–2.3, CC4.1, CC3.x) with zero findings by month 8.
- All pilot and rollout incidents captured with complete golden incident files (timeline, roles, communications, postmortem, actions, closure evidence) by month 3 onwards; audit dry run identifies zero critical gaps by month 7.
- All 28 teams transitioned to new process by month 5; all legacy alert tools decommissioned; single source of truth for incidents established and sustained.
Steps (22):
1. Executive charter, governance structure, and evidence clock
Turn the CEO email into a funded, authorized program with clear ownership and documented evidence collection for SOC 2, starting today.
- Appoint a Director of Incident Management reporting to CTO, with dotted line to COO (customer impact) and Head of Compliance (audit readiness).
- Publish a one-page charter: scope (all customer-impacting, payment-path, data-integrity incidents across 28 teams and 2 regions), decision rights (IC may freeze changes, override team preferences during incidents), and authority to mandate process participation.
- Secure annual budget for tooling, training, on-call compensation ($500–800K estimated), and resilience work. Connect funding to avoided SLA credits ($1.3M baseline).
- Establish standing Incident Management Steering Group: CTO, VP Engineering, Head of Support, Head of Compliance, one engineering manager per region. Meet monthly.
- **Start the SOC 2 evidence clock on day 1.** An audit in eight months means operating-period evidence begins now; design the process to capture evidence continuously, not retroactively.
- Make incident-response participation a documented performance expectation for all engineering managers and team leads.
- Publish timeline: working process in month 2, all 28 teams in month 5, audit-ready in month 7.
2. Baseline measurement, incident register, and evidence pack (depends on: 1)
Establish defensible baseline metrics and identify structural gaps that explain the 40% customer-detected rate and 22-minute detection time.
- Build a 12-month incident register with all 31 customer-impacting incidents: date, detection source, detection time, mitigation time, customer count, services involved, SLA credits paid, root cause class.
- Audit the current alert estate: total volume per tool, volume per team, volume per service, page-to-action ratio, top 50 noisiest rules, off-hours interruptions per engineer.
- Construct a **silent-failure register**: incidents with no internal alert fired at all. This explains the 40% customer-detected rate.
- Reconstruct the two command-ambiguity incidents minute by minute: exactly when did ownership become unclear, how long, what was the decision bottleneck.
- Survey on-call engineers (target >70% response): burden, fairness, pay expectations, escalation clarity, willingness to stay.
- Interview Support and Account Management: how do customers discover incidents, what do they complain about, how do they contact you.
- Publish the problem statement internally; retain all artifacts for SOC 2 audit evidence. This is the baseline against which all improvements are measured.
3. Control mapping and evidence architecture (depends on: 1, 2)
Design the process to generate SOC 2-compliant evidence automatically, from the first real incident, so the audit clock ticks in your favour.
- Map the new process to Trust Services Criteria CC7.1–7.5 (incident identification, response, evaluation, containment, communication), CC2.2–2.3 (authorization), CC4.1 (change management), CC3.x (information availability).
- For each control, write a one-paragraph plain-language statement, name a single owner, and specify the evidence artifact (incident record, timeline, communications log, postmortem, action tracker, training record).
- Define the **golden incident file**: one single-click export per incident containing severity, timeline, roles assigned, decisions made, communications sent, postmortem, and action items. This is the audit unit.
- Specify data retention, immutability, access control, and storage location (not laptops, not chat history that expires). Ensure evidence is searchable and organized by incident date.
- Keep a gap register with owners and dates; review fortnightly in the steering group. Identify which controls are blocked by incomplete rollout and when they unblock.
- Run an early design walkthrough with an experienced SOC 2 readiness partner inside month 1 to stress-test control design before building on it.
4. Severity and response class taxonomy (depends on: 2)
Define four severity levels and four response classes so every decision—paging, communications, postmortem, compensation—keys off a defensible rule, not a judgment call.
- **Severity by impact scope**: SEV-1 (total payment-path loss, data corruption, or >5% transaction failure for >5 min); SEV-2 (significant degradation or single region loss); SEV-3 (limited impact with workaround available); SEV-4 (internal issue or cosmetic); SEV-0 (reserved for security/regulatory/privacy events).
- **Response class** (orthogonal to severity): Availability, Performance, Data Integrity & Ledger, Security & Privacy. **Key rule: class can raise severity, never lower it.** A SEV-3 data-integrity incident gets SEV-1 response posture because integrity is not recoverable by moving faster.
- Automatic triggers: loss of one AWS region → SEV-1 or SEV-2 (class-dependent); ledger write failures → SEV-1; replication lag >10s → escalation review; payment success rate <99% for >5 min → SEV-1/2; missed settlement window → SEV-1; total external API unavailability → SEV-1.
- Who may declare: any engineer, Support, account manager (based on observed customer impact). Who may downgrade: IC only, after investigation.
- Map each level to SLA credit exposure and to customer-facing status-page state.
- Include worked examples from the last 12 months so all 28 teams recognize their own incidents in the taxonomy. Re-validate quarterly against real declarations.
5. Incident lifecycle, Triage Owner rule, and escalation policy (depends on: 4)
Eliminate the "nobody was in charge for over an hour" problem by assigning ownership the moment a page is acknowledged.
- Define lifecycle states with clear entry/exit criteria: Detected (alert fired) → Triaged (is this real and customer-impacting?) → Declared (severity assigned) → Mitigated (core issue resolved) → Resolved (all verifications done) → Postmortem (review scheduled) → Closed (action items tracked or dismissed).
- **Introduce the Triage Owner rule**: the person who acknowledges the page owns the incident until an IC is assigned or the incident is stood down. There is never an unowned gap between first page and declaration. Triage Owner's sole job: decide within 15 minutes whether this requires an IC or a direct stand-down.
- Set aggressive acknowledgement and declaration targets: page acknowledged in 5 min; triage decision (is this real?) in 15 min; severity declaration in 30 min for any customer-facing incident.
- Implement automatic escalation ladders with no human judgment required: if responder does not acknowledge in 5 min, escalate to service owner; if no ack in 10 min, escalate to team manager; if no ack in 15 min, escalate to IC on-call. Escalation is never criticized.
- Define unresponsive-team path: if a service's on-call is unreachable for 30 min, IC may direct any available engineer from any team to engage.
- For SEV-1 and SEV-2: change freeze until IC declares mitigation confirmed; IC unfreezes changes explicitly.
- Enforce one incident, one record. Timeline auto-captured from Slack channel and bridge; never written from memory later.
6. Incident roles, command structure, and decision rights (depends on: 5)
Define clear roles with one-page responsibility cards published and linked from every paging notification.
- **Incident Commander**: owns incident outcome, not the fix. Declares severity, decides escalation, approves all customer communications, freezes changes, calls responders, hands off in shifts. Non-technical ICs are acceptable; technical depth is not required.
- **Deputy IC**: assigned within 15 min of declaration; shadows IC; takes over if IC unavailable or after 4-hour shift on SEV-1. Maximum IC shift: 4 hours on SEV-1, 6 hours on SEV-2.
- **Triage Owner** (new role): owns incident from first page acknowledgement until IC takes over or stand-down decision is made. Required for all incidents.
- **Communications Lead**: owns internal Slack updates and status-page messaging; shields IC from customer contact and interruptions.
- **Scribe**: records real-time timeline with decisions, actions, and key timestamps; not responsible for fixing.
- **Subject-Matter Responders**: engineers with service context; take IC direction; report only to IC; no side channels or parallel debugging.
- **Operations Lead** (SEV-1 only): coordinates multiple responders, manages incident bridge, maintains escalation list.
- Minimum viable staffing: SEV-1 requires all roles; SEV-2 requires IC, Deputy, Comms, Scribe, SMEs; SEV-3 requires Triage Owner and IC.
- Create laminated role cards for every on-call shift location (office, home, printed in pockets).
7. Three on-call rotations: Service, Platform, and Incident Commander (depends on: 5, 6)
Directly address the "carrying a pager for another team's code" objection by making it structurally impossible.
- **Service On-Call rotation** (federated): each of the 28 teams maintains a rotation for their own services only. No engineer is paged for code their team does not own. The answer to "why am I carrying a pager?" is now simply: "for your team's code."
- **Platform Duty rotation** (centrally staffed): shared PostgreSQL cluster, Kubernetes, networking, CI/CD, observability, and incident management tooling. Nobody's product code, so it gets its own dedicated rotation. Staffed from platform teams plus volunteers from other teams; paid at premium rate.
- **Incident Commander roster** (24x7): 12–16 certified senior engineers from across all 28 teams, on one-week primary shifts with secondary backup. Covers every hour with no single point of failure and no uncovered holiday week.
- **Consequences and gates**: 16 of 28 teams have no on-call today. Each must either (a) build a Service On-Call rotation of at least 6 engineers, or (b) formally transfer service ownership to a team that will, with transfer documented and dated. No exceptions, no waivers. Unowned services are decommissioned or transferred by end of month 2.
- Merge small or low-traffic teams into shared rotations where service ownership is unclear (e.g., shared analytics, testing infrastructure).
- Enforce scheduling limits in the tooling: no engineer on-call more than 2 weeks per quarter, automatically enforced by configuration, not negotiation.
- Publish a coverage matrix for all 28 teams showing services, owners, rotation size, gaps, and monthly status.
8. On-call compensation, rest policy, and sustainability (depends on: 7)
Settle compensation before rollout, not during negotiations. Make on-call sustainable and valued.
- **Paid on-call**: effective immediately upon joining a rotation. Weekly stipend while on shift (benchmark to New York market: $600–1,000 per week per engineer), regardless of incident volume.
- **Event-based compensation**: 1.5× hourly rate for time spent mitigating out-of-hours incidents, minimum one-hour block per callout. Tracked by incident record (auto-capture from timeline).
- **Compensatory rest**: no engineer works a normal 8-hour business day after a night incident requiring >2 hours mitigation. Rest day is documented policy, not a favour granted by manager.
- **Intrusion cap**: maximum 3 unscheduled pages per week per engineer. Exceed the cap in a week and trigger an immediate review; exceed in a month and escalate to VP Engineering. Breaches are structural signal that alert quality or service stability has a problem.
- **Voluntary opt-out**: an engineer may exit a rotation; their team must hire or buy replacement coverage from paid pool at published internal rate ($X per shift). This converts culture debate into visible budget decision.
- **Amnesty policy**: incident records, near-miss reports, and false declarations are never used in performance reviews or compensation discussion. Only failure to report is a performance issue.
- **Policy publication**: publish compensation structure and effective date before any team is asked to join a rotation, and include on-call expectations in job descriptions and hiring conversations.
- **Semi-annual review**: reassess compensation and caps every six months against actual page volumes, attrition rates, and survey feedback.
9. Detection strategy: SLOs, synthetic monitoring, and customer-report intake (depends on: 4, 7)
Close the 40% customer-detected gap by monitoring customer journeys instead of infrastructure metrics.
- **SLO-based alerting**: Define SLIs and SLOs for the top 20 customer journeys (payment initiation, authorization, settlement, ledger read/write, API availability, webhook delivery, payout). Measure per region. Alert on SLO breach, not on infrastructure metric (e.g., alert on "payment success rate <99%" not "database CPU >80%").
- **Synthetic transaction monitoring**: deploy synthetic transactions from outside AWS in both regions plus a third geographic location, one-minute cadence, for all money-moving paths. These are your first alarm bell.
- **Ledger-critical signals**: PostgreSQL replication lag (target: <1s, alert >5s), connection saturation, write latency (p95), lock-wait time, transaction ID exhaustion proximity, checkpoint pressure, table bloat. These are separate alerts on shared-database health.
- **Customer-report intake** (new detection channel): Support and Account Managers can raise an incident directly in the platform. Every customer report creates an incident record automatically, and the "customer report" detection source is counted in all metrics. This is a legitimate detection method, not a failure.
- **Detection-gap rule**: whenever a customer reports an incident before internal monitoring fires, auto-create a ticket in the owning service's backlog with root cause: "Monitoring gap on [journey]."
- **Detection contract per service**: every one of the 180 services needs a named owner, at least one symptom-based alert mapped to a SLO, and a documented expected detect time (target: <5 min for payment path, <10 min for others). Published on wiki and reviewed monthly.
- **Detection drills**: run a quarterly drill per team: simulate a broken service in staging and verify it triggers a page before a human notices.
- **Resilience roadmap separation**: detection improvements do not protect against ledger corruption or multi-region failure. Fund a separate resilience roadmap to reduce shared-database blast radius and improve failover safety.
10. Alert quality standards and noise-reduction program (depends on: 9)
Cut the 3,400 monthly alerts (85% noise) to <600 with 95% signal. This is the price of admission for on-call buy-in.
- **Paging contract**: every page must satisfy all of (1) symptom-based (customer impact, not infrastructure cause), (2) actionable (linked runbook with immediate next step), (3) owned (named team responsible), (4) severity-mapped (SEV-1/2/3/4), (5) SLO-linked where applicable. **No runbook, no page.** Enforce with CI check on alert definition.
- **Separation rule**: only customer-impacting or imminent-impact signals trigger pages. Everything else becomes a dashboard entry, a ticket, or a log line. Noisy infrastructure metrics go to dashboards, not pagers.
- **Page budget per service**: no service may exceed 2 pages per on-call shift per month. Exceeding budget auto-opens a remediation ticket in the owning team's backlog (with alert-quality review assigned to tech lead).
- **Automatic suppression rules**: (1) silence alerts if service auto-recovered within 30s, (2) suppress known maintenance windows, (3) group flapping alerts (>5 in 2 min) into one page, (4) rate-limit noisy services (max 1 page per 5 min until condition clears). All suppression rules must have an expiry date; no permanent silence without a ticket.
- **Probation for new alerts**: new alert rules run as tickets only and alert to a Slack channel; after two weeks of proving actionability (every alert resulted in human action), they graduate to pager.
- **Noise sprint**: run a focused 90-day program with a public burn-down of the top 100 noisiest rules. Assign each to a named manager. Default action: fix root cause, tune threshold, or delete within 10 working days. Deletion is a legitimate successful outcome (celebrate it).
- **Correlation and deduplication**: consolidate alert sources at ingest pipeline so one outage triggering 40 alerts produces one page, not 40.
- **Alert ownership**: every alert must have an owning team and a maintenance contact. Update monthly.
11. Incident tooling consolidation and integration (depends on: 5, 10)
Replace six alert tools and ad-hoc incident records with a single source of truth that unifies paging, escalation, timeline, and audit evidence.
- **Tool selection**: choose an incident-management platform (e.g., PagerDuty, Incident.io, Opsgenie) that integrates paging schedules, escalation policies, incident records, postmortem workflow, and status-page APIs. Decision gate: month 1.
- **Event pipeline consolidation**: route all alerts from the six legacy tools into a single event pipeline that feeds the incident platform. Apply deduplication, correlation, severity/class mapping, and rate-limiting at ingest.
- **Observability integration**: connect the incident platform to your Kubernetes dashboards, PostgreSQL monitoring, distributed tracing, and logs so responders see context in one pane. Link runbooks directly into incident records.
- **Slack and bridge integration**: auto-create incident Slack channels, auto-invite roles, auto-capture timeline from channel transcript and voice-bridge recording. Timeline is not written from memory; it is auto-captured.
- **Golden incident file**: implement the export defined in S3. One click produces a complete, immutable, audit-ready PDF: severity, timeline, roles, decisions, communications, postmortem, action items, and closure evidence.
- **Dual-run period**: run both legacy and new platform in parallel for two weeks. Define rollback criteria (e.g.,
12. Escalation automation and incident lifecycle enforcement (depends on: 5, 11)
Eliminate judgment calls from the worst moments. Escalation is automatic, mechanical, and blameless.
- **Automatic escalation ladders**: page responder → if no ack in 5 min, page service owner → if no ack in 10 min, page team manager → if no ack in 15 min, page IC on-call + call them immediately (phone + SMS + Slack). No human decides to escalate; timers fire escalations.
- **Severity-based escalation tempo**: SEV-1 uses faster timers (2 min for IC on-call), SEV-2 uses moderate timers (5–10 min), SEV-3 uses slower timers (15–30 min). Configured in tooling, reviewed quarterly.
- **Dual IC rule**: if a second SEV-1 incident is detected while the first is active, immediately page and assign a separate IC. ICs never run two incidents in parallel.
- **Change freeze and rollback authority**: SEV-1 and SEV-2 trigger automatic deploy freeze. Only the IC (with CTO/VP Eng notification) may unfreeze. Freeze lifts only when IC explicitly declares mitigation confirmed and verifies no new incident symptoms for 5 min.
- **Unresponsive team escalation**: if service's on-call does not acknowledge in 30 min, IC may direct any engineer from any team (volunteers first, then rotated) to engage. This is documented and reported in monthly review (escalation = signal of rotation problem).
- **One incident, one record**: all decisions logged in the incident platform. Auto-capture from Slack, bridge, status-page updates. Timeline is the source of truth; postmortem is written from timeline, never constructed after the fact.
- **Ambiguity rule**: if two responders disagree about whether an incident should be declared, it is declared. False declarations (stand-downs within 30 min of declaration) are tracked as metrics and closed without blame.
- **Watch state**: an unconfirmed incident can live in "Watch" state for max 30 min; after that, either declare it or stand it down explicitly.
13. Internal, customer, and regulatory communications workflows (depends on: 6, 12)
Define who informs whom, in what order, via what channel, with explicit timings and pre-approved templates.
- **Internal cadence**: first update to #incidents Slack channel within 3 min of declaration (even if "Investigating"). Then updates every 5 min (SEV-1), 15 min (SEV-2), or 30 min (SEV-3), or immediately on material change (e.g., mitigation achieved, scope widened). **Comms Lead owns the update; IC must not be interrupted.**
- **Executive notification**: IC calls CTO and VP Eng within 1 min of SEV-1 declaration (not email, not Slack, call). Incident declared in Slack with severity label, IC name, and affected service. Escalation channel lead auto-pinged.
- **Customer communication channels**: status page (all 2,100 customers), direct email to affected customers (top-tier accounts and customers affected by SEV-1), account-manager calls (top 50 accounts on SEV-1).
- **Status page timings**: update within 3 min of SEV-1 declaration, 10 min of SEV-2, 30 min of SEV-3 (even if root cause unknown; use "Investigating" with next-update ETA). Updates every 5–30 min depending on severity. Always include next-update time.
- **Pre-approved templates**: draft customer-facing language for each severity and class in advance with Legal and Compliance. Templates specify impact language ("some of your transactions are delayed" not "our database failed"), workarounds if available, and next-update commitment. Never speculate on cause in customer communication.
- **Regulatory notification path**: identify incidents requiring regulator notification (NYDFS Part 500, money-transmitter rules, payment-card-network rules, securities disclosure). Build a clock matrix: event type → regulator → notification window → signer. Compliance owns all regulatory notifications (never Engineering). Pre-clear templates. Flag incidents to Compliance immediately upon declaration.
- **Account manager playbook**: contact tree for top 50 accounts, templated talking points (facts only, never speculation), escalation path if customer escalates, what to offer (service credit, technical deep-dive call).
- **Closing communication**: resolution notice, SLA credit impact, commitment date for written root-cause report, customer action required (none, or security update, etc.).
14. Status page infrastructure and customer-impact ledger (depends on: 13)
Make the status page reliable, customer-centric, and audit-ready. Track customer impact in a single durable record.
- **Status page decoupling**: host status page outside production failure domain (separate cloud, separate infrastructure, separate database). Integrate incident platform with status page so incident record drives all public updates. Status page survives total platform outage.
- **Component-to-journey mapping**: status page components map to customer journeys ("Payments", "Settlements", "Payouts", "Ledger API") not to internal services. Allow customers to subscribe to components; notify by email or webhook.
- **One-click update templates**: pre-fill status-page template with severity, impact language, next-update time, and estimated resolution. Comms Lead types minimal new info ("Root cause identified" or "Workaround available"), and updates auto-post.
- **Customer-impact ledger** (one record per incident): which customer accounts affected, which journey(s) impacted, exact start and end time of impact, estimated SLA-credit exposure. Use this single record for customer communications, credit computation, regulatory reporting, and annual review. No reconciliation of two versions of the same outage.
- **SLA credit automation**: compute credit based on duration × severity × customer tier → auto-generate customer notification → auto-post to finance system. Reconcile accrued vs. paid credits monthly and report in executive review.
- **Testing during game days**: simulate status-page outage and verify alerts continue to fire; test total region loss and confirm status page remains updated; drill runbook for manually updating status page if platform is down.
15. Postmortem policy: mandatory, blameless, three-level framework (depends on: 6, 13)
Turn postmortems from a writing exercise (11 of 64 action items closed) into the learning engine of the system.
- **Mandatory postmortems**: all SEV-0, SEV-1, and SEV-2 incidents; all SEV-3 with customer impact or repeat pattern; any near-miss IC flags; any incident where the process itself failed (IC unreachable, Comms Lead unavailable, false declaration, missed update SLA).
- **Three-level framework** (proportionate to weight): (1) lightweight async review for SEV-4 and low-impact SEV-3 (10 min template in shared doc, owner + IC review), (2) standard facilitated postmortem for SEV-2 and impactful SEV-3 (full template, facilitated by trained neutral party, published within 10 days), (3) full executive postmortem for every SEV-1 and every security incident (executive sponsor assigned, full investigation, published within 15 days, customer-facing variant prepared).
- **Fixed timeline**: draft postmortem within 5 business days, blameless review within 10 days, internal publication within 15 days.
- **Single template**: impact (who, how many, how long, financial exposure), timeline (detection through resolution), root cause (not "human error" but system condition that enabled error; what was the gap?), contributing factors (tooling, process, organization, knowledge, monitoring), what went well, what went badly, action items (≤3, rest go to reliability backlog).
- **Blameless facilitation**: train a pool of blameless postmortem facilitators (target: 10+ engineers). Require a trained, neutral facilitator for every SEV-1 and SEV-2 review. Prohibit counterfactual language ("if the engineer had"), blame language, and the phrase "human error" as a root cause.
- **Publication rule**: publish all postmortems internally by default; security review only for genuinely sensitive material (e.g., unpatched vulnerability details or customer PII in logs). Create a customer-facing root-cause report for every SEV-1, especially for regulated customers, with legal and compliance sign-off.
- **Searchability**: store postmortems in a searchable wiki or issue tracker with tags (service, class, root cause category) so teams can learn from similar incidents without repeating them.
16. Action item tracking, reliability backlog, and repeat-incident design rule (depends on: 15)
Close the loop on incident learning by enforcing verifiable, tracked action items and breaking cycles of repeat incidents.
- **Action item capping**: each postmortem generates a maximum of 3 action items. Anything beyond 3 goes into a ranked reliability backlog, not into the postmortem, to prevent overwhelming teams.
- **Action item requirements**: each item must have (1) a named human owner (not a team), (2) a due date (≤60 days, target ≤30 days), (3) a definition of done (merged code change, tested alert, audit evidence, architectural decision, new runbook, training completed) not self-reported status.
- **Single reliability backlog**: create one backlog in your engineering tracker (Jira, Linear, etc.) with mandatory label (e.g., `incident-action`), link to originating incident, and link to postmortem. Track progress weekly.
- **Closure sign-off**: Incident Commander or postmortem facilitator must sign off on closure, verifying artifact exists (code merged, alert tested in drill, runbook verified).
- **Repeat-incident rule**: if the same service or component has a second incident with the same contributing factor, **do not create another action item**. Instead, escalate immediately to an architect or tech lead and trigger a design review (not a task, a review). This breaks the cycle of repeated patches; the system needs a structure change.
- **Capacity protection**: reserve a fixed percentage of each team's sprint capacity (10–15%) for reliability work. Track unspent capacity and report to VP Engineering monthly; if a team is not spending it, work with them to identify and fix blockers.
- **Ageing and escalation**: run a weekly review of open actions; escalate anything >30 days overdue to team lead and VP Engineering. Monthly report: completion rate (target >90% within 60 days) and median action age (target <30 days).
17. Training, certification, and exercise program (depends on: 6, 13, 15, 16)
Build skills before deploying the process. Run ongoing drills so the system is tested, not guessed at.
- **Curriculum**: (1) All engineers (30-min async video): severity taxonomy, communication expectations, postmortem format, when to declare an incident, where to find runbooks. (2) On-call responders (2-hr workshop): alert tool walkthrough, playbooks by severity, escalation paths and timers, when to call manager, mobile app walkthrough, case studies from the last 12 months. (3) Incident Commanders (4-hr bootcamp + test): leadership under pressure, decision-making (severity, escalation, rollback), communicating with executives, status-page discipline, postmortem facilitation, handling ambiguity, live simulated incident (pass/fail certification). (4) Communications Leads (2-hr training): templates per severity and class, customer-communication rules (no speculation, no blame), update timings, how to shield IC, regulatory triggers.
- **IC certification**: written assessment (75% pass required) plus live simulated incident (role-play with facilitator, graded on severity declaration, escalation decisions, communication, handover). Certification valid for 12 months; recertify via annual refresher or another live sim.
- **Depth across teams**: certify at least 2 ICs per team or team group so central roster is not siloed in one group; no holiday week is uncovered.
- **Async content**: record all training videos so async teams can learn on their schedule. Create quick-reference cards (laminated, pocket-sized) for roles and playbooks; distribute to on-call locations (office, home).
- **Monthly tabletop exercises**: drawn from real incidents from the last 12 months (region loss, ledger write failure, missed settlement window, cascading failures). Facilitator describes scenario; 3–4 responders play out response (Triage Owner, IC, Comms) as if real. Run 30 min; retro for 15 min afterward.
- **Quarterly game days**: deliberately inject failures into production (database failover, status-page outage, alerting-pipeline outage, dual SEV-1 incidents). All on-call roles engage. Run 2–3 hours; measure response times, decision quality, and communication. Document findings and create action items for identified gaps.
- **Drill the process's own failure modes**: IC unreachable (on-call unavailable, phone broken), Comms Lead on PTO, two simultaneous SEV-1s, paging storm (100+ alerts), false alarm that consumes an hour. Test escalation paths, deputy takeover, and recovery.
- **New-engineer onboarding**: add incident-response module to all engineering onboarding (completion tracked, audit-ready). All engineers must complete within 30 days of joining or transferring in.
18. Metrics, dashboards, and review cadence (depends on: 2, 12, 16, 17)
Measure to prove the system works. Publish live dashboards so every engineer sees the scoreboard and the system is transparent.
- **Outcome metrics**: Median Time to Detect by source (target: <5 min internally detected, <10 min customer journeys); Median Time to Mitigate for SEV-1/2 (target: <60 min SEV-1); customer-detected incidents as % of total (target: <5%); incidents by severity (should be mostly SEV-3/4, few SEV-1); SLA credits paid (target: <$100K/year by month 12); annual credit avoidance vs. program cost.
- **Process metrics**: IC assigned within 5 min (target: >95% of incidents); page acknowledgement rate (target: >98% within 5 min); first-update timeliness (target: >95% within SLA); postmortem timeliness (target: 100% of mandatory postmortems published on time); IC roster coverage (zero uncovered hours, monitored weekly).
- **Health metrics**: alert volume and signal-to-noise ratio per team (trending toward target); off-hours pages per engineer per month (trend, cap enforcement); on-call satisfaction survey (target: >7/10); training completion by team (target: 100% within 30 days); % of services with active detection contract (target: 100%).
- **Never publish incident count as a team metric.** Reward hiding. Instead publish detection metrics (near-misses reported per team, detection gaps closed, false declarations made).
- **Live dashboards**: build dashboards visible to all engineers (not just managers) showing outcome, process, and health metrics. Auto-populate from incident platform and alert tool. Update daily. Link from Slack and internal wiki.
- **Baseline all metrics against S2 evidence pack.** Set 90-day and 12-month targets for each metric. Publish targets and progress monthly.
- **Review cadence**: (1) weekly 30-min operational review (incident by incident from prior week: what went well, what hurt, actions); (2) monthly 60-min reliability review (trends, top causes, action-item aging, alert quality per team); (3) quarterly 60-min executive review (CEO's office: customer impact, SLA credits, top five systemic causes, program ROI).
- **Quarterly process review**: what in the process wasted responder time, what confused people, what should be deleted. Solicit feedback from ICs, Comms Leads, and responders. Document changes and reasoning.
19. Pilot program with 3–4 volunteer teams (depends on: 6, 8, 11, 12, 13, 14, 15, 16, 17, 18)
Do not roll out untested to 28 teams. Run the entire process end-to-end with a small cohort using real incidents as the primary training material.
- **Team selection**: recruit 3–4 volunteers spanning criticality: one payment-path team, one ledger-adjacent team, one shared infrastructure team (platform or Kubernetes), one low-traffic team. Volunteers see early adoption and influence.
- **Full process in pilot**: new severity and class taxonomy (S4), consolidated tooling (S11), roles and Triage Owner (S5–6), three rotations (S7), escalation automation (S12), communications (S13–14), postmortems (S15), action tracking (S16), paid on-call (S8), training (S17), metrics (S18). This is not a partial test; it is the complete system.
- **Real incidents are the training**: hold a retro within 48 hours of each pilot incident (while memory is fresh). Process Owner facilitates. Discuss: what worked, what hurt, how is the runbook, is the alert tuned, did Comms template work, did roles work, was timeline auto-captured correctly. Document feedback and iterate.
- **Instrument against baseline**: measure all S18 metrics in pilot; compare results with S2 baseline. If MTTD dropped from 22 min to 8 min, that is the number you carry into the next conversation.
- **Weekly retros with pilot teams**: structured 30-min sessions reviewing policies, tooling, training, templates, and process changes. Change the written policy or tooling in response; document what changed and why so later teams see the evolution, not just the finished state.
- **Explicit exit criteria**: (1) rotation coverage sustained (≥6 engineers per rotation, zero unacknowledged pages over 3 weeks), (2) postmortems delivered on time (100% of mandatory postmortems published within 15 days), (3) action tracking established (100% of action items in backlog with owner and date), (4) metrics live (dashboards updated daily, first weekly review completed).
- **Pilot report**: document before/after numbers (MTTD, MTTR, alert noise, action completion rate, on-call satisfaction) and key process learnings. This report is the foundation for every conversation in the rollout.
20. Phased rollout sequenced by cost of failure (depends on: 16, 18, 19)
Rollout is a staged migration with readiness gates and sequencing that prioritizes visible impact, not ease.
- **Wave sequencing**: divide 28 teams into 4 waves of ~7 teams each, **ordered by incident density and customer-journey ownership** (highest-cost-of-failure teams first). Teams with the most SLA credits at stake go first; their improvement is the proof.
- **Wave spacing**: three weeks between waves. This gives each wave time to stabilize and find problems before the next cohort joins.
- **Readiness checklist per team**: (1) all services mapped and owned (no unowned services), (2) alerts cleaned to paging contract (runbook linked, severity mapped), (3) playbooks updated and tested in staging, (4) rotation staffed to ≥6 engineers, (5) team completes training module, (6) manager briefed on policy, (7) on-call compensation in effect.
- **Gate review before each wave**: Process Owner holds gate review with target teams. Move unready teams to next wave with a dated remediation plan. No exceptions, no waivers; readiness is non-negotiable.
- **Wave champion**: assign a named engineer per wave to champion the rollout, answer questions, escalate issues to Process Owner. Champions are not representatives; they are advocates and feedback collectors.
- **Communication cadence**: weekly all-hands or newsletter for 4 weeks before each wave. Explain why (owned-code-owned-pager rule, paid on-call, SLA credit savings). Use pilot numbers. Answer FAQs. Announce champion and escalation path.
- **First incident under new process**: hold a retro within 48 hours. Feed accepted process changes back through change control.
- **Retire legacy tools and processes**: at end of each wave, retire legacy alert tools, informal escalation lists, ad-hoc status-page process. No parallel processes running for >3 weeks; this prevents confusion and half-learning.
- **Sequence to avoid audit collision**: ensure no team is rolling out in the same week as the audit dry run (S21).
21. SOC 2 dry run and evidence review (depends on: 3, 20)
Convert a good working process into a provable one. Test control evidence a few months before auditors arrive, when you can still fix gaps.
- **Dry run timing**: run 6 weeks before audit window (around month 7 of this program).
- **Scope**: sample 10–15 real incidents from pilot and early rollout waves. For each incident, verify evidence artifact exists and is complete: incident record, timeline (auto-captured), severity and class declaration, roles assigned and logged, communications log (Slack + status page), postmortem (if mandatory), action items in tracker with due dates, action completion evidence (code, alert test, training record).
- **Control walkthrough**: walk through each control statement from S3 with a checklist. Is the evidence artifact present? Is it immutable? Is it searchable? Is retention adequate? Is access logged?
- **Gap remediation**: for every gap found, estimate time to fix and prioritize by audit risk. Anything risking a qualified opinion (e.g., missing postmortem, no timeline evidence) must be fixed before the audit. Test the remediation against a new incident or a resample.
- **Interview readiness**: brief 10–15 engineers who may be interviewed by auditors (ICs, Comms Leads, Process Owner, team managers). Ask them to describe the process as they actually practice it, not as written. Listen for confusion or gaps in understanding. Correct them.
- **Auditor package preparation**: assemble process documentation, sample incident records (5–10 complete golden files), training records, on-call schedules, alert quality metrics, action tracker register, and status-page archive. Organize by control. Create a table of contents and index.
- **Single audit liaison**: designate Process Owner or a small dedicated compliance person as sole point of contact for audit requests. Prevents requests scattering across 28 teams.
- **Rehearsal**: conduct mock interview with an IC and a Comms Lead. Auditors ask tough questions under pressure ("How do you know the timeline is accurate?", "What happens when both ICs are unavailable?", "Show me how you proved the alert was actionable."). Practice answering.
22. Standing governance and continuous improvement (depends on: 20, 21)
Lock in durable improvement. The classic post-audit failure is the process freezing and then decaying. This step prevents that.
- **Standing Incident Management Council**: chaired by Process Owner, monthly meetings, attendees: CTO, VP Engineering, Head of Support, Head of Compliance, one engineering manager per region, one IC, one team manager from recent wave. Agenda: metrics review, policy changes, gaps from recent incidents, escalation for contentious issues.
- **Change control mandate**: give Process Owner documented authority to change severity taxonomy, response classes, roles, communications timings, and compensation policy. Any change requires: written justification, steering group approval (monthly cadence), and documented effective date before implementation. This prevents silent drift and ensures changes are deliberate.
- **Quarterly validation**: re-validate severity and response class taxonomy against all incidents from the prior quarter. Ask: did our taxonomy correctly predict response posture? Did we misclassify? Update taxonomy if patterns emerge.
- **Annual metric re-baselining**: every 12 months, re-run measurements from S2 (alert census, incident register) to reset targets. System should improve; targets should tighten.
- **Resilience roadmap separation**: fund a distinct architectural or platform roadmap for incident prevention (reduce shared-database blast radius, multi-region failover, deploy safety, observability investments). Better incident response does not protect a single-ledger corruption or unplanned failover. These are separate problems.
- **Quarterly executive report**: CTO and VP Eng report to CEO/CFO on metric set (detection time, mitigation time, SLA credits, customer-detected %), top five systemic causes of incidents, program cost vs. credit avoidance, and strategic architecture changes in flight.
- **Public backlog of improvement ideas**: teams and engineers propose process improvements via Slack or wiki. Process Owner reviews quarterly and implements accepted ideas (e.g., "add a dashboard for detection gaps", "update postmortem template"). Publish what changed and why.
- **Celebration and learning**: share wins publicly each quarter ("We reduced MTTD from 22 min to 5 min", "Customer-detected incidents down 80%", "$800K SLA credits avoided"). Refresh training and tabletop program annually and immediately after any SEV-1 to keep the system sharp and responsive to new scenarios.
--- PROPOSAL 2 ---
Proposal ID: abac560b-e42c-412b-9ecc-a60d55c29045
Content:
Estimated Complexity: high
Success Metrics: - Median time to detect falls from 22 minutes to under 8 minutes by month 6, and under 5 minutes for money-moving journeys by month 9.
- Share of customer-impacting incidents first detected by customers falls from 40% to below 10% by month 9.
- Median time to mitigate for SEV1 falls from 3 h 10 min to under 60 minutes by month 9.
- 100% of SEV1 and SEV2 incidents have a named Triage Owner within 5 minutes and a named Incident Commander within 15 minutes, from month 4.
- Zero incidents with unclear command authority lasting more than 15 minutes, measured monthly from month 4.
- Monthly page volume falls from 3,400 to under 600, with a false-positive rate below 15%, by month 5.
- No service exceeds two pages per on-call shift for three consecutive months by month 6.
- 100% of the 180 services have a named owner, a detection contract and at least one symptom-based alert by month 7.
- All 28 teams have a documented rotation, a service ownership map and at least one trained on-call engineer by month 6.
- 100% of on-call shifts are paid under a published policy from month 2, with zero on-call-attributed voluntary attrition by month 6.
- On-call satisfaction scores 7 out of 10 or better in quarterly surveys from month 6.
- The IC roster holds at least 12 certified ICs covering 24x7 with no uncovered week from month 4.
- 100% of SEV0, SEV1 and SEV2 postmortems are published internally within 15 business days from month 5.
- Action items closed within 60 days rise from 17% to above 90%, with a median action age under 30 days, by month 6.
- Status page first update is posted within 30 minutes on at least 95% of SEV1 incidents from month 4.
- Zero missed regulatory notification windows on any incident requiring notification.
- The customer-impact record is complete for 100% of customer-impacting incidents from month 5, with SLA credits reconciled monthly.
- Annual SLA credits paid fall at least 50% from the $1.3M baseline within 12 months.
- Customer-impacting incidents fall at least 40% year over year from the 31-incident baseline.
- Review cadence is sustained: weekly operational review in at least 90% of weeks, 12 of 12 monthly reliability reviews, 4 of 4 quarterly executive reviews.
- Every page has a recorded disposition — fixed, tuned or deleted — within 10 working days, from month 4.
- SOC 2 Type II is passed with zero findings related to incident response.
- 100% of new engineers complete the incident-response onboarding module within 30 days of joining.
Steps (24):
1. Mandate, one owner, and the evidence clock
This step turns the CEO's email into a funded programme with a single accountable owner, and it starts the SOC 2 clock on day one.
- Appoint a Director of Incident Management reporting to the CTO, with a dotted line to the COO for customer and SLA matters and to the Head of Compliance for audit readiness.
- Publish a one-page charter: scope covers every customer-impacting, money-moving, security and data-integrity event across all 28 teams and both regions.
- Grant explicit authority to declare an incident, set severity, freeze deploys, page any engineer in the company, and approve customer messaging.
- Fund the envelope up front: tooling, training and drill time, on-call compensation, and a small programme team, roughly $500–700K a year against $1.3M in credits paid.
- State the return plainly to the steering group: credits avoided, churn avoided, and audit findings avoided.
**Start the evidence clock now.** A SOC 2 Type II report covers a period of operating effectiveness, so evidence cannot be backfilled; an audit in eight months means the operating period begins in the first weeks of this programme.
- Stand up a fortnightly steering group of the CTO, VP Engineering, Head of Support, Head of Compliance and two engineering managers.
- Open a programme risk register with the top risks, owners and review dates, and revisit it at every steering group.
- Make participation in the incident process a documented performance expectation for every engineering manager, not an optional extra.
2. Baseline evidence pack and cost-of-downtime model (depends on: 1)
You cannot demonstrate improvement without a defensible starting point, and you cannot win the argument about noise without numbers.
- Build a 12-month incident register and re-classify all 31 customer-impacting incidents by detection source, duration, customers affected and credits paid.
- Build the silent-failure register: incidents where no internal alert fired at all, which is the number that explains the 40% customer-detected rate.
- Run an alert census per tool, per team and per service: volume, page-to-action ratio, off-hours interruptions, and the 50 noisiest rules with a named owner each.
- Survey every on-call engineer and manager on burden, fairness, escalation clarity and pay expectations, targeting above 70% response.
- Interview Support, Account Management and Sales about how customers learn of incidents and what they are promised.
- Reconstruct the two command-ambiguity incidents minute by minute to find exactly where ownership lapsed.
- Build a cost-of-downtime model: dollars per minute of impact per customer journey, used later to sequence teams and justify funding.
- Publish the pack internally as the problem statement, and retain every artifact as management-review evidence for the audit.
3. SOC 2 control mapping and evidence architecture (depends on: 1)
Most programmes leave compliance to the end; this one maps controls in month one, because the mapping decides what the process must capture from day one.
- Map the process to the Trust Services Criteria for incident identification, evaluation, response, containment, recovery and communication (CC7.1–CC7.5, CC2.2, CC2.3, CC4.1, CC3.x).
- Write each control in plain language, with one named owner and its evidence artifact.
- Define the golden incident file: one auto-assembled, single-click export per incident containing timeline, roles, severity, communications, postmortem, actions and closure proof.
- Set retention, storage location and immutability so no control depends on a laptop, a private channel, or chat history that expires.
- Keep a gap register with owners and dates, reviewed fortnightly by the steering group.
- Meet the auditor's readiness team inside the first 90 days to test the control design before anything is built on top of it.
4. Severity and class taxonomy with the trigger matrix (depends on: 2)
Severity decides how loud the response is; class decides which playbook runs and what must never be traded away.
- Define severity in customer and money terms: customers affected, dollars at risk, regions lost, ledger integrity, settlement delay, data loss and SLA exposure.
- Use SEV0 for security, privacy or regulatory events; SEV1 for total or material loss of a payment path; SEV2 for degradation or single-region loss; SEV3 for limited impact with a workaround; SEV4 for internal-only issues and near-misses.
- Define response classes: Availability, Performance, Data integrity and Ledger, Security and Privacy, Third-party dependency, Settlement and Financial, Process failure.
**Class can raise the response but never lower it.** A SEV2 data-integrity incident gets SEV1 posture, because integrity failures are not recoverable by moving faster.
- Specify automatic triggers: loss of one AWS region, ledger write failures, replication lag above threshold, a missed settlement cut-off, payment success rate below threshold for five minutes.
- State who may declare — any engineer, Support agent or account manager — and who may downgrade: the Incident Commander alone.
- Map every level to its SLA credit exposure and to the customer-visible status page state.
- Include worked examples from the last 12 months so each of the 28 teams recognises its own incidents, and revalidate the taxonomy quarterly against real declarations.
5. Day-one operating rules and the minimum viable process (depends on: 1, 4)
The full process will take months; the first useful version must be live in two weeks using the tools that already exist.
- Publish ten day-one rules that need no procurement: a named owner within five minutes, one channel per incident, one register entry per incident, one person speaking to customers.
- Apply the ambiguity rule: if two responders disagree about whether this is an incident, it is an incident.
- Make declaring free: a false alarm is closed as a false declaration, tracked as a metric, and never criticised.
- Require a register entry within 24 hours for every customer-impacting incident, even a minimal one.
- Ban silent incidents: if we know, the customer hears it from us rather than from their own reconciliation.
- Run the first 30 days on manual command, with a rotating duty Incident Commander drawn from the 12 teams that already have on-call.
- Hold a 15-minute daily incident stand-up during month one to catch friction while it is still fresh.
6. Roles, command structure and the no-unowned-minute rule (depends on: 4)
The two hour-long command failures did not happen at declaration; they happened in the gap before it, when an alert had fired and nobody owned it.
- Publish one-page role cards: Triage Owner, Incident Commander, Deputy IC, Operations Lead, Internal Comms Lead, Customer Comms Lead, Scribe, Subject-Matter Responders, and Executive Sponsor for SEV1 only.
- Introduce the Triage Owner rule: whoever acknowledges the page owns the incident until an IC takes over or the incident is stood down.
**The IC owns the incident, not the fix, and does not debug.** An IC who starts troubleshooting has abandoned command.
- Give the IC explicit decision rights: declare and escalate severity, freeze changes, halt deploys, approve customer messaging, and pull in any engineer.
- Define minimum viable staffing per severity: SEV1 fills every role; SEV2 staffs IC, scribe, comms and responders; SEV3 staffs an IC and a scribe.
- Set handover discipline: four-hour maximum IC shifts on SEV1 with a written handover, and a deputy named within 15 minutes of declaration.
- Set responder behaviour: one channel, one bridge, no side channels, and every request phrased with a named owner and a time.
- Link the role cards from every paging notification so they are one tap away at 3 AM.
7. Lifecycle, declaration and escalation policy (depends on: 4, 6)
This step defines the mechanical path from an alert to a declared incident and back to normal service, removing judgment calls from the worst moments.
- Define states with entry and exit criteria: Detected, Triaged, Declared, Mitigated, Resolved, Postmortem, Closed, plus a Watch state with a hard 30-minute timer.
- Set targets: page acknowledged within five minutes, triage decision within fifteen, declaration within thirty for anything customer-facing.
- Build automatic escalation ladders with timers at every layer: responder, service owner, team manager, IC on-call, VP Engineering.
- Define the unresponsive-team path: fifteen minutes escalates to the team manager, thirty minutes to the IC, who may then direct any available engineer.
- Define change freeze and rollback authority during SEV1 and SEV2, with the single named condition that lifts the freeze.
- Enforce one incident, one record, with the timeline captured automatically from the channel and bridge rather than written from memory afterwards.
- Add the two-simultaneous-incidents rule: a second IC is designated, and no IC ever runs two incidents at once.
- Test every escalation path weekly with synthetic pages, and adjust the timings after the first month of real operation.
8. On-call architecture across 28 teams (depends on: 4, 6)
The objection is that engineers will not carry a pager for another team's code; the answer is to build rotations that make the objection structurally impossible.
- Run a Service On-Call rotation per team, covering only that team's own services.
- Run a Platform Duty rotation for genuinely shared infrastructure: the PostgreSQL ledger cluster, Kubernetes, networking, CI/CD and observability.
- Run a central Incident Commander roster of 12–16 certified senior engineers on one-week shifts with a primary and a secondary.
**State the consequence honestly.** Sixteen of 28 teams have no rotation today; each must build one or formally transfer service ownership to a team that has one, with the transfer dated and recorded.
- Merge small or low-traffic teams into shared rotations so no rotation has fewer than six engineers.
- Cap load in the scheduling tool: no engineer on call more than two weeks per quarter, enforced by configuration rather than negotiation.
- Publish a coverage matrix of all 28 teams showing services, owners, rotation size and gaps, reviewed monthly.
- Write on-call participation into engineering job levels and hiring criteria, so the commitment is set before someone joins.
9. Compensation, rest and the economics of opting out (depends on: 8)
Unpaid on-call is the most cited reason for resistance, so settle compensation before rollout, not during it.
- Move to paid on-call: a per-shift stipend benchmarked to the New York market, agreed with HR and Finance, published with an effective date before any team is asked to join a rotation.
- Pay for callouts at 1.5× the hourly rate for time actually spent mitigating, with a minimum block per interruption.
- Provide documented compensatory rest: no normal working day after a night incident, and the rest day is policy rather than a favour granted by a manager.
- Cap intrusion with a maximum number of off-hours pages per shift, and trigger a mandatory review whenever the cap is exceeded.
**Allow opt-out, but put a price on it.** An engineer may step out of a rotation, and their team buys coverage from the paid pool at a published internal rate, which turns a cultural argument into a visible budget decision.
- Publish amnesty: incident records, near-misses and false declarations are never used in performance reviews; only failure to report is a performance issue.
- Protect against gaming: pay attaches to shifts and callouts, never to incident counts.
- Check the New York labour, overtime and tax treatment with Legal and Finance before announcing, and review the policy every six months against real page volumes, attrition and survey results.
10. Detection strategy: journeys, synthetic signals and customer-report intake (depends on: 4)
Customers detected 40% of incidents first, which makes detection the highest-leverage business problem in this programme.
- Define SLIs and SLOs for the top 20 customer journeys, measured per region: payment initiation, authorisation, settlement, ledger read and write, API availability, webhook delivery and payout.
- Alert on symptoms against those SLOs, not on cause-based infrastructure metrics.
- Add synthetic transaction monitoring from outside AWS in both regions plus a third location, on a one-minute cadence for money-moving paths.
- Add ledger and shared-PostgreSQL signals: replication lag, connection saturation, write latency, lock waits, transaction ID exhaustion, checkpoint pressure and bloat.
- Open a customer-report intake so Support and account managers can raise an incident directly, and count that path as a detection source in all reporting.
- Apply the detection-gap rule: whenever a customer reports first, a detection-gap ticket is opened automatically and owned by the responsible team.
- Publish a detection contract per service — owner, at least one symptom alert, documented expected detect time — for all 180 services.
- Run a detection drill per team: break something in staging and see whether it pages before a human notices.
11. Paging contract and the noise-reduction programme (depends on: 2, 10)
3,400 alerts a month at 85% noise is the reason engineers resent the pager, and fixing it is the price of admission for everything else in this plan.
- Publish the paging contract: every page must be symptom-based, actionable, owned, mapped to a severity and class, and linked to a runbook.
**No runbook, no page**, enforced by a CI check on the alert definition itself.
- Separate paging from ticketing: only customer-impacting or imminent-impact signals may wake a human.
- Set a page budget per team and per service, with a remediation ticket opened automatically, owned by the engineering manager, when the budget is breached.
- Put new alerts on two-week probation as ticket-only until they have proved actionable.
- Define the default action for a noisy alert: fix, tune or delete within ten working days, with deletion treated as a legitimate and celebrated outcome and the deleted count published.
- Deduplicate and correlate at the ingest pipeline so one root cause produces one page instead of forty.
- Require an expiry date on every silencing rule and temporary threshold, so suppression cannot quietly become permanent blindness.
- Run a 90-day noise sprint with a public burn-down of the top 100 noisiest rules, each owned by a named manager.
12. Incident tooling consolidation and the golden incident file (depends on: 3, 6, 10, 11)
Six alerting tools and no single incident record are structural causes of the 22-minute detection and the three-hour mitigation.
- Choose one incident platform for paging, schedules, escalation policies, incident records, postmortem workflow and status-page integration, and choose it within four weeks.
- Consolidate the six alert sources into a single event pipeline feeding that platform, with deduplication, correlation and severity and class mapping applied at ingest.
- Integrate observability, runbooks, chat and voice bridges so responders see dashboards and runbooks inside the incident record, and the timeline is captured automatically.
- Implement the golden incident file export defined in S3, so audit evidence is one click rather than a reconstruction.
- Run a dual-run period alongside the legacy tools with a defined rollback, then switch the legacy tools off on a published cutover date.
- Host the status page outside the production failure domain so it survives a total platform outage, and prove that in a game day.
- Make the platform usable from a phone: an Incident Commander must be able to run a SEV1 from a mobile device at 3 AM.
13. Internal, customer and regulator communications (depends on: 6, 7, 12)
Today the status page is written by whoever is around; this step replaces improvisation with a clock, a named owner and pre-cleared templates.
- Set the internal cadence: first update within 15 minutes of declaration, then every 30 minutes for SEV1 and hourly for SEV2, whether or not there is progress.
- Never let an employee learn of an incident from the status page: internal communication leads, external follows.
- Set the customer cadence: status page updated within 30 minutes of a SEV1 and 60 minutes of a SEV2, with a no-new-information update still mandatory.
- Define the channel hierarchy: status page for all 2,100 customers, direct email to affected customers on SEV1, and named account-manager calls for the top 50 accounts.
- Pre-approve templates per severity and class with Legal and Compliance, each carrying its own next-update time.
- Forbid speculation: customer messages never guess at cause, never assign blame, and never commit to a root cause before the postmortem.
- Define the closing communication: resolution notice, how SLA credits will be handled, and the committed date for the written report.
- Build a regulator clock matrix covering event type, regulator, notification window, signer and the shortest applicable clock, including NYDFS Part 500, money-transmitter and banking notifications, breach notification, card-network rules and public-company disclosure.
**The regulatory clock starts at awareness, not at root cause.** Route every notification through Compliance, never Engineering, and pre-clear the templates.
14. Customer trust workstream and the SLA credit ledger (depends on: 13)
The $1.3M in credits is a symptom of having no single record of customer impact, and the CEO's inbox is a symptom of customers learning things late.
- Maintain one durable customer-impact record per incident: which customers, which journeys, from when to when, and the estimated credit.
- Use that one record for communications, credit computation, regulatory reporting and the annual review, so nobody reconciles two versions of the same outage.
- Automate the credit path end to end — computation, approval, customer notification and finance treatment — so credits stop being a manual scramble three weeks later.
- Reconcile accrued against paid credits monthly and report the result in the executive review.
- Give the status page a named product owner and map its components to customer journeys, not to internal services.
- Send a CTO-signed reliability note to the top 50 accounts and publish a quarterly reliability report to all customers.
- Give account managers a script of the facts they may state, the speculation they may not, and a path for customer escalations.
- Track credit avoidance against programme cost, so the funding case stays a number rather than an argument.
15. Postmortems: mandatory set, three levels, blameless by design (depends on: 6)
Postmortems currently happen for some incidents, in various formats; this step makes them mandatory where they matter and light where they do not.
- Make postmortems mandatory for every SEV0, SEV1 and SEV2, every SEV3 with customer impact or a repeat pattern, every near-miss the IC flags, and every incident where the process itself failed.
- Use three levels so the ritual matches the weight: a lightweight async review for SEV3 and SEV4, a facilitated postmortem for SEV2, and a full review with an executive sponsor for SEV1.
- Set deadlines: draft within five business days, blameless review within ten, internal publication within fifteen.
- Adopt one template: impact, timeline, detection, response, contributing factors across tooling, process, organisation and human factors, what went well, what went badly, and actions.
- Train a pool of blameless facilitators and require a trained one for every SEV1 review, never the IC.
**Ban blame language in the template and ban "human error" as a root cause.** The question is always what system condition made the error possible.
- Publish all postmortems internally by default, with security review only for genuinely sensitive material.
- Produce a customer-facing root-cause report for SEV1 incidents, especially for regulated and top-tier accounts.
16. Action items: capped, verifiable, with the repeat-incident rule (depends on: 15)
Eleven of 64 action items closed is not a tracking problem; it is a generation problem, because the process produces more actions than the organisation can absorb.
- Cap each postmortem at three action items, with anything beyond that going into a ranked reliability backlog.
- Require every action to have a named human owner, a due date, and a definition of done that is an artifact: a merged change, an alert that fired in a drill, or a test that fails without the fix.
- Prohibit self-reported closure; closure requires the artifact and sign-off by the process owner or the Incident Commander.
- Keep one reliability backlog in the engineering tracker with a mandatory label and a link back to the originating incident.
- Reserve a fixed share of each team's sprint capacity for reliability work, with unspent capacity visible to the vice presidents.
- Apply the repeat-incident rule: a second incident with the same contributing factor triggers a design review owned by an architect, not another action item.
- Run a weekly ageing review of open actions and escalate anything more than 30 days overdue to the VP of Engineering.
- Report action completion rate and median action age monthly, by team.
17. Metrics, dashboards and the review cadence (depends on: 2, 4, 16)
Define what good looks like, then measure it in a way that rewards reporting incidents rather than hiding them.
- Define outcome metrics: time to detect by source, time to mitigate by severity and class, share of incidents detected by customers, incidents by severity and class, SLA credits paid, and customer-impact minutes.
- Define process metrics, because the process itself has SLOs: IC assigned within five minutes, page acknowledgement rate, first-update timeliness, update-cadence adherence, postmortem timeliness, action closure rate, action median age, and IC roster coverage.
- Define health metrics: alert volume and noise ratio per team, off-hours pages per engineer, on-call satisfaction, training completion, and share of services with a detection contract.
**Never publish incident count as a team metric.** It rewards hiding incidents; publish reporting metrics instead — near-misses filed, detection gaps closed, false declarations made.
- Publish live dashboards visible to every engineer, refreshed daily, with every metric baselined against the S2 evidence pack and given a 90-day and a 12-month target.
- Run a 30-minute weekly operational review incident by incident, a monthly reliability review of trends and actions, and a quarterly executive review with the CEO's office.
- Hold a quarterly review of the process itself: what wasted time, what confused responders, and what should be deleted.
- End every review with decisions and named owners, never with numbers alone.
18. Training, certification and the drill programme (depends on: 6, 7, 13, 15)
A process that lives only on a wiki page fails on the first real page, so skills are built and tested before they are needed.
- Build a practical curriculum: how to be on call, how to triage and declare, how to run an incident as IC, how to communicate, and how to write a postmortem.
- Require certification before joining the IC roster: a written assessment plus a live simulated incident.
- Certify at least two ICs per team group so the central roster has depth across all 28 teams and no holiday week is left uncovered.
- Train communications leads separately on templates, cadences, customer language and the regulatory rules.
- Run monthly tabletops drawn from the last 12 months of real incidents, including region loss, ledger write failure and a missed settlement cut-off.
- Run quarterly game days with deliberately injected failure, including PostgreSQL failover, status-page outage and alerting-pipeline outage.
- Drill the process's own failure modes, not just technical ones: IC unreachable, comms lead on leave, two simultaneous SEV1s, a paging storm, and a false alarm that burns an hour.
- Audit the process for single points of failure: who alone can perform each critical task, and what happens in their holiday week.
- Keep a mandatory onboarding module for every engineer joining or transferring in, with audit-ready completion records.
19. Pilot with three to four teams, using real incidents (depends on: 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 18)
Do not roll out to 28 teams untested, and do not rely on drills alone when real incidents are better training material.
- Recruit three to four volunteer teams spanning criticality: a payment-path team, a ledger-adjacent team, the platform team that owns the shared PostgreSQL cluster, and a low-traffic team.
- Run the entire process end to end in the pilot: severity and class, roles, escalation, communications, postmortems, action tracking, and paid on-call.
- Treat real incidents during the pilot as the primary training material, and hold a retro within 48 hours of each one, run by the process owner while the friction is fresh.
- Instrument the pilot against the baseline and publish before-and-after numbers.
- Hold weekly retrospectives with the pilot teams, and change the written policies, the tooling and the training in response, documenting what changed and why.
- Set explicit exit criteria: rotation coverage achieved, zero unacknowledged pages over two weeks, postmortems delivered on time, and actions tracked to closure.
- Produce a pilot report that every later rollout conversation starts from.
20. Phased rollout to 28 teams, sequenced by cost of failure (depends on: 16, 18, 19)
Rollout is a staged migration with readiness gates, not an email announcement, and the sequencing matters more than the schedule.
- Split the 28 teams into four waves of roughly seven, ordered by incident density and customer-journey ownership: the teams whose incidents cost the most SLA credits go first, because their improvement is the visible proof.
- Leave three weeks between waves and define a per-team readiness checklist: services mapped and owned, alerts cleaned to the paging contract, runbooks written, rotation staffed to at least six engineers, training complete, manager briefed, compensation in effect.
- Hold a gate review with the process owner before each team joins, and move an unready team to the next wave with a dated remediation plan rather than granting an exception.
- Give each wave a named champion and explain the why using the pilot's numbers.
- Handle resistance directly and early: publish the own-your-code-own-your-pager rule and the paid on-call mechanics before each wave, never after.
- Hold a retro within 48 hours of each wave's first incident under the new process, and push accepted changes through change control.
- Retire legacy tools, informal escalation lists and the ad-hoc status page process at the end of each wave on a published cutover date, so two processes never run in parallel for long.
- Sequence the waves so no team faces both a rollout and the audit dry run in the same week.
21. Audit dry run and evidence review (depends on: 3, 20)
This step converts a good process into a provable one, about six weeks before the auditors arrive.
- Sample real incidents from the pilot and the early waves against each control's evidence requirements.
- Remediate every gap found and re-test the remediated control against the same sample, prioritising anything that risks a qualified opinion.
- Brief every engineer who may be interviewed so they can describe the process as they actually practise it, not as it is written.
- Prepare the auditor package: process documentation, sample incident records, paging logs, communication logs, training records, on-call schedules, the action register and alert quality metrics.
- Designate one audit liaison and a small evidence-request team, so requests do not land on all 28 teams at once.
- Rehearse the walkthrough with an Incident Commander and a communications lead, because auditors probe realism under pressure.
- Keep the audit liaison and the process owner as close to the same person as possible, so accountability for the control is also accountability for the evidence.
22. Standing governance and process ownership (depends on: 20, 21)
The classic post-audit failure is that the process freezes and then decays, so ownership has to outlive the programme.
- Establish a standing Incident Management Council chaired by the process owner, meeting monthly with engineering, support, compliance and product representation.
- Give the process owner a documented mandate to change standards, with formal change control for any change to severity, classes, roles, communication timings or compensation.
- Revalidate the severity and class taxonomy quarterly against real declarations, and re-baseline the entire metric set annually.
- Report quarterly to the executive team on the metric set and the top five systemic causes of incidents.
- Feed incident themes into architecture review, release engineering and hiring, so the programme improves the system and not only the response.
- Keep a public backlog of process improvement ideas so any engineer can propose a change, and act on it every quarter.
23. Resilience track for the shared ledger and multi-region blast radius (depends on: 1, 17)
Better incident management cannot save a single shared ledger during a corruption event, so this runs as a funded engineering track beside the process, not inside it.
- Reduce ledger blast radius: read replicas for read paths, connection-pool isolation per service domain, statement timeouts, write throttles and per-tenant limits.
- Build a ledger protection mode: an automatic tripwire that degrades to read-only or queues writes before corruption can spread.
- Rehearse multi-region failover with the ledger as the binding constraint, and set a measured recovery time objective per region.
- Test restore and point-in-time recovery on a schedule, and publish the timings.
- Tighten deploy safety: progressive rollout, automatic rollback on SLO burn, and change-freeze windows around settlement.
- Track a small number of resilience bets with named owners and quarterly milestones, reviewed in the same executive review as the incident metrics.
24. Culture, incentives and continuous improvement (depends on: 20, 22)
Incident management is a cultural change before it is a process change, and the incentives decide whether reporting goes up or down.
- Restate the amnesty at every rollout wave: reporting, near-misses and false declarations are safe; hiding is not.
- Recognise detection improvements, deleted alerts, closed detection gaps and well-written postmortems publicly, and reward the engineers who do them.
- Correct blame language in leadership communication on the spot, including inside the executive review.
- Watch for pager fatigue with real data, and take an engineer off rotation as soon as the intrusion cap breaks.
- Run a quarterly process retro with the whole engineering organisation and publish what changed as a result.
- Refresh the training and tabletop programme annually and immediately after any SEV1.
- Share wins publicly each quarter — faster detection, fewer customer-detected incidents, credits avoided — to hold momentum after the audit.
--- PROPOSAL 3 ---
Proposal ID: 265b9bc8-32de-46d3-8546-6dd406da4c24
Content:
Estimated Complexity: high
Success Metrics: - Median Time to Detect (MTTD) < 5 minutes.
- Median Time to Mitigate (MTTR) < 60 minutes for SEV1.
- >90% of incidents detected internally (vs. 40% previously).
- SLA credits reduced to <$100K annually.
- Monthly alert volume < 600 with >95% signal-to-noise ratio.
- Postmortem action item completion rate > 80%.
- Zero incidents with command ambiguity > 15 minutes.
- 100% of SEV1/2 postmortems published within 15 days.
- Status page updates within 30 minutes for 95% of SEV1.
- 100% of teams integrated into on-call rotations.
- On-call satisfaction score > 7/10.
- 40+ certified Incident Commanders available 24x7.
- SOC 2 Type II audit passed with zero findings.
Steps (22):
1. Executive Mandate & SOC 2 Evidence Clock
Establish clear ownership and funding to transform the CEO's concern into a governed program with immediate compliance implications.
2. Baseline Measurement & Problem Statement (depends on: 1)
Quantify the current state to build a defensible problem statement and establish the data foundation for improvement tracking.
3. Severity & Class Taxonomy (depends on: 2)
Define unambiguous incident levels based on business impact to trigger specific response protocols.
4. Incident Command Roles & Structure (depends on: 3)
Standardize roles and decision rights to eliminate command ambiguity during active incidents.
5. Tooling Consolidation & Integration (depends on: 1)
Deploy a unified incident management platform to replace fragmented alerting tools and create a single source of truth.
6. Detection Strategy & SLO Implementation (depends on: 3)
Shift monitoring focus to customer-impacting symptoms to detect issues before customers do.
7. Alert Quality Standards & Noise Reduction (depends on: 5, 6)
Implement strict alerting rules to reduce noise and restore on-call trust.
8. On-Call Architecture & Coverage Model (depends on: 4)
Design a federated, team-owned on-call structure supported by central command coverage.
9. On-Call Compensation & Wellbeing Policy (depends on: 8)
Introduce paid, sustainable on-call rotations to remove resistance and mitigate burnout.
10. Playbooks & Communication Templates (depends on: 3, 4)
Create standardized playbooks and templates for each severity level to reduce cognitive load.
11. Escalation & Incident Lifecycle Policy (depends on: 4, 5)
Define the mechanical progression from alert to resolution, including automatic escalation paths.
12. Communication Workflows (Internal & External) (depends on: 4, 10)
Establish timed protocols for internal, customer, and regulatory communications.
13. Status Page & Customer Notification Automation (depends on: 5, 12)
Integrate incident management with public status pages to ensure timely and accurate updates.
14. Postmortem Policy & Blameless Framework (depends on: 3)
Mandate and standardize postmortems to focus on system improvement rather than individual error.
15. Action Item Tracking & Governance (depends on: 14)
Ensure postmortem action items are tracked, owned, and completed to prevent recurrence.
16. Metrics, Dashboards & Review Cadence (depends on: 2, 3)
Define and monitor key performance indicators to validate the effectiveness of the incident process.
17. Training, Certification & Drills (depends on: 4, 10, 12)
Build practical skills and readiness for incident response across all roles.
18. Pilot Program with Volunteer Teams (depends on: 5, 7, 9, 11, 13, 15, 17)
Test the end-to-end process on a small scale to refine procedures before full rollout.
19. Phased Rollout to All Teams (depends on: 18)
Deploy the process to all 28 teams in manageable waves to minimize disruption and manage resistance.
20. SOC 2 Control Mapping & Evidence Framework (depends on: 1, 3)
Align the incident management process with SOC 2 requirements to ensure audit readiness.
21. SOC 2 Dry Run & Audit Support (depends on: 19, 20)
Verify evidence sufficiency and prepare the team for the external audit.
22. Governance & Continuous Improvement (depends on: 19, 21)
Establish standing governance to sustain the process and drive ongoing improvements.
Instructions:
1. Carefully evaluate each proposal considering:
- Clarity and detail of the proposed steps.
- Overall alignment with the Main Objective.
- Quality and relevance of the proposed success metrics.
- Coherence of dependencies between steps (if any).
2. Think deeply about each proposal, analyzing its strengths and weaknesses. Consider how well each proposal addresses the main objective and which one would be most effective.
After your thorough analysis, give the Proposal ID of the selected proposal and a one- or two-sentence justification of your choice.
{"selected_proposal_id": "80c18452-b9c3-4820-9dec-cf8cb9a8c7b0", "justification": "Proposal 1 delivers the most actionable and detailed implementation path, with 22 concrete, richly specified steps covering all requirements. It combines aggressive but realistic timelines (month 5 complete rollout, month 8 audit-ready), unambiguous handling of the on-call resistance problem via three explicit rotations with no exceptions, and early control mapping (step 3) that positions for SOC 2 success. While Proposal 2 is nearly equivalent and slightly stronger on resilience separation and cultural sustainability, Proposal 1's clarity, specificity, and urgency-matched pacing directly serve the objectives of a company facing an immediate audit and critical operational failures."} |
| deepseek-flash_voter_2 | deepseek-flash · deepseek/deepseek-flash |
Proposal B · deepseek-flash_refine_2 winner | Proposal 2 matches Proposal 1's depth and specificity but edges ahead on the brief's hardest asks: its P5 'day-one operating rules' delivers a live minimum-viable process in two weeks, directly answering 'introduce it without waiting for the audit,' and its dedicated resilience track plus culture/incentives steps address the single shared PostgreSQL ledger and the on-call resistance that a pure process plan would miss. Its 23 concrete, time-bound success metrics and regulator clock matrix are at least as complete as Proposal 1's, while Proposal 3 is markedly thinner with one-line step descriptions and fewer metrics. |
19.2k in · 1.3k out · 7 s | show[SYSTEM]
You are an expert and objective evaluator of project plan proposals.
Your task is to select the BEST proposal based on criteria of completeness, clarity, and alignment with the main objective.
Use your internal reasoning processes to thoroughly analyze each proposal, considering all aspects and implications.
Take as much time and space as you need to evaluate each proposal in depth before making your decision.
Every text field you write will be read by a busy person, so write for them: short sentences, one idea each, in short paragraphs. Text fields accept Markdown: separate paragraphs with a blank line, use a bullet list when you enumerate things and plain paragraphs when you explain or argue, and bold at most one key phrase per paragraph. A text of more than three sentences must be split into paragraphs; never deliver a single unbroken block of text.
[HUMAN]
Main Objective: "A B2B payments platform in New York (260 engineers in 28 teams, 2,100 customers, $4B processed a month, a 99.95% SLA with service credits) runs about 180 services on Kubernetes across two AWS regions, with a shared PostgreSQL cluster for the ledger. In the last 12 months: 31 customer-impacting incidents, median time to detect 22 minutes (customers detected 40% of them first), median time to mitigate 3 h 10 min, $1.3M paid in SLA credits, two incidents in which nobody was sure who was in charge for over an hour, and a CEO email about "outages we hear about from clients". On-call exists in 12 of the 28 teams, unpaid, fed by alerts from six different tools — 3,400 alerts a month, 85% of them noise; there is no severity scale, status-page updates are written by whoever is around, and postmortems happen for some incidents, in various formats, with action items that are rarely tracked (11 of 64 closed). A SOC 2 Type II audit in eight months will test incident response. Engineers push back against "carrying a pager for other teams' code".
Define the incident management process: severity levels and what each one triggers; roles (incident commander, communications, scribe, subject-matter responders) and how they are staffed 24x7 across 28 teams; the on-call structure, rotations, compensation and the rules for alert quality; detection and escalation paths; internal and customer communications (status page, account managers, regulators when required) with their timings; postmortems (when mandatory, format, blameless review, ownership and tracking of actions); the metrics and reviews that show whether it works; and how it is introduced across the teams without waiting for the audit."
Proposals to Evaluate:
--- PROPOSAL 1 ---
Proposal ID: 80c18452-b9c3-4820-9dec-cf8cb9a8c7b0
Content:
Estimated Complexity: high
Success Metrics: - Median Time to Detect from 22 minutes to <5 minutes for internal detection and <10 minutes for customer journeys by month 6; <2% of customer-impacting incidents first detected by customers (vs. 40% currently) by month 9.
- Median Time to Mitigate for SEV-1 from 3h 10min to <45 minutes by month 9; SEV-2 to <90 minutes by month 9.
- Monthly alert volume from 3,400 to <400 with >95% signal-to-noise ratio by month 5; all alerts satisfy paging contract (symptom-based, owned, runbook-linked, severity-mapped).
- Zero incidents with command ambiguity lasting >15 minutes from month 2 onwards; 100% of SEV-1 and SEV-2 incidents have named IC logged in timeline within 5 minutes; 100% of Triage Owners assigned within acknowledgement.
- 100% of the 28 teams have a documented Service On-Call rotation or a signed service-ownership transfer by month 4; Platform Duty rotation and IC roster both staffed and live by month 3.
- On-call compensation policy published and effective by month 1; on-call satisfaction score ≥7/10 by month 6; zero on-call-attributed voluntary attrition by month 6.
- SLA credits paid from $1.3M annually to <$100K by month 12; credit avoidance (prevented credits) tracked and reported monthly.
- 100% of mandatory postmortems (SEV-0, SEV-1, SEV-2, and repeat incidents) published internally within 15 business days by month 4.
- Postmortem action item completion rate from 17% (11 of 64) to >90% within 60 days by month 6; median action age <30 days; zero repeat incidents caused by the same contributing factor without a design review.
- 100% of the 180 services have a named owner, a detection contract, and at least one symptom-based alert by month 6.
- Status-page first update published within 30 minutes for ≥95% of SEV-1 incidents by month 4.
- 100% of new engineers complete incident-response onboarding within 30 days of joining; IC certification includes written exam and live simulation; ≥2 certified ICs per team group; zero uncovered hours in 24x7 IC roster.
- Weekly operational review held in ≥90% of weeks; 12 of 12 monthly reliability reviews; 4 of 4 quarterly executive reviews; all reviews end with documented decisions and owners.
- SOC 2 Type II audit passes all incident-response controls (CC7.1–7.5, CC2.2–2.3, CC4.1, CC3.x) with zero findings by month 8.
- All pilot and rollout incidents captured with complete golden incident files (timeline, roles, communications, postmortem, actions, closure evidence) by month 3 onwards; audit dry run identifies zero critical gaps by month 7.
- All 28 teams transitioned to new process by month 5; all legacy alert tools decommissioned; single source of truth for incidents established and sustained.
Steps (22):
1. Executive charter, governance structure, and evidence clock
Turn the CEO email into a funded, authorized program with clear ownership and documented evidence collection for SOC 2, starting today.
- Appoint a Director of Incident Management reporting to CTO, with dotted line to COO (customer impact) and Head of Compliance (audit readiness).
- Publish a one-page charter: scope (all customer-impacting, payment-path, data-integrity incidents across 28 teams and 2 regions), decision rights (IC may freeze changes, override team preferences during incidents), and authority to mandate process participation.
- Secure annual budget for tooling, training, on-call compensation ($500–800K estimated), and resilience work. Connect funding to avoided SLA credits ($1.3M baseline).
- Establish standing Incident Management Steering Group: CTO, VP Engineering, Head of Support, Head of Compliance, one engineering manager per region. Meet monthly.
- **Start the SOC 2 evidence clock on day 1.** An audit in eight months means operating-period evidence begins now; design the process to capture evidence continuously, not retroactively.
- Make incident-response participation a documented performance expectation for all engineering managers and team leads.
- Publish timeline: working process in month 2, all 28 teams in month 5, audit-ready in month 7.
2. Baseline measurement, incident register, and evidence pack (depends on: 1)
Establish defensible baseline metrics and identify structural gaps that explain the 40% customer-detected rate and 22-minute detection time.
- Build a 12-month incident register with all 31 customer-impacting incidents: date, detection source, detection time, mitigation time, customer count, services involved, SLA credits paid, root cause class.
- Audit the current alert estate: total volume per tool, volume per team, volume per service, page-to-action ratio, top 50 noisiest rules, off-hours interruptions per engineer.
- Construct a **silent-failure register**: incidents with no internal alert fired at all. This explains the 40% customer-detected rate.
- Reconstruct the two command-ambiguity incidents minute by minute: exactly when did ownership become unclear, how long, what was the decision bottleneck.
- Survey on-call engineers (target >70% response): burden, fairness, pay expectations, escalation clarity, willingness to stay.
- Interview Support and Account Management: how do customers discover incidents, what do they complain about, how do they contact you.
- Publish the problem statement internally; retain all artifacts for SOC 2 audit evidence. This is the baseline against which all improvements are measured.
3. Control mapping and evidence architecture (depends on: 1, 2)
Design the process to generate SOC 2-compliant evidence automatically, from the first real incident, so the audit clock ticks in your favour.
- Map the new process to Trust Services Criteria CC7.1–7.5 (incident identification, response, evaluation, containment, communication), CC2.2–2.3 (authorization), CC4.1 (change management), CC3.x (information availability).
- For each control, write a one-paragraph plain-language statement, name a single owner, and specify the evidence artifact (incident record, timeline, communications log, postmortem, action tracker, training record).
- Define the **golden incident file**: one single-click export per incident containing severity, timeline, roles assigned, decisions made, communications sent, postmortem, and action items. This is the audit unit.
- Specify data retention, immutability, access control, and storage location (not laptops, not chat history that expires). Ensure evidence is searchable and organized by incident date.
- Keep a gap register with owners and dates; review fortnightly in the steering group. Identify which controls are blocked by incomplete rollout and when they unblock.
- Run an early design walkthrough with an experienced SOC 2 readiness partner inside month 1 to stress-test control design before building on it.
4. Severity and response class taxonomy (depends on: 2)
Define four severity levels and four response classes so every decision—paging, communications, postmortem, compensation—keys off a defensible rule, not a judgment call.
- **Severity by impact scope**: SEV-1 (total payment-path loss, data corruption, or >5% transaction failure for >5 min); SEV-2 (significant degradation or single region loss); SEV-3 (limited impact with workaround available); SEV-4 (internal issue or cosmetic); SEV-0 (reserved for security/regulatory/privacy events).
- **Response class** (orthogonal to severity): Availability, Performance, Data Integrity & Ledger, Security & Privacy. **Key rule: class can raise severity, never lower it.** A SEV-3 data-integrity incident gets SEV-1 response posture because integrity is not recoverable by moving faster.
- Automatic triggers: loss of one AWS region → SEV-1 or SEV-2 (class-dependent); ledger write failures → SEV-1; replication lag >10s → escalation review; payment success rate <99% for >5 min → SEV-1/2; missed settlement window → SEV-1; total external API unavailability → SEV-1.
- Who may declare: any engineer, Support, account manager (based on observed customer impact). Who may downgrade: IC only, after investigation.
- Map each level to SLA credit exposure and to customer-facing status-page state.
- Include worked examples from the last 12 months so all 28 teams recognize their own incidents in the taxonomy. Re-validate quarterly against real declarations.
5. Incident lifecycle, Triage Owner rule, and escalation policy (depends on: 4)
Eliminate the "nobody was in charge for over an hour" problem by assigning ownership the moment a page is acknowledged.
- Define lifecycle states with clear entry/exit criteria: Detected (alert fired) → Triaged (is this real and customer-impacting?) → Declared (severity assigned) → Mitigated (core issue resolved) → Resolved (all verifications done) → Postmortem (review scheduled) → Closed (action items tracked or dismissed).
- **Introduce the Triage Owner rule**: the person who acknowledges the page owns the incident until an IC is assigned or the incident is stood down. There is never an unowned gap between first page and declaration. Triage Owner's sole job: decide within 15 minutes whether this requires an IC or a direct stand-down.
- Set aggressive acknowledgement and declaration targets: page acknowledged in 5 min; triage decision (is this real?) in 15 min; severity declaration in 30 min for any customer-facing incident.
- Implement automatic escalation ladders with no human judgment required: if responder does not acknowledge in 5 min, escalate to service owner; if no ack in 10 min, escalate to team manager; if no ack in 15 min, escalate to IC on-call. Escalation is never criticized.
- Define unresponsive-team path: if a service's on-call is unreachable for 30 min, IC may direct any available engineer from any team to engage.
- For SEV-1 and SEV-2: change freeze until IC declares mitigation confirmed; IC unfreezes changes explicitly.
- Enforce one incident, one record. Timeline auto-captured from Slack channel and bridge; never written from memory later.
6. Incident roles, command structure, and decision rights (depends on: 5)
Define clear roles with one-page responsibility cards published and linked from every paging notification.
- **Incident Commander**: owns incident outcome, not the fix. Declares severity, decides escalation, approves all customer communications, freezes changes, calls responders, hands off in shifts. Non-technical ICs are acceptable; technical depth is not required.
- **Deputy IC**: assigned within 15 min of declaration; shadows IC; takes over if IC unavailable or after 4-hour shift on SEV-1. Maximum IC shift: 4 hours on SEV-1, 6 hours on SEV-2.
- **Triage Owner** (new role): owns incident from first page acknowledgement until IC takes over or stand-down decision is made. Required for all incidents.
- **Communications Lead**: owns internal Slack updates and status-page messaging; shields IC from customer contact and interruptions.
- **Scribe**: records real-time timeline with decisions, actions, and key timestamps; not responsible for fixing.
- **Subject-Matter Responders**: engineers with service context; take IC direction; report only to IC; no side channels or parallel debugging.
- **Operations Lead** (SEV-1 only): coordinates multiple responders, manages incident bridge, maintains escalation list.
- Minimum viable staffing: SEV-1 requires all roles; SEV-2 requires IC, Deputy, Comms, Scribe, SMEs; SEV-3 requires Triage Owner and IC.
- Create laminated role cards for every on-call shift location (office, home, printed in pockets).
7. Three on-call rotations: Service, Platform, and Incident Commander (depends on: 5, 6)
Directly address the "carrying a pager for another team's code" objection by making it structurally impossible.
- **Service On-Call rotation** (federated): each of the 28 teams maintains a rotation for their own services only. No engineer is paged for code their team does not own. The answer to "why am I carrying a pager?" is now simply: "for your team's code."
- **Platform Duty rotation** (centrally staffed): shared PostgreSQL cluster, Kubernetes, networking, CI/CD, observability, and incident management tooling. Nobody's product code, so it gets its own dedicated rotation. Staffed from platform teams plus volunteers from other teams; paid at premium rate.
- **Incident Commander roster** (24x7): 12–16 certified senior engineers from across all 28 teams, on one-week primary shifts with secondary backup. Covers every hour with no single point of failure and no uncovered holiday week.
- **Consequences and gates**: 16 of 28 teams have no on-call today. Each must either (a) build a Service On-Call rotation of at least 6 engineers, or (b) formally transfer service ownership to a team that will, with transfer documented and dated. No exceptions, no waivers. Unowned services are decommissioned or transferred by end of month 2.
- Merge small or low-traffic teams into shared rotations where service ownership is unclear (e.g., shared analytics, testing infrastructure).
- Enforce scheduling limits in the tooling: no engineer on-call more than 2 weeks per quarter, automatically enforced by configuration, not negotiation.
- Publish a coverage matrix for all 28 teams showing services, owners, rotation size, gaps, and monthly status.
8. On-call compensation, rest policy, and sustainability (depends on: 7)
Settle compensation before rollout, not during negotiations. Make on-call sustainable and valued.
- **Paid on-call**: effective immediately upon joining a rotation. Weekly stipend while on shift (benchmark to New York market: $600–1,000 per week per engineer), regardless of incident volume.
- **Event-based compensation**: 1.5× hourly rate for time spent mitigating out-of-hours incidents, minimum one-hour block per callout. Tracked by incident record (auto-capture from timeline).
- **Compensatory rest**: no engineer works a normal 8-hour business day after a night incident requiring >2 hours mitigation. Rest day is documented policy, not a favour granted by manager.
- **Intrusion cap**: maximum 3 unscheduled pages per week per engineer. Exceed the cap in a week and trigger an immediate review; exceed in a month and escalate to VP Engineering. Breaches are structural signal that alert quality or service stability has a problem.
- **Voluntary opt-out**: an engineer may exit a rotation; their team must hire or buy replacement coverage from paid pool at published internal rate ($X per shift). This converts culture debate into visible budget decision.
- **Amnesty policy**: incident records, near-miss reports, and false declarations are never used in performance reviews or compensation discussion. Only failure to report is a performance issue.
- **Policy publication**: publish compensation structure and effective date before any team is asked to join a rotation, and include on-call expectations in job descriptions and hiring conversations.
- **Semi-annual review**: reassess compensation and caps every six months against actual page volumes, attrition rates, and survey feedback.
9. Detection strategy: SLOs, synthetic monitoring, and customer-report intake (depends on: 4, 7)
Close the 40% customer-detected gap by monitoring customer journeys instead of infrastructure metrics.
- **SLO-based alerting**: Define SLIs and SLOs for the top 20 customer journeys (payment initiation, authorization, settlement, ledger read/write, API availability, webhook delivery, payout). Measure per region. Alert on SLO breach, not on infrastructure metric (e.g., alert on "payment success rate <99%" not "database CPU >80%").
- **Synthetic transaction monitoring**: deploy synthetic transactions from outside AWS in both regions plus a third geographic location, one-minute cadence, for all money-moving paths. These are your first alarm bell.
- **Ledger-critical signals**: PostgreSQL replication lag (target: <1s, alert >5s), connection saturation, write latency (p95), lock-wait time, transaction ID exhaustion proximity, checkpoint pressure, table bloat. These are separate alerts on shared-database health.
- **Customer-report intake** (new detection channel): Support and Account Managers can raise an incident directly in the platform. Every customer report creates an incident record automatically, and the "customer report" detection source is counted in all metrics. This is a legitimate detection method, not a failure.
- **Detection-gap rule**: whenever a customer reports an incident before internal monitoring fires, auto-create a ticket in the owning service's backlog with root cause: "Monitoring gap on [journey]."
- **Detection contract per service**: every one of the 180 services needs a named owner, at least one symptom-based alert mapped to a SLO, and a documented expected detect time (target: <5 min for payment path, <10 min for others). Published on wiki and reviewed monthly.
- **Detection drills**: run a quarterly drill per team: simulate a broken service in staging and verify it triggers a page before a human notices.
- **Resilience roadmap separation**: detection improvements do not protect against ledger corruption or multi-region failure. Fund a separate resilience roadmap to reduce shared-database blast radius and improve failover safety.
10. Alert quality standards and noise-reduction program (depends on: 9)
Cut the 3,400 monthly alerts (85% noise) to <600 with 95% signal. This is the price of admission for on-call buy-in.
- **Paging contract**: every page must satisfy all of (1) symptom-based (customer impact, not infrastructure cause), (2) actionable (linked runbook with immediate next step), (3) owned (named team responsible), (4) severity-mapped (SEV-1/2/3/4), (5) SLO-linked where applicable. **No runbook, no page.** Enforce with CI check on alert definition.
- **Separation rule**: only customer-impacting or imminent-impact signals trigger pages. Everything else becomes a dashboard entry, a ticket, or a log line. Noisy infrastructure metrics go to dashboards, not pagers.
- **Page budget per service**: no service may exceed 2 pages per on-call shift per month. Exceeding budget auto-opens a remediation ticket in the owning team's backlog (with alert-quality review assigned to tech lead).
- **Automatic suppression rules**: (1) silence alerts if service auto-recovered within 30s, (2) suppress known maintenance windows, (3) group flapping alerts (>5 in 2 min) into one page, (4) rate-limit noisy services (max 1 page per 5 min until condition clears). All suppression rules must have an expiry date; no permanent silence without a ticket.
- **Probation for new alerts**: new alert rules run as tickets only and alert to a Slack channel; after two weeks of proving actionability (every alert resulted in human action), they graduate to pager.
- **Noise sprint**: run a focused 90-day program with a public burn-down of the top 100 noisiest rules. Assign each to a named manager. Default action: fix root cause, tune threshold, or delete within 10 working days. Deletion is a legitimate successful outcome (celebrate it).
- **Correlation and deduplication**: consolidate alert sources at ingest pipeline so one outage triggering 40 alerts produces one page, not 40.
- **Alert ownership**: every alert must have an owning team and a maintenance contact. Update monthly.
11. Incident tooling consolidation and integration (depends on: 5, 10)
Replace six alert tools and ad-hoc incident records with a single source of truth that unifies paging, escalation, timeline, and audit evidence.
- **Tool selection**: choose an incident-management platform (e.g., PagerDuty, Incident.io, Opsgenie) that integrates paging schedules, escalation policies, incident records, postmortem workflow, and status-page APIs. Decision gate: month 1.
- **Event pipeline consolidation**: route all alerts from the six legacy tools into a single event pipeline that feeds the incident platform. Apply deduplication, correlation, severity/class mapping, and rate-limiting at ingest.
- **Observability integration**: connect the incident platform to your Kubernetes dashboards, PostgreSQL monitoring, distributed tracing, and logs so responders see context in one pane. Link runbooks directly into incident records.
- **Slack and bridge integration**: auto-create incident Slack channels, auto-invite roles, auto-capture timeline from channel transcript and voice-bridge recording. Timeline is not written from memory; it is auto-captured.
- **Golden incident file**: implement the export defined in S3. One click produces a complete, immutable, audit-ready PDF: severity, timeline, roles, decisions, communications, postmortem, action items, and closure evidence.
- **Dual-run period**: run both legacy and new platform in parallel for two weeks. Define rollback criteria (e.g.,
12. Escalation automation and incident lifecycle enforcement (depends on: 5, 11)
Eliminate judgment calls from the worst moments. Escalation is automatic, mechanical, and blameless.
- **Automatic escalation ladders**: page responder → if no ack in 5 min, page service owner → if no ack in 10 min, page team manager → if no ack in 15 min, page IC on-call + call them immediately (phone + SMS + Slack). No human decides to escalate; timers fire escalations.
- **Severity-based escalation tempo**: SEV-1 uses faster timers (2 min for IC on-call), SEV-2 uses moderate timers (5–10 min), SEV-3 uses slower timers (15–30 min). Configured in tooling, reviewed quarterly.
- **Dual IC rule**: if a second SEV-1 incident is detected while the first is active, immediately page and assign a separate IC. ICs never run two incidents in parallel.
- **Change freeze and rollback authority**: SEV-1 and SEV-2 trigger automatic deploy freeze. Only the IC (with CTO/VP Eng notification) may unfreeze. Freeze lifts only when IC explicitly declares mitigation confirmed and verifies no new incident symptoms for 5 min.
- **Unresponsive team escalation**: if service's on-call does not acknowledge in 30 min, IC may direct any engineer from any team (volunteers first, then rotated) to engage. This is documented and reported in monthly review (escalation = signal of rotation problem).
- **One incident, one record**: all decisions logged in the incident platform. Auto-capture from Slack, bridge, status-page updates. Timeline is the source of truth; postmortem is written from timeline, never constructed after the fact.
- **Ambiguity rule**: if two responders disagree about whether an incident should be declared, it is declared. False declarations (stand-downs within 30 min of declaration) are tracked as metrics and closed without blame.
- **Watch state**: an unconfirmed incident can live in "Watch" state for max 30 min; after that, either declare it or stand it down explicitly.
13. Internal, customer, and regulatory communications workflows (depends on: 6, 12)
Define who informs whom, in what order, via what channel, with explicit timings and pre-approved templates.
- **Internal cadence**: first update to #incidents Slack channel within 3 min of declaration (even if "Investigating"). Then updates every 5 min (SEV-1), 15 min (SEV-2), or 30 min (SEV-3), or immediately on material change (e.g., mitigation achieved, scope widened). **Comms Lead owns the update; IC must not be interrupted.**
- **Executive notification**: IC calls CTO and VP Eng within 1 min of SEV-1 declaration (not email, not Slack, call). Incident declared in Slack with severity label, IC name, and affected service. Escalation channel lead auto-pinged.
- **Customer communication channels**: status page (all 2,100 customers), direct email to affected customers (top-tier accounts and customers affected by SEV-1), account-manager calls (top 50 accounts on SEV-1).
- **Status page timings**: update within 3 min of SEV-1 declaration, 10 min of SEV-2, 30 min of SEV-3 (even if root cause unknown; use "Investigating" with next-update ETA). Updates every 5–30 min depending on severity. Always include next-update time.
- **Pre-approved templates**: draft customer-facing language for each severity and class in advance with Legal and Compliance. Templates specify impact language ("some of your transactions are delayed" not "our database failed"), workarounds if available, and next-update commitment. Never speculate on cause in customer communication.
- **Regulatory notification path**: identify incidents requiring regulator notification (NYDFS Part 500, money-transmitter rules, payment-card-network rules, securities disclosure). Build a clock matrix: event type → regulator → notification window → signer. Compliance owns all regulatory notifications (never Engineering). Pre-clear templates. Flag incidents to Compliance immediately upon declaration.
- **Account manager playbook**: contact tree for top 50 accounts, templated talking points (facts only, never speculation), escalation path if customer escalates, what to offer (service credit, technical deep-dive call).
- **Closing communication**: resolution notice, SLA credit impact, commitment date for written root-cause report, customer action required (none, or security update, etc.).
14. Status page infrastructure and customer-impact ledger (depends on: 13)
Make the status page reliable, customer-centric, and audit-ready. Track customer impact in a single durable record.
- **Status page decoupling**: host status page outside production failure domain (separate cloud, separate infrastructure, separate database). Integrate incident platform with status page so incident record drives all public updates. Status page survives total platform outage.
- **Component-to-journey mapping**: status page components map to customer journeys ("Payments", "Settlements", "Payouts", "Ledger API") not to internal services. Allow customers to subscribe to components; notify by email or webhook.
- **One-click update templates**: pre-fill status-page template with severity, impact language, next-update time, and estimated resolution. Comms Lead types minimal new info ("Root cause identified" or "Workaround available"), and updates auto-post.
- **Customer-impact ledger** (one record per incident): which customer accounts affected, which journey(s) impacted, exact start and end time of impact, estimated SLA-credit exposure. Use this single record for customer communications, credit computation, regulatory reporting, and annual review. No reconciliation of two versions of the same outage.
- **SLA credit automation**: compute credit based on duration × severity × customer tier → auto-generate customer notification → auto-post to finance system. Reconcile accrued vs. paid credits monthly and report in executive review.
- **Testing during game days**: simulate status-page outage and verify alerts continue to fire; test total region loss and confirm status page remains updated; drill runbook for manually updating status page if platform is down.
15. Postmortem policy: mandatory, blameless, three-level framework (depends on: 6, 13)
Turn postmortems from a writing exercise (11 of 64 action items closed) into the learning engine of the system.
- **Mandatory postmortems**: all SEV-0, SEV-1, and SEV-2 incidents; all SEV-3 with customer impact or repeat pattern; any near-miss IC flags; any incident where the process itself failed (IC unreachable, Comms Lead unavailable, false declaration, missed update SLA).
- **Three-level framework** (proportionate to weight): (1) lightweight async review for SEV-4 and low-impact SEV-3 (10 min template in shared doc, owner + IC review), (2) standard facilitated postmortem for SEV-2 and impactful SEV-3 (full template, facilitated by trained neutral party, published within 10 days), (3) full executive postmortem for every SEV-1 and every security incident (executive sponsor assigned, full investigation, published within 15 days, customer-facing variant prepared).
- **Fixed timeline**: draft postmortem within 5 business days, blameless review within 10 days, internal publication within 15 days.
- **Single template**: impact (who, how many, how long, financial exposure), timeline (detection through resolution), root cause (not "human error" but system condition that enabled error; what was the gap?), contributing factors (tooling, process, organization, knowledge, monitoring), what went well, what went badly, action items (≤3, rest go to reliability backlog).
- **Blameless facilitation**: train a pool of blameless postmortem facilitators (target: 10+ engineers). Require a trained, neutral facilitator for every SEV-1 and SEV-2 review. Prohibit counterfactual language ("if the engineer had"), blame language, and the phrase "human error" as a root cause.
- **Publication rule**: publish all postmortems internally by default; security review only for genuinely sensitive material (e.g., unpatched vulnerability details or customer PII in logs). Create a customer-facing root-cause report for every SEV-1, especially for regulated customers, with legal and compliance sign-off.
- **Searchability**: store postmortems in a searchable wiki or issue tracker with tags (service, class, root cause category) so teams can learn from similar incidents without repeating them.
16. Action item tracking, reliability backlog, and repeat-incident design rule (depends on: 15)
Close the loop on incident learning by enforcing verifiable, tracked action items and breaking cycles of repeat incidents.
- **Action item capping**: each postmortem generates a maximum of 3 action items. Anything beyond 3 goes into a ranked reliability backlog, not into the postmortem, to prevent overwhelming teams.
- **Action item requirements**: each item must have (1) a named human owner (not a team), (2) a due date (≤60 days, target ≤30 days), (3) a definition of done (merged code change, tested alert, audit evidence, architectural decision, new runbook, training completed) not self-reported status.
- **Single reliability backlog**: create one backlog in your engineering tracker (Jira, Linear, etc.) with mandatory label (e.g., `incident-action`), link to originating incident, and link to postmortem. Track progress weekly.
- **Closure sign-off**: Incident Commander or postmortem facilitator must sign off on closure, verifying artifact exists (code merged, alert tested in drill, runbook verified).
- **Repeat-incident rule**: if the same service or component has a second incident with the same contributing factor, **do not create another action item**. Instead, escalate immediately to an architect or tech lead and trigger a design review (not a task, a review). This breaks the cycle of repeated patches; the system needs a structure change.
- **Capacity protection**: reserve a fixed percentage of each team's sprint capacity (10–15%) for reliability work. Track unspent capacity and report to VP Engineering monthly; if a team is not spending it, work with them to identify and fix blockers.
- **Ageing and escalation**: run a weekly review of open actions; escalate anything >30 days overdue to team lead and VP Engineering. Monthly report: completion rate (target >90% within 60 days) and median action age (target <30 days).
17. Training, certification, and exercise program (depends on: 6, 13, 15, 16)
Build skills before deploying the process. Run ongoing drills so the system is tested, not guessed at.
- **Curriculum**: (1) All engineers (30-min async video): severity taxonomy, communication expectations, postmortem format, when to declare an incident, where to find runbooks. (2) On-call responders (2-hr workshop): alert tool walkthrough, playbooks by severity, escalation paths and timers, when to call manager, mobile app walkthrough, case studies from the last 12 months. (3) Incident Commanders (4-hr bootcamp + test): leadership under pressure, decision-making (severity, escalation, rollback), communicating with executives, status-page discipline, postmortem facilitation, handling ambiguity, live simulated incident (pass/fail certification). (4) Communications Leads (2-hr training): templates per severity and class, customer-communication rules (no speculation, no blame), update timings, how to shield IC, regulatory triggers.
- **IC certification**: written assessment (75% pass required) plus live simulated incident (role-play with facilitator, graded on severity declaration, escalation decisions, communication, handover). Certification valid for 12 months; recertify via annual refresher or another live sim.
- **Depth across teams**: certify at least 2 ICs per team or team group so central roster is not siloed in one group; no holiday week is uncovered.
- **Async content**: record all training videos so async teams can learn on their schedule. Create quick-reference cards (laminated, pocket-sized) for roles and playbooks; distribute to on-call locations (office, home).
- **Monthly tabletop exercises**: drawn from real incidents from the last 12 months (region loss, ledger write failure, missed settlement window, cascading failures). Facilitator describes scenario; 3–4 responders play out response (Triage Owner, IC, Comms) as if real. Run 30 min; retro for 15 min afterward.
- **Quarterly game days**: deliberately inject failures into production (database failover, status-page outage, alerting-pipeline outage, dual SEV-1 incidents). All on-call roles engage. Run 2–3 hours; measure response times, decision quality, and communication. Document findings and create action items for identified gaps.
- **Drill the process's own failure modes**: IC unreachable (on-call unavailable, phone broken), Comms Lead on PTO, two simultaneous SEV-1s, paging storm (100+ alerts), false alarm that consumes an hour. Test escalation paths, deputy takeover, and recovery.
- **New-engineer onboarding**: add incident-response module to all engineering onboarding (completion tracked, audit-ready). All engineers must complete within 30 days of joining or transferring in.
18. Metrics, dashboards, and review cadence (depends on: 2, 12, 16, 17)
Measure to prove the system works. Publish live dashboards so every engineer sees the scoreboard and the system is transparent.
- **Outcome metrics**: Median Time to Detect by source (target: <5 min internally detected, <10 min customer journeys); Median Time to Mitigate for SEV-1/2 (target: <60 min SEV-1); customer-detected incidents as % of total (target: <5%); incidents by severity (should be mostly SEV-3/4, few SEV-1); SLA credits paid (target: <$100K/year by month 12); annual credit avoidance vs. program cost.
- **Process metrics**: IC assigned within 5 min (target: >95% of incidents); page acknowledgement rate (target: >98% within 5 min); first-update timeliness (target: >95% within SLA); postmortem timeliness (target: 100% of mandatory postmortems published on time); IC roster coverage (zero uncovered hours, monitored weekly).
- **Health metrics**: alert volume and signal-to-noise ratio per team (trending toward target); off-hours pages per engineer per month (trend, cap enforcement); on-call satisfaction survey (target: >7/10); training completion by team (target: 100% within 30 days); % of services with active detection contract (target: 100%).
- **Never publish incident count as a team metric.** Reward hiding. Instead publish detection metrics (near-misses reported per team, detection gaps closed, false declarations made).
- **Live dashboards**: build dashboards visible to all engineers (not just managers) showing outcome, process, and health metrics. Auto-populate from incident platform and alert tool. Update daily. Link from Slack and internal wiki.
- **Baseline all metrics against S2 evidence pack.** Set 90-day and 12-month targets for each metric. Publish targets and progress monthly.
- **Review cadence**: (1) weekly 30-min operational review (incident by incident from prior week: what went well, what hurt, actions); (2) monthly 60-min reliability review (trends, top causes, action-item aging, alert quality per team); (3) quarterly 60-min executive review (CEO's office: customer impact, SLA credits, top five systemic causes, program ROI).
- **Quarterly process review**: what in the process wasted responder time, what confused people, what should be deleted. Solicit feedback from ICs, Comms Leads, and responders. Document changes and reasoning.
19. Pilot program with 3–4 volunteer teams (depends on: 6, 8, 11, 12, 13, 14, 15, 16, 17, 18)
Do not roll out untested to 28 teams. Run the entire process end-to-end with a small cohort using real incidents as the primary training material.
- **Team selection**: recruit 3–4 volunteers spanning criticality: one payment-path team, one ledger-adjacent team, one shared infrastructure team (platform or Kubernetes), one low-traffic team. Volunteers see early adoption and influence.
- **Full process in pilot**: new severity and class taxonomy (S4), consolidated tooling (S11), roles and Triage Owner (S5–6), three rotations (S7), escalation automation (S12), communications (S13–14), postmortems (S15), action tracking (S16), paid on-call (S8), training (S17), metrics (S18). This is not a partial test; it is the complete system.
- **Real incidents are the training**: hold a retro within 48 hours of each pilot incident (while memory is fresh). Process Owner facilitates. Discuss: what worked, what hurt, how is the runbook, is the alert tuned, did Comms template work, did roles work, was timeline auto-captured correctly. Document feedback and iterate.
- **Instrument against baseline**: measure all S18 metrics in pilot; compare results with S2 baseline. If MTTD dropped from 22 min to 8 min, that is the number you carry into the next conversation.
- **Weekly retros with pilot teams**: structured 30-min sessions reviewing policies, tooling, training, templates, and process changes. Change the written policy or tooling in response; document what changed and why so later teams see the evolution, not just the finished state.
- **Explicit exit criteria**: (1) rotation coverage sustained (≥6 engineers per rotation, zero unacknowledged pages over 3 weeks), (2) postmortems delivered on time (100% of mandatory postmortems published within 15 days), (3) action tracking established (100% of action items in backlog with owner and date), (4) metrics live (dashboards updated daily, first weekly review completed).
- **Pilot report**: document before/after numbers (MTTD, MTTR, alert noise, action completion rate, on-call satisfaction) and key process learnings. This report is the foundation for every conversation in the rollout.
20. Phased rollout sequenced by cost of failure (depends on: 16, 18, 19)
Rollout is a staged migration with readiness gates and sequencing that prioritizes visible impact, not ease.
- **Wave sequencing**: divide 28 teams into 4 waves of ~7 teams each, **ordered by incident density and customer-journey ownership** (highest-cost-of-failure teams first). Teams with the most SLA credits at stake go first; their improvement is the proof.
- **Wave spacing**: three weeks between waves. This gives each wave time to stabilize and find problems before the next cohort joins.
- **Readiness checklist per team**: (1) all services mapped and owned (no unowned services), (2) alerts cleaned to paging contract (runbook linked, severity mapped), (3) playbooks updated and tested in staging, (4) rotation staffed to ≥6 engineers, (5) team completes training module, (6) manager briefed on policy, (7) on-call compensation in effect.
- **Gate review before each wave**: Process Owner holds gate review with target teams. Move unready teams to next wave with a dated remediation plan. No exceptions, no waivers; readiness is non-negotiable.
- **Wave champion**: assign a named engineer per wave to champion the rollout, answer questions, escalate issues to Process Owner. Champions are not representatives; they are advocates and feedback collectors.
- **Communication cadence**: weekly all-hands or newsletter for 4 weeks before each wave. Explain why (owned-code-owned-pager rule, paid on-call, SLA credit savings). Use pilot numbers. Answer FAQs. Announce champion and escalation path.
- **First incident under new process**: hold a retro within 48 hours. Feed accepted process changes back through change control.
- **Retire legacy tools and processes**: at end of each wave, retire legacy alert tools, informal escalation lists, ad-hoc status-page process. No parallel processes running for >3 weeks; this prevents confusion and half-learning.
- **Sequence to avoid audit collision**: ensure no team is rolling out in the same week as the audit dry run (S21).
21. SOC 2 dry run and evidence review (depends on: 3, 20)
Convert a good working process into a provable one. Test control evidence a few months before auditors arrive, when you can still fix gaps.
- **Dry run timing**: run 6 weeks before audit window (around month 7 of this program).
- **Scope**: sample 10–15 real incidents from pilot and early rollout waves. For each incident, verify evidence artifact exists and is complete: incident record, timeline (auto-captured), severity and class declaration, roles assigned and logged, communications log (Slack + status page), postmortem (if mandatory), action items in tracker with due dates, action completion evidence (code, alert test, training record).
- **Control walkthrough**: walk through each control statement from S3 with a checklist. Is the evidence artifact present? Is it immutable? Is it searchable? Is retention adequate? Is access logged?
- **Gap remediation**: for every gap found, estimate time to fix and prioritize by audit risk. Anything risking a qualified opinion (e.g., missing postmortem, no timeline evidence) must be fixed before the audit. Test the remediation against a new incident or a resample.
- **Interview readiness**: brief 10–15 engineers who may be interviewed by auditors (ICs, Comms Leads, Process Owner, team managers). Ask them to describe the process as they actually practice it, not as written. Listen for confusion or gaps in understanding. Correct them.
- **Auditor package preparation**: assemble process documentation, sample incident records (5–10 complete golden files), training records, on-call schedules, alert quality metrics, action tracker register, and status-page archive. Organize by control. Create a table of contents and index.
- **Single audit liaison**: designate Process Owner or a small dedicated compliance person as sole point of contact for audit requests. Prevents requests scattering across 28 teams.
- **Rehearsal**: conduct mock interview with an IC and a Comms Lead. Auditors ask tough questions under pressure ("How do you know the timeline is accurate?", "What happens when both ICs are unavailable?", "Show me how you proved the alert was actionable."). Practice answering.
22. Standing governance and continuous improvement (depends on: 20, 21)
Lock in durable improvement. The classic post-audit failure is the process freezing and then decaying. This step prevents that.
- **Standing Incident Management Council**: chaired by Process Owner, monthly meetings, attendees: CTO, VP Engineering, Head of Support, Head of Compliance, one engineering manager per region, one IC, one team manager from recent wave. Agenda: metrics review, policy changes, gaps from recent incidents, escalation for contentious issues.
- **Change control mandate**: give Process Owner documented authority to change severity taxonomy, response classes, roles, communications timings, and compensation policy. Any change requires: written justification, steering group approval (monthly cadence), and documented effective date before implementation. This prevents silent drift and ensures changes are deliberate.
- **Quarterly validation**: re-validate severity and response class taxonomy against all incidents from the prior quarter. Ask: did our taxonomy correctly predict response posture? Did we misclassify? Update taxonomy if patterns emerge.
- **Annual metric re-baselining**: every 12 months, re-run measurements from S2 (alert census, incident register) to reset targets. System should improve; targets should tighten.
- **Resilience roadmap separation**: fund a distinct architectural or platform roadmap for incident prevention (reduce shared-database blast radius, multi-region failover, deploy safety, observability investments). Better incident response does not protect a single-ledger corruption or unplanned failover. These are separate problems.
- **Quarterly executive report**: CTO and VP Eng report to CEO/CFO on metric set (detection time, mitigation time, SLA credits, customer-detected %), top five systemic causes of incidents, program cost vs. credit avoidance, and strategic architecture changes in flight.
- **Public backlog of improvement ideas**: teams and engineers propose process improvements via Slack or wiki. Process Owner reviews quarterly and implements accepted ideas (e.g., "add a dashboard for detection gaps", "update postmortem template"). Publish what changed and why.
- **Celebration and learning**: share wins publicly each quarter ("We reduced MTTD from 22 min to 5 min", "Customer-detected incidents down 80%", "$800K SLA credits avoided"). Refresh training and tabletop program annually and immediately after any SEV-1 to keep the system sharp and responsive to new scenarios.
--- PROPOSAL 2 ---
Proposal ID: abac560b-e42c-412b-9ecc-a60d55c29045
Content:
Estimated Complexity: high
Success Metrics: - Median time to detect falls from 22 minutes to under 8 minutes by month 6, and under 5 minutes for money-moving journeys by month 9.
- Share of customer-impacting incidents first detected by customers falls from 40% to below 10% by month 9.
- Median time to mitigate for SEV1 falls from 3 h 10 min to under 60 minutes by month 9.
- 100% of SEV1 and SEV2 incidents have a named Triage Owner within 5 minutes and a named Incident Commander within 15 minutes, from month 4.
- Zero incidents with unclear command authority lasting more than 15 minutes, measured monthly from month 4.
- Monthly page volume falls from 3,400 to under 600, with a false-positive rate below 15%, by month 5.
- No service exceeds two pages per on-call shift for three consecutive months by month 6.
- 100% of the 180 services have a named owner, a detection contract and at least one symptom-based alert by month 7.
- All 28 teams have a documented rotation, a service ownership map and at least one trained on-call engineer by month 6.
- 100% of on-call shifts are paid under a published policy from month 2, with zero on-call-attributed voluntary attrition by month 6.
- On-call satisfaction scores 7 out of 10 or better in quarterly surveys from month 6.
- The IC roster holds at least 12 certified ICs covering 24x7 with no uncovered week from month 4.
- 100% of SEV0, SEV1 and SEV2 postmortems are published internally within 15 business days from month 5.
- Action items closed within 60 days rise from 17% to above 90%, with a median action age under 30 days, by month 6.
- Status page first update is posted within 30 minutes on at least 95% of SEV1 incidents from month 4.
- Zero missed regulatory notification windows on any incident requiring notification.
- The customer-impact record is complete for 100% of customer-impacting incidents from month 5, with SLA credits reconciled monthly.
- Annual SLA credits paid fall at least 50% from the $1.3M baseline within 12 months.
- Customer-impacting incidents fall at least 40% year over year from the 31-incident baseline.
- Review cadence is sustained: weekly operational review in at least 90% of weeks, 12 of 12 monthly reliability reviews, 4 of 4 quarterly executive reviews.
- Every page has a recorded disposition — fixed, tuned or deleted — within 10 working days, from month 4.
- SOC 2 Type II is passed with zero findings related to incident response.
- 100% of new engineers complete the incident-response onboarding module within 30 days of joining.
Steps (24):
1. Mandate, one owner, and the evidence clock
This step turns the CEO's email into a funded programme with a single accountable owner, and it starts the SOC 2 clock on day one.
- Appoint a Director of Incident Management reporting to the CTO, with a dotted line to the COO for customer and SLA matters and to the Head of Compliance for audit readiness.
- Publish a one-page charter: scope covers every customer-impacting, money-moving, security and data-integrity event across all 28 teams and both regions.
- Grant explicit authority to declare an incident, set severity, freeze deploys, page any engineer in the company, and approve customer messaging.
- Fund the envelope up front: tooling, training and drill time, on-call compensation, and a small programme team, roughly $500–700K a year against $1.3M in credits paid.
- State the return plainly to the steering group: credits avoided, churn avoided, and audit findings avoided.
**Start the evidence clock now.** A SOC 2 Type II report covers a period of operating effectiveness, so evidence cannot be backfilled; an audit in eight months means the operating period begins in the first weeks of this programme.
- Stand up a fortnightly steering group of the CTO, VP Engineering, Head of Support, Head of Compliance and two engineering managers.
- Open a programme risk register with the top risks, owners and review dates, and revisit it at every steering group.
- Make participation in the incident process a documented performance expectation for every engineering manager, not an optional extra.
2. Baseline evidence pack and cost-of-downtime model (depends on: 1)
You cannot demonstrate improvement without a defensible starting point, and you cannot win the argument about noise without numbers.
- Build a 12-month incident register and re-classify all 31 customer-impacting incidents by detection source, duration, customers affected and credits paid.
- Build the silent-failure register: incidents where no internal alert fired at all, which is the number that explains the 40% customer-detected rate.
- Run an alert census per tool, per team and per service: volume, page-to-action ratio, off-hours interruptions, and the 50 noisiest rules with a named owner each.
- Survey every on-call engineer and manager on burden, fairness, escalation clarity and pay expectations, targeting above 70% response.
- Interview Support, Account Management and Sales about how customers learn of incidents and what they are promised.
- Reconstruct the two command-ambiguity incidents minute by minute to find exactly where ownership lapsed.
- Build a cost-of-downtime model: dollars per minute of impact per customer journey, used later to sequence teams and justify funding.
- Publish the pack internally as the problem statement, and retain every artifact as management-review evidence for the audit.
3. SOC 2 control mapping and evidence architecture (depends on: 1)
Most programmes leave compliance to the end; this one maps controls in month one, because the mapping decides what the process must capture from day one.
- Map the process to the Trust Services Criteria for incident identification, evaluation, response, containment, recovery and communication (CC7.1–CC7.5, CC2.2, CC2.3, CC4.1, CC3.x).
- Write each control in plain language, with one named owner and its evidence artifact.
- Define the golden incident file: one auto-assembled, single-click export per incident containing timeline, roles, severity, communications, postmortem, actions and closure proof.
- Set retention, storage location and immutability so no control depends on a laptop, a private channel, or chat history that expires.
- Keep a gap register with owners and dates, reviewed fortnightly by the steering group.
- Meet the auditor's readiness team inside the first 90 days to test the control design before anything is built on top of it.
4. Severity and class taxonomy with the trigger matrix (depends on: 2)
Severity decides how loud the response is; class decides which playbook runs and what must never be traded away.
- Define severity in customer and money terms: customers affected, dollars at risk, regions lost, ledger integrity, settlement delay, data loss and SLA exposure.
- Use SEV0 for security, privacy or regulatory events; SEV1 for total or material loss of a payment path; SEV2 for degradation or single-region loss; SEV3 for limited impact with a workaround; SEV4 for internal-only issues and near-misses.
- Define response classes: Availability, Performance, Data integrity and Ledger, Security and Privacy, Third-party dependency, Settlement and Financial, Process failure.
**Class can raise the response but never lower it.** A SEV2 data-integrity incident gets SEV1 posture, because integrity failures are not recoverable by moving faster.
- Specify automatic triggers: loss of one AWS region, ledger write failures, replication lag above threshold, a missed settlement cut-off, payment success rate below threshold for five minutes.
- State who may declare — any engineer, Support agent or account manager — and who may downgrade: the Incident Commander alone.
- Map every level to its SLA credit exposure and to the customer-visible status page state.
- Include worked examples from the last 12 months so each of the 28 teams recognises its own incidents, and revalidate the taxonomy quarterly against real declarations.
5. Day-one operating rules and the minimum viable process (depends on: 1, 4)
The full process will take months; the first useful version must be live in two weeks using the tools that already exist.
- Publish ten day-one rules that need no procurement: a named owner within five minutes, one channel per incident, one register entry per incident, one person speaking to customers.
- Apply the ambiguity rule: if two responders disagree about whether this is an incident, it is an incident.
- Make declaring free: a false alarm is closed as a false declaration, tracked as a metric, and never criticised.
- Require a register entry within 24 hours for every customer-impacting incident, even a minimal one.
- Ban silent incidents: if we know, the customer hears it from us rather than from their own reconciliation.
- Run the first 30 days on manual command, with a rotating duty Incident Commander drawn from the 12 teams that already have on-call.
- Hold a 15-minute daily incident stand-up during month one to catch friction while it is still fresh.
6. Roles, command structure and the no-unowned-minute rule (depends on: 4)
The two hour-long command failures did not happen at declaration; they happened in the gap before it, when an alert had fired and nobody owned it.
- Publish one-page role cards: Triage Owner, Incident Commander, Deputy IC, Operations Lead, Internal Comms Lead, Customer Comms Lead, Scribe, Subject-Matter Responders, and Executive Sponsor for SEV1 only.
- Introduce the Triage Owner rule: whoever acknowledges the page owns the incident until an IC takes over or the incident is stood down.
**The IC owns the incident, not the fix, and does not debug.** An IC who starts troubleshooting has abandoned command.
- Give the IC explicit decision rights: declare and escalate severity, freeze changes, halt deploys, approve customer messaging, and pull in any engineer.
- Define minimum viable staffing per severity: SEV1 fills every role; SEV2 staffs IC, scribe, comms and responders; SEV3 staffs an IC and a scribe.
- Set handover discipline: four-hour maximum IC shifts on SEV1 with a written handover, and a deputy named within 15 minutes of declaration.
- Set responder behaviour: one channel, one bridge, no side channels, and every request phrased with a named owner and a time.
- Link the role cards from every paging notification so they are one tap away at 3 AM.
7. Lifecycle, declaration and escalation policy (depends on: 4, 6)
This step defines the mechanical path from an alert to a declared incident and back to normal service, removing judgment calls from the worst moments.
- Define states with entry and exit criteria: Detected, Triaged, Declared, Mitigated, Resolved, Postmortem, Closed, plus a Watch state with a hard 30-minute timer.
- Set targets: page acknowledged within five minutes, triage decision within fifteen, declaration within thirty for anything customer-facing.
- Build automatic escalation ladders with timers at every layer: responder, service owner, team manager, IC on-call, VP Engineering.
- Define the unresponsive-team path: fifteen minutes escalates to the team manager, thirty minutes to the IC, who may then direct any available engineer.
- Define change freeze and rollback authority during SEV1 and SEV2, with the single named condition that lifts the freeze.
- Enforce one incident, one record, with the timeline captured automatically from the channel and bridge rather than written from memory afterwards.
- Add the two-simultaneous-incidents rule: a second IC is designated, and no IC ever runs two incidents at once.
- Test every escalation path weekly with synthetic pages, and adjust the timings after the first month of real operation.
8. On-call architecture across 28 teams (depends on: 4, 6)
The objection is that engineers will not carry a pager for another team's code; the answer is to build rotations that make the objection structurally impossible.
- Run a Service On-Call rotation per team, covering only that team's own services.
- Run a Platform Duty rotation for genuinely shared infrastructure: the PostgreSQL ledger cluster, Kubernetes, networking, CI/CD and observability.
- Run a central Incident Commander roster of 12–16 certified senior engineers on one-week shifts with a primary and a secondary.
**State the consequence honestly.** Sixteen of 28 teams have no rotation today; each must build one or formally transfer service ownership to a team that has one, with the transfer dated and recorded.
- Merge small or low-traffic teams into shared rotations so no rotation has fewer than six engineers.
- Cap load in the scheduling tool: no engineer on call more than two weeks per quarter, enforced by configuration rather than negotiation.
- Publish a coverage matrix of all 28 teams showing services, owners, rotation size and gaps, reviewed monthly.
- Write on-call participation into engineering job levels and hiring criteria, so the commitment is set before someone joins.
9. Compensation, rest and the economics of opting out (depends on: 8)
Unpaid on-call is the most cited reason for resistance, so settle compensation before rollout, not during it.
- Move to paid on-call: a per-shift stipend benchmarked to the New York market, agreed with HR and Finance, published with an effective date before any team is asked to join a rotation.
- Pay for callouts at 1.5× the hourly rate for time actually spent mitigating, with a minimum block per interruption.
- Provide documented compensatory rest: no normal working day after a night incident, and the rest day is policy rather than a favour granted by a manager.
- Cap intrusion with a maximum number of off-hours pages per shift, and trigger a mandatory review whenever the cap is exceeded.
**Allow opt-out, but put a price on it.** An engineer may step out of a rotation, and their team buys coverage from the paid pool at a published internal rate, which turns a cultural argument into a visible budget decision.
- Publish amnesty: incident records, near-misses and false declarations are never used in performance reviews; only failure to report is a performance issue.
- Protect against gaming: pay attaches to shifts and callouts, never to incident counts.
- Check the New York labour, overtime and tax treatment with Legal and Finance before announcing, and review the policy every six months against real page volumes, attrition and survey results.
10. Detection strategy: journeys, synthetic signals and customer-report intake (depends on: 4)
Customers detected 40% of incidents first, which makes detection the highest-leverage business problem in this programme.
- Define SLIs and SLOs for the top 20 customer journeys, measured per region: payment initiation, authorisation, settlement, ledger read and write, API availability, webhook delivery and payout.
- Alert on symptoms against those SLOs, not on cause-based infrastructure metrics.
- Add synthetic transaction monitoring from outside AWS in both regions plus a third location, on a one-minute cadence for money-moving paths.
- Add ledger and shared-PostgreSQL signals: replication lag, connection saturation, write latency, lock waits, transaction ID exhaustion, checkpoint pressure and bloat.
- Open a customer-report intake so Support and account managers can raise an incident directly, and count that path as a detection source in all reporting.
- Apply the detection-gap rule: whenever a customer reports first, a detection-gap ticket is opened automatically and owned by the responsible team.
- Publish a detection contract per service — owner, at least one symptom alert, documented expected detect time — for all 180 services.
- Run a detection drill per team: break something in staging and see whether it pages before a human notices.
11. Paging contract and the noise-reduction programme (depends on: 2, 10)
3,400 alerts a month at 85% noise is the reason engineers resent the pager, and fixing it is the price of admission for everything else in this plan.
- Publish the paging contract: every page must be symptom-based, actionable, owned, mapped to a severity and class, and linked to a runbook.
**No runbook, no page**, enforced by a CI check on the alert definition itself.
- Separate paging from ticketing: only customer-impacting or imminent-impact signals may wake a human.
- Set a page budget per team and per service, with a remediation ticket opened automatically, owned by the engineering manager, when the budget is breached.
- Put new alerts on two-week probation as ticket-only until they have proved actionable.
- Define the default action for a noisy alert: fix, tune or delete within ten working days, with deletion treated as a legitimate and celebrated outcome and the deleted count published.
- Deduplicate and correlate at the ingest pipeline so one root cause produces one page instead of forty.
- Require an expiry date on every silencing rule and temporary threshold, so suppression cannot quietly become permanent blindness.
- Run a 90-day noise sprint with a public burn-down of the top 100 noisiest rules, each owned by a named manager.
12. Incident tooling consolidation and the golden incident file (depends on: 3, 6, 10, 11)
Six alerting tools and no single incident record are structural causes of the 22-minute detection and the three-hour mitigation.
- Choose one incident platform for paging, schedules, escalation policies, incident records, postmortem workflow and status-page integration, and choose it within four weeks.
- Consolidate the six alert sources into a single event pipeline feeding that platform, with deduplication, correlation and severity and class mapping applied at ingest.
- Integrate observability, runbooks, chat and voice bridges so responders see dashboards and runbooks inside the incident record, and the timeline is captured automatically.
- Implement the golden incident file export defined in S3, so audit evidence is one click rather than a reconstruction.
- Run a dual-run period alongside the legacy tools with a defined rollback, then switch the legacy tools off on a published cutover date.
- Host the status page outside the production failure domain so it survives a total platform outage, and prove that in a game day.
- Make the platform usable from a phone: an Incident Commander must be able to run a SEV1 from a mobile device at 3 AM.
13. Internal, customer and regulator communications (depends on: 6, 7, 12)
Today the status page is written by whoever is around; this step replaces improvisation with a clock, a named owner and pre-cleared templates.
- Set the internal cadence: first update within 15 minutes of declaration, then every 30 minutes for SEV1 and hourly for SEV2, whether or not there is progress.
- Never let an employee learn of an incident from the status page: internal communication leads, external follows.
- Set the customer cadence: status page updated within 30 minutes of a SEV1 and 60 minutes of a SEV2, with a no-new-information update still mandatory.
- Define the channel hierarchy: status page for all 2,100 customers, direct email to affected customers on SEV1, and named account-manager calls for the top 50 accounts.
- Pre-approve templates per severity and class with Legal and Compliance, each carrying its own next-update time.
- Forbid speculation: customer messages never guess at cause, never assign blame, and never commit to a root cause before the postmortem.
- Define the closing communication: resolution notice, how SLA credits will be handled, and the committed date for the written report.
- Build a regulator clock matrix covering event type, regulator, notification window, signer and the shortest applicable clock, including NYDFS Part 500, money-transmitter and banking notifications, breach notification, card-network rules and public-company disclosure.
**The regulatory clock starts at awareness, not at root cause.** Route every notification through Compliance, never Engineering, and pre-clear the templates.
14. Customer trust workstream and the SLA credit ledger (depends on: 13)
The $1.3M in credits is a symptom of having no single record of customer impact, and the CEO's inbox is a symptom of customers learning things late.
- Maintain one durable customer-impact record per incident: which customers, which journeys, from when to when, and the estimated credit.
- Use that one record for communications, credit computation, regulatory reporting and the annual review, so nobody reconciles two versions of the same outage.
- Automate the credit path end to end — computation, approval, customer notification and finance treatment — so credits stop being a manual scramble three weeks later.
- Reconcile accrued against paid credits monthly and report the result in the executive review.
- Give the status page a named product owner and map its components to customer journeys, not to internal services.
- Send a CTO-signed reliability note to the top 50 accounts and publish a quarterly reliability report to all customers.
- Give account managers a script of the facts they may state, the speculation they may not, and a path for customer escalations.
- Track credit avoidance against programme cost, so the funding case stays a number rather than an argument.
15. Postmortems: mandatory set, three levels, blameless by design (depends on: 6)
Postmortems currently happen for some incidents, in various formats; this step makes them mandatory where they matter and light where they do not.
- Make postmortems mandatory for every SEV0, SEV1 and SEV2, every SEV3 with customer impact or a repeat pattern, every near-miss the IC flags, and every incident where the process itself failed.
- Use three levels so the ritual matches the weight: a lightweight async review for SEV3 and SEV4, a facilitated postmortem for SEV2, and a full review with an executive sponsor for SEV1.
- Set deadlines: draft within five business days, blameless review within ten, internal publication within fifteen.
- Adopt one template: impact, timeline, detection, response, contributing factors across tooling, process, organisation and human factors, what went well, what went badly, and actions.
- Train a pool of blameless facilitators and require a trained one for every SEV1 review, never the IC.
**Ban blame language in the template and ban "human error" as a root cause.** The question is always what system condition made the error possible.
- Publish all postmortems internally by default, with security review only for genuinely sensitive material.
- Produce a customer-facing root-cause report for SEV1 incidents, especially for regulated and top-tier accounts.
16. Action items: capped, verifiable, with the repeat-incident rule (depends on: 15)
Eleven of 64 action items closed is not a tracking problem; it is a generation problem, because the process produces more actions than the organisation can absorb.
- Cap each postmortem at three action items, with anything beyond that going into a ranked reliability backlog.
- Require every action to have a named human owner, a due date, and a definition of done that is an artifact: a merged change, an alert that fired in a drill, or a test that fails without the fix.
- Prohibit self-reported closure; closure requires the artifact and sign-off by the process owner or the Incident Commander.
- Keep one reliability backlog in the engineering tracker with a mandatory label and a link back to the originating incident.
- Reserve a fixed share of each team's sprint capacity for reliability work, with unspent capacity visible to the vice presidents.
- Apply the repeat-incident rule: a second incident with the same contributing factor triggers a design review owned by an architect, not another action item.
- Run a weekly ageing review of open actions and escalate anything more than 30 days overdue to the VP of Engineering.
- Report action completion rate and median action age monthly, by team.
17. Metrics, dashboards and the review cadence (depends on: 2, 4, 16)
Define what good looks like, then measure it in a way that rewards reporting incidents rather than hiding them.
- Define outcome metrics: time to detect by source, time to mitigate by severity and class, share of incidents detected by customers, incidents by severity and class, SLA credits paid, and customer-impact minutes.
- Define process metrics, because the process itself has SLOs: IC assigned within five minutes, page acknowledgement rate, first-update timeliness, update-cadence adherence, postmortem timeliness, action closure rate, action median age, and IC roster coverage.
- Define health metrics: alert volume and noise ratio per team, off-hours pages per engineer, on-call satisfaction, training completion, and share of services with a detection contract.
**Never publish incident count as a team metric.** It rewards hiding incidents; publish reporting metrics instead — near-misses filed, detection gaps closed, false declarations made.
- Publish live dashboards visible to every engineer, refreshed daily, with every metric baselined against the S2 evidence pack and given a 90-day and a 12-month target.
- Run a 30-minute weekly operational review incident by incident, a monthly reliability review of trends and actions, and a quarterly executive review with the CEO's office.
- Hold a quarterly review of the process itself: what wasted time, what confused responders, and what should be deleted.
- End every review with decisions and named owners, never with numbers alone.
18. Training, certification and the drill programme (depends on: 6, 7, 13, 15)
A process that lives only on a wiki page fails on the first real page, so skills are built and tested before they are needed.
- Build a practical curriculum: how to be on call, how to triage and declare, how to run an incident as IC, how to communicate, and how to write a postmortem.
- Require certification before joining the IC roster: a written assessment plus a live simulated incident.
- Certify at least two ICs per team group so the central roster has depth across all 28 teams and no holiday week is left uncovered.
- Train communications leads separately on templates, cadences, customer language and the regulatory rules.
- Run monthly tabletops drawn from the last 12 months of real incidents, including region loss, ledger write failure and a missed settlement cut-off.
- Run quarterly game days with deliberately injected failure, including PostgreSQL failover, status-page outage and alerting-pipeline outage.
- Drill the process's own failure modes, not just technical ones: IC unreachable, comms lead on leave, two simultaneous SEV1s, a paging storm, and a false alarm that burns an hour.
- Audit the process for single points of failure: who alone can perform each critical task, and what happens in their holiday week.
- Keep a mandatory onboarding module for every engineer joining or transferring in, with audit-ready completion records.
19. Pilot with three to four teams, using real incidents (depends on: 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 18)
Do not roll out to 28 teams untested, and do not rely on drills alone when real incidents are better training material.
- Recruit three to four volunteer teams spanning criticality: a payment-path team, a ledger-adjacent team, the platform team that owns the shared PostgreSQL cluster, and a low-traffic team.
- Run the entire process end to end in the pilot: severity and class, roles, escalation, communications, postmortems, action tracking, and paid on-call.
- Treat real incidents during the pilot as the primary training material, and hold a retro within 48 hours of each one, run by the process owner while the friction is fresh.
- Instrument the pilot against the baseline and publish before-and-after numbers.
- Hold weekly retrospectives with the pilot teams, and change the written policies, the tooling and the training in response, documenting what changed and why.
- Set explicit exit criteria: rotation coverage achieved, zero unacknowledged pages over two weeks, postmortems delivered on time, and actions tracked to closure.
- Produce a pilot report that every later rollout conversation starts from.
20. Phased rollout to 28 teams, sequenced by cost of failure (depends on: 16, 18, 19)
Rollout is a staged migration with readiness gates, not an email announcement, and the sequencing matters more than the schedule.
- Split the 28 teams into four waves of roughly seven, ordered by incident density and customer-journey ownership: the teams whose incidents cost the most SLA credits go first, because their improvement is the visible proof.
- Leave three weeks between waves and define a per-team readiness checklist: services mapped and owned, alerts cleaned to the paging contract, runbooks written, rotation staffed to at least six engineers, training complete, manager briefed, compensation in effect.
- Hold a gate review with the process owner before each team joins, and move an unready team to the next wave with a dated remediation plan rather than granting an exception.
- Give each wave a named champion and explain the why using the pilot's numbers.
- Handle resistance directly and early: publish the own-your-code-own-your-pager rule and the paid on-call mechanics before each wave, never after.
- Hold a retro within 48 hours of each wave's first incident under the new process, and push accepted changes through change control.
- Retire legacy tools, informal escalation lists and the ad-hoc status page process at the end of each wave on a published cutover date, so two processes never run in parallel for long.
- Sequence the waves so no team faces both a rollout and the audit dry run in the same week.
21. Audit dry run and evidence review (depends on: 3, 20)
This step converts a good process into a provable one, about six weeks before the auditors arrive.
- Sample real incidents from the pilot and the early waves against each control's evidence requirements.
- Remediate every gap found and re-test the remediated control against the same sample, prioritising anything that risks a qualified opinion.
- Brief every engineer who may be interviewed so they can describe the process as they actually practise it, not as it is written.
- Prepare the auditor package: process documentation, sample incident records, paging logs, communication logs, training records, on-call schedules, the action register and alert quality metrics.
- Designate one audit liaison and a small evidence-request team, so requests do not land on all 28 teams at once.
- Rehearse the walkthrough with an Incident Commander and a communications lead, because auditors probe realism under pressure.
- Keep the audit liaison and the process owner as close to the same person as possible, so accountability for the control is also accountability for the evidence.
22. Standing governance and process ownership (depends on: 20, 21)
The classic post-audit failure is that the process freezes and then decays, so ownership has to outlive the programme.
- Establish a standing Incident Management Council chaired by the process owner, meeting monthly with engineering, support, compliance and product representation.
- Give the process owner a documented mandate to change standards, with formal change control for any change to severity, classes, roles, communication timings or compensation.
- Revalidate the severity and class taxonomy quarterly against real declarations, and re-baseline the entire metric set annually.
- Report quarterly to the executive team on the metric set and the top five systemic causes of incidents.
- Feed incident themes into architecture review, release engineering and hiring, so the programme improves the system and not only the response.
- Keep a public backlog of process improvement ideas so any engineer can propose a change, and act on it every quarter.
23. Resilience track for the shared ledger and multi-region blast radius (depends on: 1, 17)
Better incident management cannot save a single shared ledger during a corruption event, so this runs as a funded engineering track beside the process, not inside it.
- Reduce ledger blast radius: read replicas for read paths, connection-pool isolation per service domain, statement timeouts, write throttles and per-tenant limits.
- Build a ledger protection mode: an automatic tripwire that degrades to read-only or queues writes before corruption can spread.
- Rehearse multi-region failover with the ledger as the binding constraint, and set a measured recovery time objective per region.
- Test restore and point-in-time recovery on a schedule, and publish the timings.
- Tighten deploy safety: progressive rollout, automatic rollback on SLO burn, and change-freeze windows around settlement.
- Track a small number of resilience bets with named owners and quarterly milestones, reviewed in the same executive review as the incident metrics.
24. Culture, incentives and continuous improvement (depends on: 20, 22)
Incident management is a cultural change before it is a process change, and the incentives decide whether reporting goes up or down.
- Restate the amnesty at every rollout wave: reporting, near-misses and false declarations are safe; hiding is not.
- Recognise detection improvements, deleted alerts, closed detection gaps and well-written postmortems publicly, and reward the engineers who do them.
- Correct blame language in leadership communication on the spot, including inside the executive review.
- Watch for pager fatigue with real data, and take an engineer off rotation as soon as the intrusion cap breaks.
- Run a quarterly process retro with the whole engineering organisation and publish what changed as a result.
- Refresh the training and tabletop programme annually and immediately after any SEV1.
- Share wins publicly each quarter — faster detection, fewer customer-detected incidents, credits avoided — to hold momentum after the audit.
--- PROPOSAL 3 ---
Proposal ID: 265b9bc8-32de-46d3-8546-6dd406da4c24
Content:
Estimated Complexity: high
Success Metrics: - Median Time to Detect (MTTD) < 5 minutes.
- Median Time to Mitigate (MTTR) < 60 minutes for SEV1.
- >90% of incidents detected internally (vs. 40% previously).
- SLA credits reduced to <$100K annually.
- Monthly alert volume < 600 with >95% signal-to-noise ratio.
- Postmortem action item completion rate > 80%.
- Zero incidents with command ambiguity > 15 minutes.
- 100% of SEV1/2 postmortems published within 15 days.
- Status page updates within 30 minutes for 95% of SEV1.
- 100% of teams integrated into on-call rotations.
- On-call satisfaction score > 7/10.
- 40+ certified Incident Commanders available 24x7.
- SOC 2 Type II audit passed with zero findings.
Steps (22):
1. Executive Mandate & SOC 2 Evidence Clock
Establish clear ownership and funding to transform the CEO's concern into a governed program with immediate compliance implications.
2. Baseline Measurement & Problem Statement (depends on: 1)
Quantify the current state to build a defensible problem statement and establish the data foundation for improvement tracking.
3. Severity & Class Taxonomy (depends on: 2)
Define unambiguous incident levels based on business impact to trigger specific response protocols.
4. Incident Command Roles & Structure (depends on: 3)
Standardize roles and decision rights to eliminate command ambiguity during active incidents.
5. Tooling Consolidation & Integration (depends on: 1)
Deploy a unified incident management platform to replace fragmented alerting tools and create a single source of truth.
6. Detection Strategy & SLO Implementation (depends on: 3)
Shift monitoring focus to customer-impacting symptoms to detect issues before customers do.
7. Alert Quality Standards & Noise Reduction (depends on: 5, 6)
Implement strict alerting rules to reduce noise and restore on-call trust.
8. On-Call Architecture & Coverage Model (depends on: 4)
Design a federated, team-owned on-call structure supported by central command coverage.
9. On-Call Compensation & Wellbeing Policy (depends on: 8)
Introduce paid, sustainable on-call rotations to remove resistance and mitigate burnout.
10. Playbooks & Communication Templates (depends on: 3, 4)
Create standardized playbooks and templates for each severity level to reduce cognitive load.
11. Escalation & Incident Lifecycle Policy (depends on: 4, 5)
Define the mechanical progression from alert to resolution, including automatic escalation paths.
12. Communication Workflows (Internal & External) (depends on: 4, 10)
Establish timed protocols for internal, customer, and regulatory communications.
13. Status Page & Customer Notification Automation (depends on: 5, 12)
Integrate incident management with public status pages to ensure timely and accurate updates.
14. Postmortem Policy & Blameless Framework (depends on: 3)
Mandate and standardize postmortems to focus on system improvement rather than individual error.
15. Action Item Tracking & Governance (depends on: 14)
Ensure postmortem action items are tracked, owned, and completed to prevent recurrence.
16. Metrics, Dashboards & Review Cadence (depends on: 2, 3)
Define and monitor key performance indicators to validate the effectiveness of the incident process.
17. Training, Certification & Drills (depends on: 4, 10, 12)
Build practical skills and readiness for incident response across all roles.
18. Pilot Program with Volunteer Teams (depends on: 5, 7, 9, 11, 13, 15, 17)
Test the end-to-end process on a small scale to refine procedures before full rollout.
19. Phased Rollout to All Teams (depends on: 18)
Deploy the process to all 28 teams in manageable waves to minimize disruption and manage resistance.
20. SOC 2 Control Mapping & Evidence Framework (depends on: 1, 3)
Align the incident management process with SOC 2 requirements to ensure audit readiness.
21. SOC 2 Dry Run & Audit Support (depends on: 19, 20)
Verify evidence sufficiency and prepare the team for the external audit.
22. Governance & Continuous Improvement (depends on: 19, 21)
Establish standing governance to sustain the process and drive ongoing improvements.
Instructions:
1. Carefully evaluate each proposal considering:
- Clarity and detail of the proposed steps.
- Overall alignment with the Main Objective.
- Quality and relevance of the proposed success metrics.
- Coherence of dependencies between steps (if any).
2. Think deeply about each proposal, analyzing its strengths and weaknesses. Consider how well each proposal addresses the main objective and which one would be most effective.
After your thorough analysis, give the Proposal ID of the selected proposal and a one- or two-sentence justification of your choice.
Answer with a single JSON object, and nothing else, that conforms to this JSON schema:
{"description": "What a voting agent must answer with.", "properties": {"selected_proposal_id": {"title": "Selected Proposal Id", "type": "string", "enum": ["80c18452-b9c3-4820-9dec-cf8cb9a8c7b0", "abac560b-e42c-412b-9ecc-a60d55c29045", "265b9bc8-32de-46d3-8546-6dd406da4c24"]}, "justification": {"title": "Justification", "type": "string"}}, "required": ["selected_proposal_id", "justification"], "title": "VoteAnswer", "type": "object", "additionalProperties": false}
{"selected_proposal_id": "abac560b-e42c-412b-9ecc-a60d55c29045", "justification": "Proposal 2 matches Proposal 1's depth and specificity but edges ahead on the brief's hardest asks: its P5 'day-one operating rules' delivers a live minimum-viable process in two weeks, directly answering 'introduce it without waiting for the audit,' and its dedicated resilience track plus culture/incentives steps address the single shared PostgreSQL ledger and the on-call resistance that a pure process plan would miss. Its 23 concrete, time-bound success metrics and regulator clock matrix are at least as complete as Proposal 1's, while Proposal 3 is markedly thinner with one-line step descriptions and fewer metrics."} |
| qwen3.8-flash_voter_3 | qwen3.8-flash · alibaba/qwen3.8-flash |
Proposal B · deepseek-flash_refine_2 winner | This proposal offers the most realistic and detailed roadmap by introducing a 'minimum viable process' to bridge the gap between current chaos and full implementation. It explicitly addresses the specific cultural friction regarding pager ownership through structural separation of duties and financial incentives, ensuring higher adoption rates. |
20.0k in · 2.6k out · 35 s | show[SYSTEM]
You are an expert and objective evaluator of project plan proposals.
Your task is to select the BEST proposal based on criteria of completeness, clarity, and alignment with the main objective.
Use your internal reasoning processes to thoroughly analyze each proposal, considering all aspects and implications.
Take as much time and space as you need to evaluate each proposal in depth before making your decision.
Every text field you write will be read by a busy person, so write for them: short sentences, one idea each, in short paragraphs. Text fields accept Markdown: separate paragraphs with a blank line, use a bullet list when you enumerate things and plain paragraphs when you explain or argue, and bold at most one key phrase per paragraph. A text of more than three sentences must be split into paragraphs; never deliver a single unbroken block of text.
[HUMAN]
Main Objective: "A B2B payments platform in New York (260 engineers in 28 teams, 2,100 customers, $4B processed a month, a 99.95% SLA with service credits) runs about 180 services on Kubernetes across two AWS regions, with a shared PostgreSQL cluster for the ledger. In the last 12 months: 31 customer-impacting incidents, median time to detect 22 minutes (customers detected 40% of them first), median time to mitigate 3 h 10 min, $1.3M paid in SLA credits, two incidents in which nobody was sure who was in charge for over an hour, and a CEO email about "outages we hear about from clients". On-call exists in 12 of the 28 teams, unpaid, fed by alerts from six different tools — 3,400 alerts a month, 85% of them noise; there is no severity scale, status-page updates are written by whoever is around, and postmortems happen for some incidents, in various formats, with action items that are rarely tracked (11 of 64 closed). A SOC 2 Type II audit in eight months will test incident response. Engineers push back against "carrying a pager for other teams' code".
Define the incident management process: severity levels and what each one triggers; roles (incident commander, communications, scribe, subject-matter responders) and how they are staffed 24x7 across 28 teams; the on-call structure, rotations, compensation and the rules for alert quality; detection and escalation paths; internal and customer communications (status page, account managers, regulators when required) with their timings; postmortems (when mandatory, format, blameless review, ownership and tracking of actions); the metrics and reviews that show whether it works; and how it is introduced across the teams without waiting for the audit."
Proposals to Evaluate:
--- PROPOSAL 1 ---
Proposal ID: 80c18452-b9c3-4820-9dec-cf8cb9a8c7b0
Content:
Estimated Complexity: high
Success Metrics: - Median Time to Detect from 22 minutes to <5 minutes for internal detection and <10 minutes for customer journeys by month 6; <2% of customer-impacting incidents first detected by customers (vs. 40% currently) by month 9.
- Median Time to Mitigate for SEV-1 from 3h 10min to <45 minutes by month 9; SEV-2 to <90 minutes by month 9.
- Monthly alert volume from 3,400 to <400 with >95% signal-to-noise ratio by month 5; all alerts satisfy paging contract (symptom-based, owned, runbook-linked, severity-mapped).
- Zero incidents with command ambiguity lasting >15 minutes from month 2 onwards; 100% of SEV-1 and SEV-2 incidents have named IC logged in timeline within 5 minutes; 100% of Triage Owners assigned within acknowledgement.
- 100% of the 28 teams have a documented Service On-Call rotation or a signed service-ownership transfer by month 4; Platform Duty rotation and IC roster both staffed and live by month 3.
- On-call compensation policy published and effective by month 1; on-call satisfaction score ≥7/10 by month 6; zero on-call-attributed voluntary attrition by month 6.
- SLA credits paid from $1.3M annually to <$100K by month 12; credit avoidance (prevented credits) tracked and reported monthly.
- 100% of mandatory postmortems (SEV-0, SEV-1, SEV-2, and repeat incidents) published internally within 15 business days by month 4.
- Postmortem action item completion rate from 17% (11 of 64) to >90% within 60 days by month 6; median action age <30 days; zero repeat incidents caused by the same contributing factor without a design review.
- 100% of the 180 services have a named owner, a detection contract, and at least one symptom-based alert by month 6.
- Status-page first update published within 30 minutes for ≥95% of SEV-1 incidents by month 4.
- 100% of new engineers complete incident-response onboarding within 30 days of joining; IC certification includes written exam and live simulation; ≥2 certified ICs per team group; zero uncovered hours in 24x7 IC roster.
- Weekly operational review held in ≥90% of weeks; 12 of 12 monthly reliability reviews; 4 of 4 quarterly executive reviews; all reviews end with documented decisions and owners.
- SOC 2 Type II audit passes all incident-response controls (CC7.1–7.5, CC2.2–2.3, CC4.1, CC3.x) with zero findings by month 8.
- All pilot and rollout incidents captured with complete golden incident files (timeline, roles, communications, postmortem, actions, closure evidence) by month 3 onwards; audit dry run identifies zero critical gaps by month 7.
- All 28 teams transitioned to new process by month 5; all legacy alert tools decommissioned; single source of truth for incidents established and sustained.
Steps (22):
1. Executive charter, governance structure, and evidence clock
Turn the CEO email into a funded, authorized program with clear ownership and documented evidence collection for SOC 2, starting today.
- Appoint a Director of Incident Management reporting to CTO, with dotted line to COO (customer impact) and Head of Compliance (audit readiness).
- Publish a one-page charter: scope (all customer-impacting, payment-path, data-integrity incidents across 28 teams and 2 regions), decision rights (IC may freeze changes, override team preferences during incidents), and authority to mandate process participation.
- Secure annual budget for tooling, training, on-call compensation ($500–800K estimated), and resilience work. Connect funding to avoided SLA credits ($1.3M baseline).
- Establish standing Incident Management Steering Group: CTO, VP Engineering, Head of Support, Head of Compliance, one engineering manager per region. Meet monthly.
- **Start the SOC 2 evidence clock on day 1.** An audit in eight months means operating-period evidence begins now; design the process to capture evidence continuously, not retroactively.
- Make incident-response participation a documented performance expectation for all engineering managers and team leads.
- Publish timeline: working process in month 2, all 28 teams in month 5, audit-ready in month 7.
2. Baseline measurement, incident register, and evidence pack (depends on: 1)
Establish defensible baseline metrics and identify structural gaps that explain the 40% customer-detected rate and 22-minute detection time.
- Build a 12-month incident register with all 31 customer-impacting incidents: date, detection source, detection time, mitigation time, customer count, services involved, SLA credits paid, root cause class.
- Audit the current alert estate: total volume per tool, volume per team, volume per service, page-to-action ratio, top 50 noisiest rules, off-hours interruptions per engineer.
- Construct a **silent-failure register**: incidents with no internal alert fired at all. This explains the 40% customer-detected rate.
- Reconstruct the two command-ambiguity incidents minute by minute: exactly when did ownership become unclear, how long, what was the decision bottleneck.
- Survey on-call engineers (target >70% response): burden, fairness, pay expectations, escalation clarity, willingness to stay.
- Interview Support and Account Management: how do customers discover incidents, what do they complain about, how do they contact you.
- Publish the problem statement internally; retain all artifacts for SOC 2 audit evidence. This is the baseline against which all improvements are measured.
3. Control mapping and evidence architecture (depends on: 1, 2)
Design the process to generate SOC 2-compliant evidence automatically, from the first real incident, so the audit clock ticks in your favour.
- Map the new process to Trust Services Criteria CC7.1–7.5 (incident identification, response, evaluation, containment, communication), CC2.2–2.3 (authorization), CC4.1 (change management), CC3.x (information availability).
- For each control, write a one-paragraph plain-language statement, name a single owner, and specify the evidence artifact (incident record, timeline, communications log, postmortem, action tracker, training record).
- Define the **golden incident file**: one single-click export per incident containing severity, timeline, roles assigned, decisions made, communications sent, postmortem, and action items. This is the audit unit.
- Specify data retention, immutability, access control, and storage location (not laptops, not chat history that expires). Ensure evidence is searchable and organized by incident date.
- Keep a gap register with owners and dates; review fortnightly in the steering group. Identify which controls are blocked by incomplete rollout and when they unblock.
- Run an early design walkthrough with an experienced SOC 2 readiness partner inside month 1 to stress-test control design before building on it.
4. Severity and response class taxonomy (depends on: 2)
Define four severity levels and four response classes so every decision—paging, communications, postmortem, compensation—keys off a defensible rule, not a judgment call.
- **Severity by impact scope**: SEV-1 (total payment-path loss, data corruption, or >5% transaction failure for >5 min); SEV-2 (significant degradation or single region loss); SEV-3 (limited impact with workaround available); SEV-4 (internal issue or cosmetic); SEV-0 (reserved for security/regulatory/privacy events).
- **Response class** (orthogonal to severity): Availability, Performance, Data Integrity & Ledger, Security & Privacy. **Key rule: class can raise severity, never lower it.** A SEV-3 data-integrity incident gets SEV-1 response posture because integrity is not recoverable by moving faster.
- Automatic triggers: loss of one AWS region → SEV-1 or SEV-2 (class-dependent); ledger write failures → SEV-1; replication lag >10s → escalation review; payment success rate <99% for >5 min → SEV-1/2; missed settlement window → SEV-1; total external API unavailability → SEV-1.
- Who may declare: any engineer, Support, account manager (based on observed customer impact). Who may downgrade: IC only, after investigation.
- Map each level to SLA credit exposure and to customer-facing status-page state.
- Include worked examples from the last 12 months so all 28 teams recognize their own incidents in the taxonomy. Re-validate quarterly against real declarations.
5. Incident lifecycle, Triage Owner rule, and escalation policy (depends on: 4)
Eliminate the "nobody was in charge for over an hour" problem by assigning ownership the moment a page is acknowledged.
- Define lifecycle states with clear entry/exit criteria: Detected (alert fired) → Triaged (is this real and customer-impacting?) → Declared (severity assigned) → Mitigated (core issue resolved) → Resolved (all verifications done) → Postmortem (review scheduled) → Closed (action items tracked or dismissed).
- **Introduce the Triage Owner rule**: the person who acknowledges the page owns the incident until an IC is assigned or the incident is stood down. There is never an unowned gap between first page and declaration. Triage Owner's sole job: decide within 15 minutes whether this requires an IC or a direct stand-down.
- Set aggressive acknowledgement and declaration targets: page acknowledged in 5 min; triage decision (is this real?) in 15 min; severity declaration in 30 min for any customer-facing incident.
- Implement automatic escalation ladders with no human judgment required: if responder does not acknowledge in 5 min, escalate to service owner; if no ack in 10 min, escalate to team manager; if no ack in 15 min, escalate to IC on-call. Escalation is never criticized.
- Define unresponsive-team path: if a service's on-call is unreachable for 30 min, IC may direct any available engineer from any team to engage.
- For SEV-1 and SEV-2: change freeze until IC declares mitigation confirmed; IC unfreezes changes explicitly.
- Enforce one incident, one record. Timeline auto-captured from Slack channel and bridge; never written from memory later.
6. Incident roles, command structure, and decision rights (depends on: 5)
Define clear roles with one-page responsibility cards published and linked from every paging notification.
- **Incident Commander**: owns incident outcome, not the fix. Declares severity, decides escalation, approves all customer communications, freezes changes, calls responders, hands off in shifts. Non-technical ICs are acceptable; technical depth is not required.
- **Deputy IC**: assigned within 15 min of declaration; shadows IC; takes over if IC unavailable or after 4-hour shift on SEV-1. Maximum IC shift: 4 hours on SEV-1, 6 hours on SEV-2.
- **Triage Owner** (new role): owns incident from first page acknowledgement until IC takes over or stand-down decision is made. Required for all incidents.
- **Communications Lead**: owns internal Slack updates and status-page messaging; shields IC from customer contact and interruptions.
- **Scribe**: records real-time timeline with decisions, actions, and key timestamps; not responsible for fixing.
- **Subject-Matter Responders**: engineers with service context; take IC direction; report only to IC; no side channels or parallel debugging.
- **Operations Lead** (SEV-1 only): coordinates multiple responders, manages incident bridge, maintains escalation list.
- Minimum viable staffing: SEV-1 requires all roles; SEV-2 requires IC, Deputy, Comms, Scribe, SMEs; SEV-3 requires Triage Owner and IC.
- Create laminated role cards for every on-call shift location (office, home, printed in pockets).
7. Three on-call rotations: Service, Platform, and Incident Commander (depends on: 5, 6)
Directly address the "carrying a pager for another team's code" objection by making it structurally impossible.
- **Service On-Call rotation** (federated): each of the 28 teams maintains a rotation for their own services only. No engineer is paged for code their team does not own. The answer to "why am I carrying a pager?" is now simply: "for your team's code."
- **Platform Duty rotation** (centrally staffed): shared PostgreSQL cluster, Kubernetes, networking, CI/CD, observability, and incident management tooling. Nobody's product code, so it gets its own dedicated rotation. Staffed from platform teams plus volunteers from other teams; paid at premium rate.
- **Incident Commander roster** (24x7): 12–16 certified senior engineers from across all 28 teams, on one-week primary shifts with secondary backup. Covers every hour with no single point of failure and no uncovered holiday week.
- **Consequences and gates**: 16 of 28 teams have no on-call today. Each must either (a) build a Service On-Call rotation of at least 6 engineers, or (b) formally transfer service ownership to a team that will, with transfer documented and dated. No exceptions, no waivers. Unowned services are decommissioned or transferred by end of month 2.
- Merge small or low-traffic teams into shared rotations where service ownership is unclear (e.g., shared analytics, testing infrastructure).
- Enforce scheduling limits in the tooling: no engineer on-call more than 2 weeks per quarter, automatically enforced by configuration, not negotiation.
- Publish a coverage matrix for all 28 teams showing services, owners, rotation size, gaps, and monthly status.
8. On-call compensation, rest policy, and sustainability (depends on: 7)
Settle compensation before rollout, not during negotiations. Make on-call sustainable and valued.
- **Paid on-call**: effective immediately upon joining a rotation. Weekly stipend while on shift (benchmark to New York market: $600–1,000 per week per engineer), regardless of incident volume.
- **Event-based compensation**: 1.5× hourly rate for time spent mitigating out-of-hours incidents, minimum one-hour block per callout. Tracked by incident record (auto-capture from timeline).
- **Compensatory rest**: no engineer works a normal 8-hour business day after a night incident requiring >2 hours mitigation. Rest day is documented policy, not a favour granted by manager.
- **Intrusion cap**: maximum 3 unscheduled pages per week per engineer. Exceed the cap in a week and trigger an immediate review; exceed in a month and escalate to VP Engineering. Breaches are structural signal that alert quality or service stability has a problem.
- **Voluntary opt-out**: an engineer may exit a rotation; their team must hire or buy replacement coverage from paid pool at published internal rate ($X per shift). This converts culture debate into visible budget decision.
- **Amnesty policy**: incident records, near-miss reports, and false declarations are never used in performance reviews or compensation discussion. Only failure to report is a performance issue.
- **Policy publication**: publish compensation structure and effective date before any team is asked to join a rotation, and include on-call expectations in job descriptions and hiring conversations.
- **Semi-annual review**: reassess compensation and caps every six months against actual page volumes, attrition rates, and survey feedback.
9. Detection strategy: SLOs, synthetic monitoring, and customer-report intake (depends on: 4, 7)
Close the 40% customer-detected gap by monitoring customer journeys instead of infrastructure metrics.
- **SLO-based alerting**: Define SLIs and SLOs for the top 20 customer journeys (payment initiation, authorization, settlement, ledger read/write, API availability, webhook delivery, payout). Measure per region. Alert on SLO breach, not on infrastructure metric (e.g., alert on "payment success rate <99%" not "database CPU >80%").
- **Synthetic transaction monitoring**: deploy synthetic transactions from outside AWS in both regions plus a third geographic location, one-minute cadence, for all money-moving paths. These are your first alarm bell.
- **Ledger-critical signals**: PostgreSQL replication lag (target: <1s, alert >5s), connection saturation, write latency (p95), lock-wait time, transaction ID exhaustion proximity, checkpoint pressure, table bloat. These are separate alerts on shared-database health.
- **Customer-report intake** (new detection channel): Support and Account Managers can raise an incident directly in the platform. Every customer report creates an incident record automatically, and the "customer report" detection source is counted in all metrics. This is a legitimate detection method, not a failure.
- **Detection-gap rule**: whenever a customer reports an incident before internal monitoring fires, auto-create a ticket in the owning service's backlog with root cause: "Monitoring gap on [journey]."
- **Detection contract per service**: every one of the 180 services needs a named owner, at least one symptom-based alert mapped to a SLO, and a documented expected detect time (target: <5 min for payment path, <10 min for others). Published on wiki and reviewed monthly.
- **Detection drills**: run a quarterly drill per team: simulate a broken service in staging and verify it triggers a page before a human notices.
- **Resilience roadmap separation**: detection improvements do not protect against ledger corruption or multi-region failure. Fund a separate resilience roadmap to reduce shared-database blast radius and improve failover safety.
10. Alert quality standards and noise-reduction program (depends on: 9)
Cut the 3,400 monthly alerts (85% noise) to <600 with 95% signal. This is the price of admission for on-call buy-in.
- **Paging contract**: every page must satisfy all of (1) symptom-based (customer impact, not infrastructure cause), (2) actionable (linked runbook with immediate next step), (3) owned (named team responsible), (4) severity-mapped (SEV-1/2/3/4), (5) SLO-linked where applicable. **No runbook, no page.** Enforce with CI check on alert definition.
- **Separation rule**: only customer-impacting or imminent-impact signals trigger pages. Everything else becomes a dashboard entry, a ticket, or a log line. Noisy infrastructure metrics go to dashboards, not pagers.
- **Page budget per service**: no service may exceed 2 pages per on-call shift per month. Exceeding budget auto-opens a remediation ticket in the owning team's backlog (with alert-quality review assigned to tech lead).
- **Automatic suppression rules**: (1) silence alerts if service auto-recovered within 30s, (2) suppress known maintenance windows, (3) group flapping alerts (>5 in 2 min) into one page, (4) rate-limit noisy services (max 1 page per 5 min until condition clears). All suppression rules must have an expiry date; no permanent silence without a ticket.
- **Probation for new alerts**: new alert rules run as tickets only and alert to a Slack channel; after two weeks of proving actionability (every alert resulted in human action), they graduate to pager.
- **Noise sprint**: run a focused 90-day program with a public burn-down of the top 100 noisiest rules. Assign each to a named manager. Default action: fix root cause, tune threshold, or delete within 10 working days. Deletion is a legitimate successful outcome (celebrate it).
- **Correlation and deduplication**: consolidate alert sources at ingest pipeline so one outage triggering 40 alerts produces one page, not 40.
- **Alert ownership**: every alert must have an owning team and a maintenance contact. Update monthly.
11. Incident tooling consolidation and integration (depends on: 5, 10)
Replace six alert tools and ad-hoc incident records with a single source of truth that unifies paging, escalation, timeline, and audit evidence.
- **Tool selection**: choose an incident-management platform (e.g., PagerDuty, Incident.io, Opsgenie) that integrates paging schedules, escalation policies, incident records, postmortem workflow, and status-page APIs. Decision gate: month 1.
- **Event pipeline consolidation**: route all alerts from the six legacy tools into a single event pipeline that feeds the incident platform. Apply deduplication, correlation, severity/class mapping, and rate-limiting at ingest.
- **Observability integration**: connect the incident platform to your Kubernetes dashboards, PostgreSQL monitoring, distributed tracing, and logs so responders see context in one pane. Link runbooks directly into incident records.
- **Slack and bridge integration**: auto-create incident Slack channels, auto-invite roles, auto-capture timeline from channel transcript and voice-bridge recording. Timeline is not written from memory; it is auto-captured.
- **Golden incident file**: implement the export defined in S3. One click produces a complete, immutable, audit-ready PDF: severity, timeline, roles, decisions, communications, postmortem, action items, and closure evidence.
- **Dual-run period**: run both legacy and new platform in parallel for two weeks. Define rollback criteria (e.g.,
12. Escalation automation and incident lifecycle enforcement (depends on: 5, 11)
Eliminate judgment calls from the worst moments. Escalation is automatic, mechanical, and blameless.
- **Automatic escalation ladders**: page responder → if no ack in 5 min, page service owner → if no ack in 10 min, page team manager → if no ack in 15 min, page IC on-call + call them immediately (phone + SMS + Slack). No human decides to escalate; timers fire escalations.
- **Severity-based escalation tempo**: SEV-1 uses faster timers (2 min for IC on-call), SEV-2 uses moderate timers (5–10 min), SEV-3 uses slower timers (15–30 min). Configured in tooling, reviewed quarterly.
- **Dual IC rule**: if a second SEV-1 incident is detected while the first is active, immediately page and assign a separate IC. ICs never run two incidents in parallel.
- **Change freeze and rollback authority**: SEV-1 and SEV-2 trigger automatic deploy freeze. Only the IC (with CTO/VP Eng notification) may unfreeze. Freeze lifts only when IC explicitly declares mitigation confirmed and verifies no new incident symptoms for 5 min.
- **Unresponsive team escalation**: if service's on-call does not acknowledge in 30 min, IC may direct any engineer from any team (volunteers first, then rotated) to engage. This is documented and reported in monthly review (escalation = signal of rotation problem).
- **One incident, one record**: all decisions logged in the incident platform. Auto-capture from Slack, bridge, status-page updates. Timeline is the source of truth; postmortem is written from timeline, never constructed after the fact.
- **Ambiguity rule**: if two responders disagree about whether an incident should be declared, it is declared. False declarations (stand-downs within 30 min of declaration) are tracked as metrics and closed without blame.
- **Watch state**: an unconfirmed incident can live in "Watch" state for max 30 min; after that, either declare it or stand it down explicitly.
13. Internal, customer, and regulatory communications workflows (depends on: 6, 12)
Define who informs whom, in what order, via what channel, with explicit timings and pre-approved templates.
- **Internal cadence**: first update to #incidents Slack channel within 3 min of declaration (even if "Investigating"). Then updates every 5 min (SEV-1), 15 min (SEV-2), or 30 min (SEV-3), or immediately on material change (e.g., mitigation achieved, scope widened). **Comms Lead owns the update; IC must not be interrupted.**
- **Executive notification**: IC calls CTO and VP Eng within 1 min of SEV-1 declaration (not email, not Slack, call). Incident declared in Slack with severity label, IC name, and affected service. Escalation channel lead auto-pinged.
- **Customer communication channels**: status page (all 2,100 customers), direct email to affected customers (top-tier accounts and customers affected by SEV-1), account-manager calls (top 50 accounts on SEV-1).
- **Status page timings**: update within 3 min of SEV-1 declaration, 10 min of SEV-2, 30 min of SEV-3 (even if root cause unknown; use "Investigating" with next-update ETA). Updates every 5–30 min depending on severity. Always include next-update time.
- **Pre-approved templates**: draft customer-facing language for each severity and class in advance with Legal and Compliance. Templates specify impact language ("some of your transactions are delayed" not "our database failed"), workarounds if available, and next-update commitment. Never speculate on cause in customer communication.
- **Regulatory notification path**: identify incidents requiring regulator notification (NYDFS Part 500, money-transmitter rules, payment-card-network rules, securities disclosure). Build a clock matrix: event type → regulator → notification window → signer. Compliance owns all regulatory notifications (never Engineering). Pre-clear templates. Flag incidents to Compliance immediately upon declaration.
- **Account manager playbook**: contact tree for top 50 accounts, templated talking points (facts only, never speculation), escalation path if customer escalates, what to offer (service credit, technical deep-dive call).
- **Closing communication**: resolution notice, SLA credit impact, commitment date for written root-cause report, customer action required (none, or security update, etc.).
14. Status page infrastructure and customer-impact ledger (depends on: 13)
Make the status page reliable, customer-centric, and audit-ready. Track customer impact in a single durable record.
- **Status page decoupling**: host status page outside production failure domain (separate cloud, separate infrastructure, separate database). Integrate incident platform with status page so incident record drives all public updates. Status page survives total platform outage.
- **Component-to-journey mapping**: status page components map to customer journeys ("Payments", "Settlements", "Payouts", "Ledger API") not to internal services. Allow customers to subscribe to components; notify by email or webhook.
- **One-click update templates**: pre-fill status-page template with severity, impact language, next-update time, and estimated resolution. Comms Lead types minimal new info ("Root cause identified" or "Workaround available"), and updates auto-post.
- **Customer-impact ledger** (one record per incident): which customer accounts affected, which journey(s) impacted, exact start and end time of impact, estimated SLA-credit exposure. Use this single record for customer communications, credit computation, regulatory reporting, and annual review. No reconciliation of two versions of the same outage.
- **SLA credit automation**: compute credit based on duration × severity × customer tier → auto-generate customer notification → auto-post to finance system. Reconcile accrued vs. paid credits monthly and report in executive review.
- **Testing during game days**: simulate status-page outage and verify alerts continue to fire; test total region loss and confirm status page remains updated; drill runbook for manually updating status page if platform is down.
15. Postmortem policy: mandatory, blameless, three-level framework (depends on: 6, 13)
Turn postmortems from a writing exercise (11 of 64 action items closed) into the learning engine of the system.
- **Mandatory postmortems**: all SEV-0, SEV-1, and SEV-2 incidents; all SEV-3 with customer impact or repeat pattern; any near-miss IC flags; any incident where the process itself failed (IC unreachable, Comms Lead unavailable, false declaration, missed update SLA).
- **Three-level framework** (proportionate to weight): (1) lightweight async review for SEV-4 and low-impact SEV-3 (10 min template in shared doc, owner + IC review), (2) standard facilitated postmortem for SEV-2 and impactful SEV-3 (full template, facilitated by trained neutral party, published within 10 days), (3) full executive postmortem for every SEV-1 and every security incident (executive sponsor assigned, full investigation, published within 15 days, customer-facing variant prepared).
- **Fixed timeline**: draft postmortem within 5 business days, blameless review within 10 days, internal publication within 15 days.
- **Single template**: impact (who, how many, how long, financial exposure), timeline (detection through resolution), root cause (not "human error" but system condition that enabled error; what was the gap?), contributing factors (tooling, process, organization, knowledge, monitoring), what went well, what went badly, action items (≤3, rest go to reliability backlog).
- **Blameless facilitation**: train a pool of blameless postmortem facilitators (target: 10+ engineers). Require a trained, neutral facilitator for every SEV-1 and SEV-2 review. Prohibit counterfactual language ("if the engineer had"), blame language, and the phrase "human error" as a root cause.
- **Publication rule**: publish all postmortems internally by default; security review only for genuinely sensitive material (e.g., unpatched vulnerability details or customer PII in logs). Create a customer-facing root-cause report for every SEV-1, especially for regulated customers, with legal and compliance sign-off.
- **Searchability**: store postmortems in a searchable wiki or issue tracker with tags (service, class, root cause category) so teams can learn from similar incidents without repeating them.
16. Action item tracking, reliability backlog, and repeat-incident design rule (depends on: 15)
Close the loop on incident learning by enforcing verifiable, tracked action items and breaking cycles of repeat incidents.
- **Action item capping**: each postmortem generates a maximum of 3 action items. Anything beyond 3 goes into a ranked reliability backlog, not into the postmortem, to prevent overwhelming teams.
- **Action item requirements**: each item must have (1) a named human owner (not a team), (2) a due date (≤60 days, target ≤30 days), (3) a definition of done (merged code change, tested alert, audit evidence, architectural decision, new runbook, training completed) not self-reported status.
- **Single reliability backlog**: create one backlog in your engineering tracker (Jira, Linear, etc.) with mandatory label (e.g., `incident-action`), link to originating incident, and link to postmortem. Track progress weekly.
- **Closure sign-off**: Incident Commander or postmortem facilitator must sign off on closure, verifying artifact exists (code merged, alert tested in drill, runbook verified).
- **Repeat-incident rule**: if the same service or component has a second incident with the same contributing factor, **do not create another action item**. Instead, escalate immediately to an architect or tech lead and trigger a design review (not a task, a review). This breaks the cycle of repeated patches; the system needs a structure change.
- **Capacity protection**: reserve a fixed percentage of each team's sprint capacity (10–15%) for reliability work. Track unspent capacity and report to VP Engineering monthly; if a team is not spending it, work with them to identify and fix blockers.
- **Ageing and escalation**: run a weekly review of open actions; escalate anything >30 days overdue to team lead and VP Engineering. Monthly report: completion rate (target >90% within 60 days) and median action age (target <30 days).
17. Training, certification, and exercise program (depends on: 6, 13, 15, 16)
Build skills before deploying the process. Run ongoing drills so the system is tested, not guessed at.
- **Curriculum**: (1) All engineers (30-min async video): severity taxonomy, communication expectations, postmortem format, when to declare an incident, where to find runbooks. (2) On-call responders (2-hr workshop): alert tool walkthrough, playbooks by severity, escalation paths and timers, when to call manager, mobile app walkthrough, case studies from the last 12 months. (3) Incident Commanders (4-hr bootcamp + test): leadership under pressure, decision-making (severity, escalation, rollback), communicating with executives, status-page discipline, postmortem facilitation, handling ambiguity, live simulated incident (pass/fail certification). (4) Communications Leads (2-hr training): templates per severity and class, customer-communication rules (no speculation, no blame), update timings, how to shield IC, regulatory triggers.
- **IC certification**: written assessment (75% pass required) plus live simulated incident (role-play with facilitator, graded on severity declaration, escalation decisions, communication, handover). Certification valid for 12 months; recertify via annual refresher or another live sim.
- **Depth across teams**: certify at least 2 ICs per team or team group so central roster is not siloed in one group; no holiday week is uncovered.
- **Async content**: record all training videos so async teams can learn on their schedule. Create quick-reference cards (laminated, pocket-sized) for roles and playbooks; distribute to on-call locations (office, home).
- **Monthly tabletop exercises**: drawn from real incidents from the last 12 months (region loss, ledger write failure, missed settlement window, cascading failures). Facilitator describes scenario; 3–4 responders play out response (Triage Owner, IC, Comms) as if real. Run 30 min; retro for 15 min afterward.
- **Quarterly game days**: deliberately inject failures into production (database failover, status-page outage, alerting-pipeline outage, dual SEV-1 incidents). All on-call roles engage. Run 2–3 hours; measure response times, decision quality, and communication. Document findings and create action items for identified gaps.
- **Drill the process's own failure modes**: IC unreachable (on-call unavailable, phone broken), Comms Lead on PTO, two simultaneous SEV-1s, paging storm (100+ alerts), false alarm that consumes an hour. Test escalation paths, deputy takeover, and recovery.
- **New-engineer onboarding**: add incident-response module to all engineering onboarding (completion tracked, audit-ready). All engineers must complete within 30 days of joining or transferring in.
18. Metrics, dashboards, and review cadence (depends on: 2, 12, 16, 17)
Measure to prove the system works. Publish live dashboards so every engineer sees the scoreboard and the system is transparent.
- **Outcome metrics**: Median Time to Detect by source (target: <5 min internally detected, <10 min customer journeys); Median Time to Mitigate for SEV-1/2 (target: <60 min SEV-1); customer-detected incidents as % of total (target: <5%); incidents by severity (should be mostly SEV-3/4, few SEV-1); SLA credits paid (target: <$100K/year by month 12); annual credit avoidance vs. program cost.
- **Process metrics**: IC assigned within 5 min (target: >95% of incidents); page acknowledgement rate (target: >98% within 5 min); first-update timeliness (target: >95% within SLA); postmortem timeliness (target: 100% of mandatory postmortems published on time); IC roster coverage (zero uncovered hours, monitored weekly).
- **Health metrics**: alert volume and signal-to-noise ratio per team (trending toward target); off-hours pages per engineer per month (trend, cap enforcement); on-call satisfaction survey (target: >7/10); training completion by team (target: 100% within 30 days); % of services with active detection contract (target: 100%).
- **Never publish incident count as a team metric.** Reward hiding. Instead publish detection metrics (near-misses reported per team, detection gaps closed, false declarations made).
- **Live dashboards**: build dashboards visible to all engineers (not just managers) showing outcome, process, and health metrics. Auto-populate from incident platform and alert tool. Update daily. Link from Slack and internal wiki.
- **Baseline all metrics against S2 evidence pack.** Set 90-day and 12-month targets for each metric. Publish targets and progress monthly.
- **Review cadence**: (1) weekly 30-min operational review (incident by incident from prior week: what went well, what hurt, actions); (2) monthly 60-min reliability review (trends, top causes, action-item aging, alert quality per team); (3) quarterly 60-min executive review (CEO's office: customer impact, SLA credits, top five systemic causes, program ROI).
- **Quarterly process review**: what in the process wasted responder time, what confused people, what should be deleted. Solicit feedback from ICs, Comms Leads, and responders. Document changes and reasoning.
19. Pilot program with 3–4 volunteer teams (depends on: 6, 8, 11, 12, 13, 14, 15, 16, 17, 18)
Do not roll out untested to 28 teams. Run the entire process end-to-end with a small cohort using real incidents as the primary training material.
- **Team selection**: recruit 3–4 volunteers spanning criticality: one payment-path team, one ledger-adjacent team, one shared infrastructure team (platform or Kubernetes), one low-traffic team. Volunteers see early adoption and influence.
- **Full process in pilot**: new severity and class taxonomy (S4), consolidated tooling (S11), roles and Triage Owner (S5–6), three rotations (S7), escalation automation (S12), communications (S13–14), postmortems (S15), action tracking (S16), paid on-call (S8), training (S17), metrics (S18). This is not a partial test; it is the complete system.
- **Real incidents are the training**: hold a retro within 48 hours of each pilot incident (while memory is fresh). Process Owner facilitates. Discuss: what worked, what hurt, how is the runbook, is the alert tuned, did Comms template work, did roles work, was timeline auto-captured correctly. Document feedback and iterate.
- **Instrument against baseline**: measure all S18 metrics in pilot; compare results with S2 baseline. If MTTD dropped from 22 min to 8 min, that is the number you carry into the next conversation.
- **Weekly retros with pilot teams**: structured 30-min sessions reviewing policies, tooling, training, templates, and process changes. Change the written policy or tooling in response; document what changed and why so later teams see the evolution, not just the finished state.
- **Explicit exit criteria**: (1) rotation coverage sustained (≥6 engineers per rotation, zero unacknowledged pages over 3 weeks), (2) postmortems delivered on time (100% of mandatory postmortems published within 15 days), (3) action tracking established (100% of action items in backlog with owner and date), (4) metrics live (dashboards updated daily, first weekly review completed).
- **Pilot report**: document before/after numbers (MTTD, MTTR, alert noise, action completion rate, on-call satisfaction) and key process learnings. This report is the foundation for every conversation in the rollout.
20. Phased rollout sequenced by cost of failure (depends on: 16, 18, 19)
Rollout is a staged migration with readiness gates and sequencing that prioritizes visible impact, not ease.
- **Wave sequencing**: divide 28 teams into 4 waves of ~7 teams each, **ordered by incident density and customer-journey ownership** (highest-cost-of-failure teams first). Teams with the most SLA credits at stake go first; their improvement is the proof.
- **Wave spacing**: three weeks between waves. This gives each wave time to stabilize and find problems before the next cohort joins.
- **Readiness checklist per team**: (1) all services mapped and owned (no unowned services), (2) alerts cleaned to paging contract (runbook linked, severity mapped), (3) playbooks updated and tested in staging, (4) rotation staffed to ≥6 engineers, (5) team completes training module, (6) manager briefed on policy, (7) on-call compensation in effect.
- **Gate review before each wave**: Process Owner holds gate review with target teams. Move unready teams to next wave with a dated remediation plan. No exceptions, no waivers; readiness is non-negotiable.
- **Wave champion**: assign a named engineer per wave to champion the rollout, answer questions, escalate issues to Process Owner. Champions are not representatives; they are advocates and feedback collectors.
- **Communication cadence**: weekly all-hands or newsletter for 4 weeks before each wave. Explain why (owned-code-owned-pager rule, paid on-call, SLA credit savings). Use pilot numbers. Answer FAQs. Announce champion and escalation path.
- **First incident under new process**: hold a retro within 48 hours. Feed accepted process changes back through change control.
- **Retire legacy tools and processes**: at end of each wave, retire legacy alert tools, informal escalation lists, ad-hoc status-page process. No parallel processes running for >3 weeks; this prevents confusion and half-learning.
- **Sequence to avoid audit collision**: ensure no team is rolling out in the same week as the audit dry run (S21).
21. SOC 2 dry run and evidence review (depends on: 3, 20)
Convert a good working process into a provable one. Test control evidence a few months before auditors arrive, when you can still fix gaps.
- **Dry run timing**: run 6 weeks before audit window (around month 7 of this program).
- **Scope**: sample 10–15 real incidents from pilot and early rollout waves. For each incident, verify evidence artifact exists and is complete: incident record, timeline (auto-captured), severity and class declaration, roles assigned and logged, communications log (Slack + status page), postmortem (if mandatory), action items in tracker with due dates, action completion evidence (code, alert test, training record).
- **Control walkthrough**: walk through each control statement from S3 with a checklist. Is the evidence artifact present? Is it immutable? Is it searchable? Is retention adequate? Is access logged?
- **Gap remediation**: for every gap found, estimate time to fix and prioritize by audit risk. Anything risking a qualified opinion (e.g., missing postmortem, no timeline evidence) must be fixed before the audit. Test the remediation against a new incident or a resample.
- **Interview readiness**: brief 10–15 engineers who may be interviewed by auditors (ICs, Comms Leads, Process Owner, team managers). Ask them to describe the process as they actually practice it, not as written. Listen for confusion or gaps in understanding. Correct them.
- **Auditor package preparation**: assemble process documentation, sample incident records (5–10 complete golden files), training records, on-call schedules, alert quality metrics, action tracker register, and status-page archive. Organize by control. Create a table of contents and index.
- **Single audit liaison**: designate Process Owner or a small dedicated compliance person as sole point of contact for audit requests. Prevents requests scattering across 28 teams.
- **Rehearsal**: conduct mock interview with an IC and a Comms Lead. Auditors ask tough questions under pressure ("How do you know the timeline is accurate?", "What happens when both ICs are unavailable?", "Show me how you proved the alert was actionable."). Practice answering.
22. Standing governance and continuous improvement (depends on: 20, 21)
Lock in durable improvement. The classic post-audit failure is the process freezing and then decaying. This step prevents that.
- **Standing Incident Management Council**: chaired by Process Owner, monthly meetings, attendees: CTO, VP Engineering, Head of Support, Head of Compliance, one engineering manager per region, one IC, one team manager from recent wave. Agenda: metrics review, policy changes, gaps from recent incidents, escalation for contentious issues.
- **Change control mandate**: give Process Owner documented authority to change severity taxonomy, response classes, roles, communications timings, and compensation policy. Any change requires: written justification, steering group approval (monthly cadence), and documented effective date before implementation. This prevents silent drift and ensures changes are deliberate.
- **Quarterly validation**: re-validate severity and response class taxonomy against all incidents from the prior quarter. Ask: did our taxonomy correctly predict response posture? Did we misclassify? Update taxonomy if patterns emerge.
- **Annual metric re-baselining**: every 12 months, re-run measurements from S2 (alert census, incident register) to reset targets. System should improve; targets should tighten.
- **Resilience roadmap separation**: fund a distinct architectural or platform roadmap for incident prevention (reduce shared-database blast radius, multi-region failover, deploy safety, observability investments). Better incident response does not protect a single-ledger corruption or unplanned failover. These are separate problems.
- **Quarterly executive report**: CTO and VP Eng report to CEO/CFO on metric set (detection time, mitigation time, SLA credits, customer-detected %), top five systemic causes of incidents, program cost vs. credit avoidance, and strategic architecture changes in flight.
- **Public backlog of improvement ideas**: teams and engineers propose process improvements via Slack or wiki. Process Owner reviews quarterly and implements accepted ideas (e.g., "add a dashboard for detection gaps", "update postmortem template"). Publish what changed and why.
- **Celebration and learning**: share wins publicly each quarter ("We reduced MTTD from 22 min to 5 min", "Customer-detected incidents down 80%", "$800K SLA credits avoided"). Refresh training and tabletop program annually and immediately after any SEV-1 to keep the system sharp and responsive to new scenarios.
--- PROPOSAL 2 ---
Proposal ID: abac560b-e42c-412b-9ecc-a60d55c29045
Content:
Estimated Complexity: high
Success Metrics: - Median time to detect falls from 22 minutes to under 8 minutes by month 6, and under 5 minutes for money-moving journeys by month 9.
- Share of customer-impacting incidents first detected by customers falls from 40% to below 10% by month 9.
- Median time to mitigate for SEV1 falls from 3 h 10 min to under 60 minutes by month 9.
- 100% of SEV1 and SEV2 incidents have a named Triage Owner within 5 minutes and a named Incident Commander within 15 minutes, from month 4.
- Zero incidents with unclear command authority lasting more than 15 minutes, measured monthly from month 4.
- Monthly page volume falls from 3,400 to under 600, with a false-positive rate below 15%, by month 5.
- No service exceeds two pages per on-call shift for three consecutive months by month 6.
- 100% of the 180 services have a named owner, a detection contract and at least one symptom-based alert by month 7.
- All 28 teams have a documented rotation, a service ownership map and at least one trained on-call engineer by month 6.
- 100% of on-call shifts are paid under a published policy from month 2, with zero on-call-attributed voluntary attrition by month 6.
- On-call satisfaction scores 7 out of 10 or better in quarterly surveys from month 6.
- The IC roster holds at least 12 certified ICs covering 24x7 with no uncovered week from month 4.
- 100% of SEV0, SEV1 and SEV2 postmortems are published internally within 15 business days from month 5.
- Action items closed within 60 days rise from 17% to above 90%, with a median action age under 30 days, by month 6.
- Status page first update is posted within 30 minutes on at least 95% of SEV1 incidents from month 4.
- Zero missed regulatory notification windows on any incident requiring notification.
- The customer-impact record is complete for 100% of customer-impacting incidents from month 5, with SLA credits reconciled monthly.
- Annual SLA credits paid fall at least 50% from the $1.3M baseline within 12 months.
- Customer-impacting incidents fall at least 40% year over year from the 31-incident baseline.
- Review cadence is sustained: weekly operational review in at least 90% of weeks, 12 of 12 monthly reliability reviews, 4 of 4 quarterly executive reviews.
- Every page has a recorded disposition — fixed, tuned or deleted — within 10 working days, from month 4.
- SOC 2 Type II is passed with zero findings related to incident response.
- 100% of new engineers complete the incident-response onboarding module within 30 days of joining.
Steps (24):
1. Mandate, one owner, and the evidence clock
This step turns the CEO's email into a funded programme with a single accountable owner, and it starts the SOC 2 clock on day one.
- Appoint a Director of Incident Management reporting to the CTO, with a dotted line to the COO for customer and SLA matters and to the Head of Compliance for audit readiness.
- Publish a one-page charter: scope covers every customer-impacting, money-moving, security and data-integrity event across all 28 teams and both regions.
- Grant explicit authority to declare an incident, set severity, freeze deploys, page any engineer in the company, and approve customer messaging.
- Fund the envelope up front: tooling, training and drill time, on-call compensation, and a small programme team, roughly $500–700K a year against $1.3M in credits paid.
- State the return plainly to the steering group: credits avoided, churn avoided, and audit findings avoided.
**Start the evidence clock now.** A SOC 2 Type II report covers a period of operating effectiveness, so evidence cannot be backfilled; an audit in eight months means the operating period begins in the first weeks of this programme.
- Stand up a fortnightly steering group of the CTO, VP Engineering, Head of Support, Head of Compliance and two engineering managers.
- Open a programme risk register with the top risks, owners and review dates, and revisit it at every steering group.
- Make participation in the incident process a documented performance expectation for every engineering manager, not an optional extra.
2. Baseline evidence pack and cost-of-downtime model (depends on: 1)
You cannot demonstrate improvement without a defensible starting point, and you cannot win the argument about noise without numbers.
- Build a 12-month incident register and re-classify all 31 customer-impacting incidents by detection source, duration, customers affected and credits paid.
- Build the silent-failure register: incidents where no internal alert fired at all, which is the number that explains the 40% customer-detected rate.
- Run an alert census per tool, per team and per service: volume, page-to-action ratio, off-hours interruptions, and the 50 noisiest rules with a named owner each.
- Survey every on-call engineer and manager on burden, fairness, escalation clarity and pay expectations, targeting above 70% response.
- Interview Support, Account Management and Sales about how customers learn of incidents and what they are promised.
- Reconstruct the two command-ambiguity incidents minute by minute to find exactly where ownership lapsed.
- Build a cost-of-downtime model: dollars per minute of impact per customer journey, used later to sequence teams and justify funding.
- Publish the pack internally as the problem statement, and retain every artifact as management-review evidence for the audit.
3. SOC 2 control mapping and evidence architecture (depends on: 1)
Most programmes leave compliance to the end; this one maps controls in month one, because the mapping decides what the process must capture from day one.
- Map the process to the Trust Services Criteria for incident identification, evaluation, response, containment, recovery and communication (CC7.1–CC7.5, CC2.2, CC2.3, CC4.1, CC3.x).
- Write each control in plain language, with one named owner and its evidence artifact.
- Define the golden incident file: one auto-assembled, single-click export per incident containing timeline, roles, severity, communications, postmortem, actions and closure proof.
- Set retention, storage location and immutability so no control depends on a laptop, a private channel, or chat history that expires.
- Keep a gap register with owners and dates, reviewed fortnightly by the steering group.
- Meet the auditor's readiness team inside the first 90 days to test the control design before anything is built on top of it.
4. Severity and class taxonomy with the trigger matrix (depends on: 2)
Severity decides how loud the response is; class decides which playbook runs and what must never be traded away.
- Define severity in customer and money terms: customers affected, dollars at risk, regions lost, ledger integrity, settlement delay, data loss and SLA exposure.
- Use SEV0 for security, privacy or regulatory events; SEV1 for total or material loss of a payment path; SEV2 for degradation or single-region loss; SEV3 for limited impact with a workaround; SEV4 for internal-only issues and near-misses.
- Define response classes: Availability, Performance, Data integrity and Ledger, Security and Privacy, Third-party dependency, Settlement and Financial, Process failure.
**Class can raise the response but never lower it.** A SEV2 data-integrity incident gets SEV1 posture, because integrity failures are not recoverable by moving faster.
- Specify automatic triggers: loss of one AWS region, ledger write failures, replication lag above threshold, a missed settlement cut-off, payment success rate below threshold for five minutes.
- State who may declare — any engineer, Support agent or account manager — and who may downgrade: the Incident Commander alone.
- Map every level to its SLA credit exposure and to the customer-visible status page state.
- Include worked examples from the last 12 months so each of the 28 teams recognises its own incidents, and revalidate the taxonomy quarterly against real declarations.
5. Day-one operating rules and the minimum viable process (depends on: 1, 4)
The full process will take months; the first useful version must be live in two weeks using the tools that already exist.
- Publish ten day-one rules that need no procurement: a named owner within five minutes, one channel per incident, one register entry per incident, one person speaking to customers.
- Apply the ambiguity rule: if two responders disagree about whether this is an incident, it is an incident.
- Make declaring free: a false alarm is closed as a false declaration, tracked as a metric, and never criticised.
- Require a register entry within 24 hours for every customer-impacting incident, even a minimal one.
- Ban silent incidents: if we know, the customer hears it from us rather than from their own reconciliation.
- Run the first 30 days on manual command, with a rotating duty Incident Commander drawn from the 12 teams that already have on-call.
- Hold a 15-minute daily incident stand-up during month one to catch friction while it is still fresh.
6. Roles, command structure and the no-unowned-minute rule (depends on: 4)
The two hour-long command failures did not happen at declaration; they happened in the gap before it, when an alert had fired and nobody owned it.
- Publish one-page role cards: Triage Owner, Incident Commander, Deputy IC, Operations Lead, Internal Comms Lead, Customer Comms Lead, Scribe, Subject-Matter Responders, and Executive Sponsor for SEV1 only.
- Introduce the Triage Owner rule: whoever acknowledges the page owns the incident until an IC takes over or the incident is stood down.
**The IC owns the incident, not the fix, and does not debug.** An IC who starts troubleshooting has abandoned command.
- Give the IC explicit decision rights: declare and escalate severity, freeze changes, halt deploys, approve customer messaging, and pull in any engineer.
- Define minimum viable staffing per severity: SEV1 fills every role; SEV2 staffs IC, scribe, comms and responders; SEV3 staffs an IC and a scribe.
- Set handover discipline: four-hour maximum IC shifts on SEV1 with a written handover, and a deputy named within 15 minutes of declaration.
- Set responder behaviour: one channel, one bridge, no side channels, and every request phrased with a named owner and a time.
- Link the role cards from every paging notification so they are one tap away at 3 AM.
7. Lifecycle, declaration and escalation policy (depends on: 4, 6)
This step defines the mechanical path from an alert to a declared incident and back to normal service, removing judgment calls from the worst moments.
- Define states with entry and exit criteria: Detected, Triaged, Declared, Mitigated, Resolved, Postmortem, Closed, plus a Watch state with a hard 30-minute timer.
- Set targets: page acknowledged within five minutes, triage decision within fifteen, declaration within thirty for anything customer-facing.
- Build automatic escalation ladders with timers at every layer: responder, service owner, team manager, IC on-call, VP Engineering.
- Define the unresponsive-team path: fifteen minutes escalates to the team manager, thirty minutes to the IC, who may then direct any available engineer.
- Define change freeze and rollback authority during SEV1 and SEV2, with the single named condition that lifts the freeze.
- Enforce one incident, one record, with the timeline captured automatically from the channel and bridge rather than written from memory afterwards.
- Add the two-simultaneous-incidents rule: a second IC is designated, and no IC ever runs two incidents at once.
- Test every escalation path weekly with synthetic pages, and adjust the timings after the first month of real operation.
8. On-call architecture across 28 teams (depends on: 4, 6)
The objection is that engineers will not carry a pager for another team's code; the answer is to build rotations that make the objection structurally impossible.
- Run a Service On-Call rotation per team, covering only that team's own services.
- Run a Platform Duty rotation for genuinely shared infrastructure: the PostgreSQL ledger cluster, Kubernetes, networking, CI/CD and observability.
- Run a central Incident Commander roster of 12–16 certified senior engineers on one-week shifts with a primary and a secondary.
**State the consequence honestly.** Sixteen of 28 teams have no rotation today; each must build one or formally transfer service ownership to a team that has one, with the transfer dated and recorded.
- Merge small or low-traffic teams into shared rotations so no rotation has fewer than six engineers.
- Cap load in the scheduling tool: no engineer on call more than two weeks per quarter, enforced by configuration rather than negotiation.
- Publish a coverage matrix of all 28 teams showing services, owners, rotation size and gaps, reviewed monthly.
- Write on-call participation into engineering job levels and hiring criteria, so the commitment is set before someone joins.
9. Compensation, rest and the economics of opting out (depends on: 8)
Unpaid on-call is the most cited reason for resistance, so settle compensation before rollout, not during it.
- Move to paid on-call: a per-shift stipend benchmarked to the New York market, agreed with HR and Finance, published with an effective date before any team is asked to join a rotation.
- Pay for callouts at 1.5× the hourly rate for time actually spent mitigating, with a minimum block per interruption.
- Provide documented compensatory rest: no normal working day after a night incident, and the rest day is policy rather than a favour granted by a manager.
- Cap intrusion with a maximum number of off-hours pages per shift, and trigger a mandatory review whenever the cap is exceeded.
**Allow opt-out, but put a price on it.** An engineer may step out of a rotation, and their team buys coverage from the paid pool at a published internal rate, which turns a cultural argument into a visible budget decision.
- Publish amnesty: incident records, near-misses and false declarations are never used in performance reviews; only failure to report is a performance issue.
- Protect against gaming: pay attaches to shifts and callouts, never to incident counts.
- Check the New York labour, overtime and tax treatment with Legal and Finance before announcing, and review the policy every six months against real page volumes, attrition and survey results.
10. Detection strategy: journeys, synthetic signals and customer-report intake (depends on: 4)
Customers detected 40% of incidents first, which makes detection the highest-leverage business problem in this programme.
- Define SLIs and SLOs for the top 20 customer journeys, measured per region: payment initiation, authorisation, settlement, ledger read and write, API availability, webhook delivery and payout.
- Alert on symptoms against those SLOs, not on cause-based infrastructure metrics.
- Add synthetic transaction monitoring from outside AWS in both regions plus a third location, on a one-minute cadence for money-moving paths.
- Add ledger and shared-PostgreSQL signals: replication lag, connection saturation, write latency, lock waits, transaction ID exhaustion, checkpoint pressure and bloat.
- Open a customer-report intake so Support and account managers can raise an incident directly, and count that path as a detection source in all reporting.
- Apply the detection-gap rule: whenever a customer reports first, a detection-gap ticket is opened automatically and owned by the responsible team.
- Publish a detection contract per service — owner, at least one symptom alert, documented expected detect time — for all 180 services.
- Run a detection drill per team: break something in staging and see whether it pages before a human notices.
11. Paging contract and the noise-reduction programme (depends on: 2, 10)
3,400 alerts a month at 85% noise is the reason engineers resent the pager, and fixing it is the price of admission for everything else in this plan.
- Publish the paging contract: every page must be symptom-based, actionable, owned, mapped to a severity and class, and linked to a runbook.
**No runbook, no page**, enforced by a CI check on the alert definition itself.
- Separate paging from ticketing: only customer-impacting or imminent-impact signals may wake a human.
- Set a page budget per team and per service, with a remediation ticket opened automatically, owned by the engineering manager, when the budget is breached.
- Put new alerts on two-week probation as ticket-only until they have proved actionable.
- Define the default action for a noisy alert: fix, tune or delete within ten working days, with deletion treated as a legitimate and celebrated outcome and the deleted count published.
- Deduplicate and correlate at the ingest pipeline so one root cause produces one page instead of forty.
- Require an expiry date on every silencing rule and temporary threshold, so suppression cannot quietly become permanent blindness.
- Run a 90-day noise sprint with a public burn-down of the top 100 noisiest rules, each owned by a named manager.
12. Incident tooling consolidation and the golden incident file (depends on: 3, 6, 10, 11)
Six alerting tools and no single incident record are structural causes of the 22-minute detection and the three-hour mitigation.
- Choose one incident platform for paging, schedules, escalation policies, incident records, postmortem workflow and status-page integration, and choose it within four weeks.
- Consolidate the six alert sources into a single event pipeline feeding that platform, with deduplication, correlation and severity and class mapping applied at ingest.
- Integrate observability, runbooks, chat and voice bridges so responders see dashboards and runbooks inside the incident record, and the timeline is captured automatically.
- Implement the golden incident file export defined in S3, so audit evidence is one click rather than a reconstruction.
- Run a dual-run period alongside the legacy tools with a defined rollback, then switch the legacy tools off on a published cutover date.
- Host the status page outside the production failure domain so it survives a total platform outage, and prove that in a game day.
- Make the platform usable from a phone: an Incident Commander must be able to run a SEV1 from a mobile device at 3 AM.
13. Internal, customer and regulator communications (depends on: 6, 7, 12)
Today the status page is written by whoever is around; this step replaces improvisation with a clock, a named owner and pre-cleared templates.
- Set the internal cadence: first update within 15 minutes of declaration, then every 30 minutes for SEV1 and hourly for SEV2, whether or not there is progress.
- Never let an employee learn of an incident from the status page: internal communication leads, external follows.
- Set the customer cadence: status page updated within 30 minutes of a SEV1 and 60 minutes of a SEV2, with a no-new-information update still mandatory.
- Define the channel hierarchy: status page for all 2,100 customers, direct email to affected customers on SEV1, and named account-manager calls for the top 50 accounts.
- Pre-approve templates per severity and class with Legal and Compliance, each carrying its own next-update time.
- Forbid speculation: customer messages never guess at cause, never assign blame, and never commit to a root cause before the postmortem.
- Define the closing communication: resolution notice, how SLA credits will be handled, and the committed date for the written report.
- Build a regulator clock matrix covering event type, regulator, notification window, signer and the shortest applicable clock, including NYDFS Part 500, money-transmitter and banking notifications, breach notification, card-network rules and public-company disclosure.
**The regulatory clock starts at awareness, not at root cause.** Route every notification through Compliance, never Engineering, and pre-clear the templates.
14. Customer trust workstream and the SLA credit ledger (depends on: 13)
The $1.3M in credits is a symptom of having no single record of customer impact, and the CEO's inbox is a symptom of customers learning things late.
- Maintain one durable customer-impact record per incident: which customers, which journeys, from when to when, and the estimated credit.
- Use that one record for communications, credit computation, regulatory reporting and the annual review, so nobody reconciles two versions of the same outage.
- Automate the credit path end to end — computation, approval, customer notification and finance treatment — so credits stop being a manual scramble three weeks later.
- Reconcile accrued against paid credits monthly and report the result in the executive review.
- Give the status page a named product owner and map its components to customer journeys, not to internal services.
- Send a CTO-signed reliability note to the top 50 accounts and publish a quarterly reliability report to all customers.
- Give account managers a script of the facts they may state, the speculation they may not, and a path for customer escalations.
- Track credit avoidance against programme cost, so the funding case stays a number rather than an argument.
15. Postmortems: mandatory set, three levels, blameless by design (depends on: 6)
Postmortems currently happen for some incidents, in various formats; this step makes them mandatory where they matter and light where they do not.
- Make postmortems mandatory for every SEV0, SEV1 and SEV2, every SEV3 with customer impact or a repeat pattern, every near-miss the IC flags, and every incident where the process itself failed.
- Use three levels so the ritual matches the weight: a lightweight async review for SEV3 and SEV4, a facilitated postmortem for SEV2, and a full review with an executive sponsor for SEV1.
- Set deadlines: draft within five business days, blameless review within ten, internal publication within fifteen.
- Adopt one template: impact, timeline, detection, response, contributing factors across tooling, process, organisation and human factors, what went well, what went badly, and actions.
- Train a pool of blameless facilitators and require a trained one for every SEV1 review, never the IC.
**Ban blame language in the template and ban "human error" as a root cause.** The question is always what system condition made the error possible.
- Publish all postmortems internally by default, with security review only for genuinely sensitive material.
- Produce a customer-facing root-cause report for SEV1 incidents, especially for regulated and top-tier accounts.
16. Action items: capped, verifiable, with the repeat-incident rule (depends on: 15)
Eleven of 64 action items closed is not a tracking problem; it is a generation problem, because the process produces more actions than the organisation can absorb.
- Cap each postmortem at three action items, with anything beyond that going into a ranked reliability backlog.
- Require every action to have a named human owner, a due date, and a definition of done that is an artifact: a merged change, an alert that fired in a drill, or a test that fails without the fix.
- Prohibit self-reported closure; closure requires the artifact and sign-off by the process owner or the Incident Commander.
- Keep one reliability backlog in the engineering tracker with a mandatory label and a link back to the originating incident.
- Reserve a fixed share of each team's sprint capacity for reliability work, with unspent capacity visible to the vice presidents.
- Apply the repeat-incident rule: a second incident with the same contributing factor triggers a design review owned by an architect, not another action item.
- Run a weekly ageing review of open actions and escalate anything more than 30 days overdue to the VP of Engineering.
- Report action completion rate and median action age monthly, by team.
17. Metrics, dashboards and the review cadence (depends on: 2, 4, 16)
Define what good looks like, then measure it in a way that rewards reporting incidents rather than hiding them.
- Define outcome metrics: time to detect by source, time to mitigate by severity and class, share of incidents detected by customers, incidents by severity and class, SLA credits paid, and customer-impact minutes.
- Define process metrics, because the process itself has SLOs: IC assigned within five minutes, page acknowledgement rate, first-update timeliness, update-cadence adherence, postmortem timeliness, action closure rate, action median age, and IC roster coverage.
- Define health metrics: alert volume and noise ratio per team, off-hours pages per engineer, on-call satisfaction, training completion, and share of services with a detection contract.
**Never publish incident count as a team metric.** It rewards hiding incidents; publish reporting metrics instead — near-misses filed, detection gaps closed, false declarations made.
- Publish live dashboards visible to every engineer, refreshed daily, with every metric baselined against the S2 evidence pack and given a 90-day and a 12-month target.
- Run a 30-minute weekly operational review incident by incident, a monthly reliability review of trends and actions, and a quarterly executive review with the CEO's office.
- Hold a quarterly review of the process itself: what wasted time, what confused responders, and what should be deleted.
- End every review with decisions and named owners, never with numbers alone.
18. Training, certification and the drill programme (depends on: 6, 7, 13, 15)
A process that lives only on a wiki page fails on the first real page, so skills are built and tested before they are needed.
- Build a practical curriculum: how to be on call, how to triage and declare, how to run an incident as IC, how to communicate, and how to write a postmortem.
- Require certification before joining the IC roster: a written assessment plus a live simulated incident.
- Certify at least two ICs per team group so the central roster has depth across all 28 teams and no holiday week is left uncovered.
- Train communications leads separately on templates, cadences, customer language and the regulatory rules.
- Run monthly tabletops drawn from the last 12 months of real incidents, including region loss, ledger write failure and a missed settlement cut-off.
- Run quarterly game days with deliberately injected failure, including PostgreSQL failover, status-page outage and alerting-pipeline outage.
- Drill the process's own failure modes, not just technical ones: IC unreachable, comms lead on leave, two simultaneous SEV1s, a paging storm, and a false alarm that burns an hour.
- Audit the process for single points of failure: who alone can perform each critical task, and what happens in their holiday week.
- Keep a mandatory onboarding module for every engineer joining or transferring in, with audit-ready completion records.
19. Pilot with three to four teams, using real incidents (depends on: 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 18)
Do not roll out to 28 teams untested, and do not rely on drills alone when real incidents are better training material.
- Recruit three to four volunteer teams spanning criticality: a payment-path team, a ledger-adjacent team, the platform team that owns the shared PostgreSQL cluster, and a low-traffic team.
- Run the entire process end to end in the pilot: severity and class, roles, escalation, communications, postmortems, action tracking, and paid on-call.
- Treat real incidents during the pilot as the primary training material, and hold a retro within 48 hours of each one, run by the process owner while the friction is fresh.
- Instrument the pilot against the baseline and publish before-and-after numbers.
- Hold weekly retrospectives with the pilot teams, and change the written policies, the tooling and the training in response, documenting what changed and why.
- Set explicit exit criteria: rotation coverage achieved, zero unacknowledged pages over two weeks, postmortems delivered on time, and actions tracked to closure.
- Produce a pilot report that every later rollout conversation starts from.
20. Phased rollout to 28 teams, sequenced by cost of failure (depends on: 16, 18, 19)
Rollout is a staged migration with readiness gates, not an email announcement, and the sequencing matters more than the schedule.
- Split the 28 teams into four waves of roughly seven, ordered by incident density and customer-journey ownership: the teams whose incidents cost the most SLA credits go first, because their improvement is the visible proof.
- Leave three weeks between waves and define a per-team readiness checklist: services mapped and owned, alerts cleaned to the paging contract, runbooks written, rotation staffed to at least six engineers, training complete, manager briefed, compensation in effect.
- Hold a gate review with the process owner before each team joins, and move an unready team to the next wave with a dated remediation plan rather than granting an exception.
- Give each wave a named champion and explain the why using the pilot's numbers.
- Handle resistance directly and early: publish the own-your-code-own-your-pager rule and the paid on-call mechanics before each wave, never after.
- Hold a retro within 48 hours of each wave's first incident under the new process, and push accepted changes through change control.
- Retire legacy tools, informal escalation lists and the ad-hoc status page process at the end of each wave on a published cutover date, so two processes never run in parallel for long.
- Sequence the waves so no team faces both a rollout and the audit dry run in the same week.
21. Audit dry run and evidence review (depends on: 3, 20)
This step converts a good process into a provable one, about six weeks before the auditors arrive.
- Sample real incidents from the pilot and the early waves against each control's evidence requirements.
- Remediate every gap found and re-test the remediated control against the same sample, prioritising anything that risks a qualified opinion.
- Brief every engineer who may be interviewed so they can describe the process as they actually practise it, not as it is written.
- Prepare the auditor package: process documentation, sample incident records, paging logs, communication logs, training records, on-call schedules, the action register and alert quality metrics.
- Designate one audit liaison and a small evidence-request team, so requests do not land on all 28 teams at once.
- Rehearse the walkthrough with an Incident Commander and a communications lead, because auditors probe realism under pressure.
- Keep the audit liaison and the process owner as close to the same person as possible, so accountability for the control is also accountability for the evidence.
22. Standing governance and process ownership (depends on: 20, 21)
The classic post-audit failure is that the process freezes and then decays, so ownership has to outlive the programme.
- Establish a standing Incident Management Council chaired by the process owner, meeting monthly with engineering, support, compliance and product representation.
- Give the process owner a documented mandate to change standards, with formal change control for any change to severity, classes, roles, communication timings or compensation.
- Revalidate the severity and class taxonomy quarterly against real declarations, and re-baseline the entire metric set annually.
- Report quarterly to the executive team on the metric set and the top five systemic causes of incidents.
- Feed incident themes into architecture review, release engineering and hiring, so the programme improves the system and not only the response.
- Keep a public backlog of process improvement ideas so any engineer can propose a change, and act on it every quarter.
23. Resilience track for the shared ledger and multi-region blast radius (depends on: 1, 17)
Better incident management cannot save a single shared ledger during a corruption event, so this runs as a funded engineering track beside the process, not inside it.
- Reduce ledger blast radius: read replicas for read paths, connection-pool isolation per service domain, statement timeouts, write throttles and per-tenant limits.
- Build a ledger protection mode: an automatic tripwire that degrades to read-only or queues writes before corruption can spread.
- Rehearse multi-region failover with the ledger as the binding constraint, and set a measured recovery time objective per region.
- Test restore and point-in-time recovery on a schedule, and publish the timings.
- Tighten deploy safety: progressive rollout, automatic rollback on SLO burn, and change-freeze windows around settlement.
- Track a small number of resilience bets with named owners and quarterly milestones, reviewed in the same executive review as the incident metrics.
24. Culture, incentives and continuous improvement (depends on: 20, 22)
Incident management is a cultural change before it is a process change, and the incentives decide whether reporting goes up or down.
- Restate the amnesty at every rollout wave: reporting, near-misses and false declarations are safe; hiding is not.
- Recognise detection improvements, deleted alerts, closed detection gaps and well-written postmortems publicly, and reward the engineers who do them.
- Correct blame language in leadership communication on the spot, including inside the executive review.
- Watch for pager fatigue with real data, and take an engineer off rotation as soon as the intrusion cap breaks.
- Run a quarterly process retro with the whole engineering organisation and publish what changed as a result.
- Refresh the training and tabletop programme annually and immediately after any SEV1.
- Share wins publicly each quarter — faster detection, fewer customer-detected incidents, credits avoided — to hold momentum after the audit.
--- PROPOSAL 3 ---
Proposal ID: 265b9bc8-32de-46d3-8546-6dd406da4c24
Content:
Estimated Complexity: high
Success Metrics: - Median Time to Detect (MTTD) < 5 minutes.
- Median Time to Mitigate (MTTR) < 60 minutes for SEV1.
- >90% of incidents detected internally (vs. 40% previously).
- SLA credits reduced to <$100K annually.
- Monthly alert volume < 600 with >95% signal-to-noise ratio.
- Postmortem action item completion rate > 80%.
- Zero incidents with command ambiguity > 15 minutes.
- 100% of SEV1/2 postmortems published within 15 days.
- Status page updates within 30 minutes for 95% of SEV1.
- 100% of teams integrated into on-call rotations.
- On-call satisfaction score > 7/10.
- 40+ certified Incident Commanders available 24x7.
- SOC 2 Type II audit passed with zero findings.
Steps (22):
1. Executive Mandate & SOC 2 Evidence Clock
Establish clear ownership and funding to transform the CEO's concern into a governed program with immediate compliance implications.
2. Baseline Measurement & Problem Statement (depends on: 1)
Quantify the current state to build a defensible problem statement and establish the data foundation for improvement tracking.
3. Severity & Class Taxonomy (depends on: 2)
Define unambiguous incident levels based on business impact to trigger specific response protocols.
4. Incident Command Roles & Structure (depends on: 3)
Standardize roles and decision rights to eliminate command ambiguity during active incidents.
5. Tooling Consolidation & Integration (depends on: 1)
Deploy a unified incident management platform to replace fragmented alerting tools and create a single source of truth.
6. Detection Strategy & SLO Implementation (depends on: 3)
Shift monitoring focus to customer-impacting symptoms to detect issues before customers do.
7. Alert Quality Standards & Noise Reduction (depends on: 5, 6)
Implement strict alerting rules to reduce noise and restore on-call trust.
8. On-Call Architecture & Coverage Model (depends on: 4)
Design a federated, team-owned on-call structure supported by central command coverage.
9. On-Call Compensation & Wellbeing Policy (depends on: 8)
Introduce paid, sustainable on-call rotations to remove resistance and mitigate burnout.
10. Playbooks & Communication Templates (depends on: 3, 4)
Create standardized playbooks and templates for each severity level to reduce cognitive load.
11. Escalation & Incident Lifecycle Policy (depends on: 4, 5)
Define the mechanical progression from alert to resolution, including automatic escalation paths.
12. Communication Workflows (Internal & External) (depends on: 4, 10)
Establish timed protocols for internal, customer, and regulatory communications.
13. Status Page & Customer Notification Automation (depends on: 5, 12)
Integrate incident management with public status pages to ensure timely and accurate updates.
14. Postmortem Policy & Blameless Framework (depends on: 3)
Mandate and standardize postmortems to focus on system improvement rather than individual error.
15. Action Item Tracking & Governance (depends on: 14)
Ensure postmortem action items are tracked, owned, and completed to prevent recurrence.
16. Metrics, Dashboards & Review Cadence (depends on: 2, 3)
Define and monitor key performance indicators to validate the effectiveness of the incident process.
17. Training, Certification & Drills (depends on: 4, 10, 12)
Build practical skills and readiness for incident response across all roles.
18. Pilot Program with Volunteer Teams (depends on: 5, 7, 9, 11, 13, 15, 17)
Test the end-to-end process on a small scale to refine procedures before full rollout.
19. Phased Rollout to All Teams (depends on: 18)
Deploy the process to all 28 teams in manageable waves to minimize disruption and manage resistance.
20. SOC 2 Control Mapping & Evidence Framework (depends on: 1, 3)
Align the incident management process with SOC 2 requirements to ensure audit readiness.
21. SOC 2 Dry Run & Audit Support (depends on: 19, 20)
Verify evidence sufficiency and prepare the team for the external audit.
22. Governance & Continuous Improvement (depends on: 19, 21)
Establish standing governance to sustain the process and drive ongoing improvements.
Instructions:
1. Carefully evaluate each proposal considering:
- Clarity and detail of the proposed steps.
- Overall alignment with the Main Objective.
- Quality and relevance of the proposed success metrics.
- Coherence of dependencies between steps (if any).
2. Think deeply about each proposal, analyzing its strengths and weaknesses. Consider how well each proposal addresses the main objective and which one would be most effective.
After your thorough analysis, give the Proposal ID of the selected proposal and a one- or two-sentence justification of your choice.
{"selected_proposal_id": "abac560b-e42c-412b-9ecc-a60d55c29045", "justification": "This proposal offers the most realistic and detailed roadmap by introducing a 'minimum viable process' to bridge the gap between current chaos and full implementation. It explicitly addresses the specific cultural friction regarding pager ownership through structural separation of duties and financial incentives, ensuring higher adoption rates."} |
Selected plan
1. Mandate, one owner, and the evidence clock
This step turns the CEO's email into a funded programme with a single accountable owner, and it starts the SOC 2 clock on day one.
- Appoint a Director of Incident Management reporting to the CTO, with a dotted line to the COO for customer and SLA matters and to the Head of Compliance for audit readiness.
- Publish a one-page charter: scope covers every customer-impacting, money-moving, security and data-integrity event across all 28 teams and both regions.
- Grant explicit authority to declare an incident, set severity, freeze deploys, page any engineer in the company, and approve customer messaging.
- Fund the envelope up front: tooling, training and drill time, on-call compensation, and a small programme team, roughly $500–700K a year against $1.3M in credits paid.
- State the return plainly to the steering group: credits avoided, churn avoided, and audit findings avoided.
Start the evidence clock now. A SOC 2 Type II report covers a period of operating effectiveness, so evidence cannot be backfilled; an audit in eight months means the operating period begins in the first weeks of this programme.
- Stand up a fortnightly steering group of the CTO, VP Engineering, Head of Support, Head of Compliance and two engineering managers.
- Open a programme risk register with the top risks, owners and review dates, and revisit it at every steering group.
- Make participation in the incident process a documented performance expectation for every engineering manager, not an optional extra.
2. Baseline evidence pack and cost-of-downtime model (after 1)
You cannot demonstrate improvement without a defensible starting point, and you cannot win the argument about noise without numbers.
- Build a 12-month incident register and re-classify all 31 customer-impacting incidents by detection source, duration, customers affected and credits paid.
- Build the silent-failure register: incidents where no internal alert fired at all, which is the number that explains the 40% customer-detected rate.
- Run an alert census per tool, per team and per service: volume, page-to-action ratio, off-hours interruptions, and the 50 noisiest rules with a named owner each.
- Survey every on-call engineer and manager on burden, fairness, escalation clarity and pay expectations, targeting above 70% response.
- Interview Support, Account Management and Sales about how customers learn of incidents and what they are promised.
- Reconstruct the two command-ambiguity incidents minute by minute to find exactly where ownership lapsed.
- Build a cost-of-downtime model: dollars per minute of impact per customer journey, used later to sequence teams and justify funding.
- Publish the pack internally as the problem statement, and retain every artifact as management-review evidence for the audit.
3. SOC 2 control mapping and evidence architecture (after 1)
Most programmes leave compliance to the end; this one maps controls in month one, because the mapping decides what the process must capture from day one.
- Map the process to the Trust Services Criteria for incident identification, evaluation, response, containment, recovery and communication (CC7.1–CC7.5, CC2.2, CC2.3, CC4.1, CC3.x).
- Write each control in plain language, with one named owner and its evidence artifact.
- Define the golden incident file: one auto-assembled, single-click export per incident containing timeline, roles, severity, communications, postmortem, actions and closure proof.
- Set retention, storage location and immutability so no control depends on a laptop, a private channel, or chat history that expires.
- Keep a gap register with owners and dates, reviewed fortnightly by the steering group.
- Meet the auditor's readiness team inside the first 90 days to test the control design before anything is built on top of it.
4. Severity and class taxonomy with the trigger matrix (after 2)
Severity decides how loud the response is; class decides which playbook runs and what must never be traded away.
- Define severity in customer and money terms: customers affected, dollars at risk, regions lost, ledger integrity, settlement delay, data loss and SLA exposure.
- Use SEV0 for security, privacy or regulatory events; SEV1 for total or material loss of a payment path; SEV2 for degradation or single-region loss; SEV3 for limited impact with a workaround; SEV4 for internal-only issues and near-misses.
- Define response classes: Availability, Performance, Data integrity and Ledger, Security and Privacy, Third-party dependency, Settlement and Financial, Process failure.
Class can raise the response but never lower it. A SEV2 data-integrity incident gets SEV1 posture, because integrity failures are not recoverable by moving faster.
- Specify automatic triggers: loss of one AWS region, ledger write failures, replication lag above threshold, a missed settlement cut-off, payment success rate below threshold for five minutes.
- State who may declare — any engineer, Support agent or account manager — and who may downgrade: the Incident Commander alone.
- Map every level to its SLA credit exposure and to the customer-visible status page state.
- Include worked examples from the last 12 months so each of the 28 teams recognises its own incidents, and revalidate the taxonomy quarterly against real declarations.
5. Day-one operating rules and the minimum viable process (after 1, 4)
The full process will take months; the first useful version must be live in two weeks using the tools that already exist.
- Publish ten day-one rules that need no procurement: a named owner within five minutes, one channel per incident, one register entry per incident, one person speaking to customers.
- Apply the ambiguity rule: if two responders disagree about whether this is an incident, it is an incident.
- Make declaring free: a false alarm is closed as a false declaration, tracked as a metric, and never criticised.
- Require a register entry within 24 hours for every customer-impacting incident, even a minimal one.
- Ban silent incidents: if we know, the customer hears it from us rather than from their own reconciliation.
- Run the first 30 days on manual command, with a rotating duty Incident Commander drawn from the 12 teams that already have on-call.
- Hold a 15-minute daily incident stand-up during month one to catch friction while it is still fresh.
6. Roles, command structure and the no-unowned-minute rule (after 4)
The two hour-long command failures did not happen at declaration; they happened in the gap before it, when an alert had fired and nobody owned it.
- Publish one-page role cards: Triage Owner, Incident Commander, Deputy IC, Operations Lead, Internal Comms Lead, Customer Comms Lead, Scribe, Subject-Matter Responders, and Executive Sponsor for SEV1 only.
- Introduce the Triage Owner rule: whoever acknowledges the page owns the incident until an IC takes over or the incident is stood down.
The IC owns the incident, not the fix, and does not debug. An IC who starts troubleshooting has abandoned command.
- Give the IC explicit decision rights: declare and escalate severity, freeze changes, halt deploys, approve customer messaging, and pull in any engineer.
- Define minimum viable staffing per severity: SEV1 fills every role; SEV2 staffs IC, scribe, comms and responders; SEV3 staffs an IC and a scribe.
- Set handover discipline: four-hour maximum IC shifts on SEV1 with a written handover, and a deputy named within 15 minutes of declaration.
- Set responder behaviour: one channel, one bridge, no side channels, and every request phrased with a named owner and a time.
- Link the role cards from every paging notification so they are one tap away at 3 AM.
7. Lifecycle, declaration and escalation policy (after 4, 6)
This step defines the mechanical path from an alert to a declared incident and back to normal service, removing judgment calls from the worst moments.
- Define states with entry and exit criteria: Detected, Triaged, Declared, Mitigated, Resolved, Postmortem, Closed, plus a Watch state with a hard 30-minute timer.
- Set targets: page acknowledged within five minutes, triage decision within fifteen, declaration within thirty for anything customer-facing.
- Build automatic escalation ladders with timers at every layer: responder, service owner, team manager, IC on-call, VP Engineering.
- Define the unresponsive-team path: fifteen minutes escalates to the team manager, thirty minutes to the IC, who may then direct any available engineer.
- Define change freeze and rollback authority during SEV1 and SEV2, with the single named condition that lifts the freeze.
- Enforce one incident, one record, with the timeline captured automatically from the channel and bridge rather than written from memory afterwards.
- Add the two-simultaneous-incidents rule: a second IC is designated, and no IC ever runs two incidents at once.
- Test every escalation path weekly with synthetic pages, and adjust the timings after the first month of real operation.
8. On-call architecture across 28 teams (after 4, 6)
The objection is that engineers will not carry a pager for another team's code; the answer is to build rotations that make the objection structurally impossible.
- Run a Service On-Call rotation per team, covering only that team's own services.
- Run a Platform Duty rotation for genuinely shared infrastructure: the PostgreSQL ledger cluster, Kubernetes, networking, CI/CD and observability.
- Run a central Incident Commander roster of 12–16 certified senior engineers on one-week shifts with a primary and a secondary.
State the consequence honestly. Sixteen of 28 teams have no rotation today; each must build one or formally transfer service ownership to a team that has one, with the transfer dated and recorded.
- Merge small or low-traffic teams into shared rotations so no rotation has fewer than six engineers.
- Cap load in the scheduling tool: no engineer on call more than two weeks per quarter, enforced by configuration rather than negotiation.
- Publish a coverage matrix of all 28 teams showing services, owners, rotation size and gaps, reviewed monthly.
- Write on-call participation into engineering job levels and hiring criteria, so the commitment is set before someone joins.
9. Compensation, rest and the economics of opting out (after 8)
Unpaid on-call is the most cited reason for resistance, so settle compensation before rollout, not during it.
- Move to paid on-call: a per-shift stipend benchmarked to the New York market, agreed with HR and Finance, published with an effective date before any team is asked to join a rotation.
- Pay for callouts at 1.5× the hourly rate for time actually spent mitigating, with a minimum block per interruption.
- Provide documented compensatory rest: no normal working day after a night incident, and the rest day is policy rather than a favour granted by a manager.
- Cap intrusion with a maximum number of off-hours pages per shift, and trigger a mandatory review whenever the cap is exceeded.
Allow opt-out, but put a price on it. An engineer may step out of a rotation, and their team buys coverage from the paid pool at a published internal rate, which turns a cultural argument into a visible budget decision.
- Publish amnesty: incident records, near-misses and false declarations are never used in performance reviews; only failure to report is a performance issue.
- Protect against gaming: pay attaches to shifts and callouts, never to incident counts.
- Check the New York labour, overtime and tax treatment with Legal and Finance before announcing, and review the policy every six months against real page volumes, attrition and survey results.
10. Detection strategy: journeys, synthetic signals and customer-report intake (after 4)
Customers detected 40% of incidents first, which makes detection the highest-leverage business problem in this programme.
- Define SLIs and SLOs for the top 20 customer journeys, measured per region: payment initiation, authorisation, settlement, ledger read and write, API availability, webhook delivery and payout.
- Alert on symptoms against those SLOs, not on cause-based infrastructure metrics.
- Add synthetic transaction monitoring from outside AWS in both regions plus a third location, on a one-minute cadence for money-moving paths.
- Add ledger and shared-PostgreSQL signals: replication lag, connection saturation, write latency, lock waits, transaction ID exhaustion, checkpoint pressure and bloat.
- Open a customer-report intake so Support and account managers can raise an incident directly, and count that path as a detection source in all reporting.
- Apply the detection-gap rule: whenever a customer reports first, a detection-gap ticket is opened automatically and owned by the responsible team.
- Publish a detection contract per service — owner, at least one symptom alert, documented expected detect time — for all 180 services.
- Run a detection drill per team: break something in staging and see whether it pages before a human notices.
11. Paging contract and the noise-reduction programme (after 2, 10)
3,400 alerts a month at 85% noise is the reason engineers resent the pager, and fixing it is the price of admission for everything else in this plan.
- Publish the paging contract: every page must be symptom-based, actionable, owned, mapped to a severity and class, and linked to a runbook.
No runbook, no page, enforced by a CI check on the alert definition itself.
- Separate paging from ticketing: only customer-impacting or imminent-impact signals may wake a human.
- Set a page budget per team and per service, with a remediation ticket opened automatically, owned by the engineering manager, when the budget is breached.
- Put new alerts on two-week probation as ticket-only until they have proved actionable.
- Define the default action for a noisy alert: fix, tune or delete within ten working days, with deletion treated as a legitimate and celebrated outcome and the deleted count published.
- Deduplicate and correlate at the ingest pipeline so one root cause produces one page instead of forty.
- Require an expiry date on every silencing rule and temporary threshold, so suppression cannot quietly become permanent blindness.
- Run a 90-day noise sprint with a public burn-down of the top 100 noisiest rules, each owned by a named manager.
12. Incident tooling consolidation and the golden incident file (after 3, 6, 10, 11)
Six alerting tools and no single incident record are structural causes of the 22-minute detection and the three-hour mitigation.
- Choose one incident platform for paging, schedules, escalation policies, incident records, postmortem workflow and status-page integration, and choose it within four weeks.
- Consolidate the six alert sources into a single event pipeline feeding that platform, with deduplication, correlation and severity and class mapping applied at ingest.
- Integrate observability, runbooks, chat and voice bridges so responders see dashboards and runbooks inside the incident record, and the timeline is captured automatically.
- Implement the golden incident file export defined in S3, so audit evidence is one click rather than a reconstruction.
- Run a dual-run period alongside the legacy tools with a defined rollback, then switch the legacy tools off on a published cutover date.
- Host the status page outside the production failure domain so it survives a total platform outage, and prove that in a game day.
- Make the platform usable from a phone: an Incident Commander must be able to run a SEV1 from a mobile device at 3 AM.
13. Internal, customer and regulator communications (after 6, 7, 12)
Today the status page is written by whoever is around; this step replaces improvisation with a clock, a named owner and pre-cleared templates.
- Set the internal cadence: first update within 15 minutes of declaration, then every 30 minutes for SEV1 and hourly for SEV2, whether or not there is progress.
- Never let an employee learn of an incident from the status page: internal communication leads, external follows.
- Set the customer cadence: status page updated within 30 minutes of a SEV1 and 60 minutes of a SEV2, with a no-new-information update still mandatory.
- Define the channel hierarchy: status page for all 2,100 customers, direct email to affected customers on SEV1, and named account-manager calls for the top 50 accounts.
- Pre-approve templates per severity and class with Legal and Compliance, each carrying its own next-update time.
- Forbid speculation: customer messages never guess at cause, never assign blame, and never commit to a root cause before the postmortem.
- Define the closing communication: resolution notice, how SLA credits will be handled, and the committed date for the written report.
- Build a regulator clock matrix covering event type, regulator, notification window, signer and the shortest applicable clock, including NYDFS Part 500, money-transmitter and banking notifications, breach notification, card-network rules and public-company disclosure.
The regulatory clock starts at awareness, not at root cause. Route every notification through Compliance, never Engineering, and pre-clear the templates.
14. Customer trust workstream and the SLA credit ledger (after 13)
The $1.3M in credits is a symptom of having no single record of customer impact, and the CEO's inbox is a symptom of customers learning things late.
- Maintain one durable customer-impact record per incident: which customers, which journeys, from when to when, and the estimated credit.
- Use that one record for communications, credit computation, regulatory reporting and the annual review, so nobody reconciles two versions of the same outage.
- Automate the credit path end to end — computation, approval, customer notification and finance treatment — so credits stop being a manual scramble three weeks later.
- Reconcile accrued against paid credits monthly and report the result in the executive review.
- Give the status page a named product owner and map its components to customer journeys, not to internal services.
- Send a CTO-signed reliability note to the top 50 accounts and publish a quarterly reliability report to all customers.
- Give account managers a script of the facts they may state, the speculation they may not, and a path for customer escalations.
- Track credit avoidance against programme cost, so the funding case stays a number rather than an argument.
15. Postmortems: mandatory set, three levels, blameless by design (after 6)
Postmortems currently happen for some incidents, in various formats; this step makes them mandatory where they matter and light where they do not.
- Make postmortems mandatory for every SEV0, SEV1 and SEV2, every SEV3 with customer impact or a repeat pattern, every near-miss the IC flags, and every incident where the process itself failed.
- Use three levels so the ritual matches the weight: a lightweight async review for SEV3 and SEV4, a facilitated postmortem for SEV2, and a full review with an executive sponsor for SEV1.
- Set deadlines: draft within five business days, blameless review within ten, internal publication within fifteen.
- Adopt one template: impact, timeline, detection, response, contributing factors across tooling, process, organisation and human factors, what went well, what went badly, and actions.
- Train a pool of blameless facilitators and require a trained one for every SEV1 review, never the IC.
Ban blame language in the template and ban "human error" as a root cause. The question is always what system condition made the error possible.
- Publish all postmortems internally by default, with security review only for genuinely sensitive material.
- Produce a customer-facing root-cause report for SEV1 incidents, especially for regulated and top-tier accounts.
16. Action items: capped, verifiable, with the repeat-incident rule (after 15)
Eleven of 64 action items closed is not a tracking problem; it is a generation problem, because the process produces more actions than the organisation can absorb.
- Cap each postmortem at three action items, with anything beyond that going into a ranked reliability backlog.
- Require every action to have a named human owner, a due date, and a definition of done that is an artifact: a merged change, an alert that fired in a drill, or a test that fails without the fix.
- Prohibit self-reported closure; closure requires the artifact and sign-off by the process owner or the Incident Commander.
- Keep one reliability backlog in the engineering tracker with a mandatory label and a link back to the originating incident.
- Reserve a fixed share of each team's sprint capacity for reliability work, with unspent capacity visible to the vice presidents.
- Apply the repeat-incident rule: a second incident with the same contributing factor triggers a design review owned by an architect, not another action item.
- Run a weekly ageing review of open actions and escalate anything more than 30 days overdue to the VP of Engineering.
- Report action completion rate and median action age monthly, by team.
17. Metrics, dashboards and the review cadence (after 2, 4, 16)
Define what good looks like, then measure it in a way that rewards reporting incidents rather than hiding them.
- Define outcome metrics: time to detect by source, time to mitigate by severity and class, share of incidents detected by customers, incidents by severity and class, SLA credits paid, and customer-impact minutes.
- Define process metrics, because the process itself has SLOs: IC assigned within five minutes, page acknowledgement rate, first-update timeliness, update-cadence adherence, postmortem timeliness, action closure rate, action median age, and IC roster coverage.
- Define health metrics: alert volume and noise ratio per team, off-hours pages per engineer, on-call satisfaction, training completion, and share of services with a detection contract.
Never publish incident count as a team metric. It rewards hiding incidents; publish reporting metrics instead — near-misses filed, detection gaps closed, false declarations made.
- Publish live dashboards visible to every engineer, refreshed daily, with every metric baselined against the S2 evidence pack and given a 90-day and a 12-month target.
- Run a 30-minute weekly operational review incident by incident, a monthly reliability review of trends and actions, and a quarterly executive review with the CEO's office.
- Hold a quarterly review of the process itself: what wasted time, what confused responders, and what should be deleted.
- End every review with decisions and named owners, never with numbers alone.
18. Training, certification and the drill programme (after 6, 7, 13, 15)
A process that lives only on a wiki page fails on the first real page, so skills are built and tested before they are needed.
- Build a practical curriculum: how to be on call, how to triage and declare, how to run an incident as IC, how to communicate, and how to write a postmortem.
- Require certification before joining the IC roster: a written assessment plus a live simulated incident.
- Certify at least two ICs per team group so the central roster has depth across all 28 teams and no holiday week is left uncovered.
- Train communications leads separately on templates, cadences, customer language and the regulatory rules.
- Run monthly tabletops drawn from the last 12 months of real incidents, including region loss, ledger write failure and a missed settlement cut-off.
- Run quarterly game days with deliberately injected failure, including PostgreSQL failover, status-page outage and alerting-pipeline outage.
- Drill the process's own failure modes, not just technical ones: IC unreachable, comms lead on leave, two simultaneous SEV1s, a paging storm, and a false alarm that burns an hour.
- Audit the process for single points of failure: who alone can perform each critical task, and what happens in their holiday week.
- Keep a mandatory onboarding module for every engineer joining or transferring in, with audit-ready completion records.
19. Pilot with three to four teams, using real incidents (after 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 18)
Do not roll out to 28 teams untested, and do not rely on drills alone when real incidents are better training material.
- Recruit three to four volunteer teams spanning criticality: a payment-path team, a ledger-adjacent team, the platform team that owns the shared PostgreSQL cluster, and a low-traffic team.
- Run the entire process end to end in the pilot: severity and class, roles, escalation, communications, postmortems, action tracking, and paid on-call.
- Treat real incidents during the pilot as the primary training material, and hold a retro within 48 hours of each one, run by the process owner while the friction is fresh.
- Instrument the pilot against the baseline and publish before-and-after numbers.
- Hold weekly retrospectives with the pilot teams, and change the written policies, the tooling and the training in response, documenting what changed and why.
- Set explicit exit criteria: rotation coverage achieved, zero unacknowledged pages over two weeks, postmortems delivered on time, and actions tracked to closure.
- Produce a pilot report that every later rollout conversation starts from.
20. Phased rollout to 28 teams, sequenced by cost of failure (after 16, 18, 19)
Rollout is a staged migration with readiness gates, not an email announcement, and the sequencing matters more than the schedule.
- Split the 28 teams into four waves of roughly seven, ordered by incident density and customer-journey ownership: the teams whose incidents cost the most SLA credits go first, because their improvement is the visible proof.
- Leave three weeks between waves and define a per-team readiness checklist: services mapped and owned, alerts cleaned to the paging contract, runbooks written, rotation staffed to at least six engineers, training complete, manager briefed, compensation in effect.
- Hold a gate review with the process owner before each team joins, and move an unready team to the next wave with a dated remediation plan rather than granting an exception.
- Give each wave a named champion and explain the why using the pilot's numbers.
- Handle resistance directly and early: publish the own-your-code-own-your-pager rule and the paid on-call mechanics before each wave, never after.
- Hold a retro within 48 hours of each wave's first incident under the new process, and push accepted changes through change control.
- Retire legacy tools, informal escalation lists and the ad-hoc status page process at the end of each wave on a published cutover date, so two processes never run in parallel for long.
- Sequence the waves so no team faces both a rollout and the audit dry run in the same week.
21. Audit dry run and evidence review (after 3, 20)
This step converts a good process into a provable one, about six weeks before the auditors arrive.
- Sample real incidents from the pilot and the early waves against each control's evidence requirements.
- Remediate every gap found and re-test the remediated control against the same sample, prioritising anything that risks a qualified opinion.
- Brief every engineer who may be interviewed so they can describe the process as they actually practise it, not as it is written.
- Prepare the auditor package: process documentation, sample incident records, paging logs, communication logs, training records, on-call schedules, the action register and alert quality metrics.
- Designate one audit liaison and a small evidence-request team, so requests do not land on all 28 teams at once.
- Rehearse the walkthrough with an Incident Commander and a communications lead, because auditors probe realism under pressure.
- Keep the audit liaison and the process owner as close to the same person as possible, so accountability for the control is also accountability for the evidence.
22. Standing governance and process ownership (after 20, 21)
The classic post-audit failure is that the process freezes and then decays, so ownership has to outlive the programme.
- Establish a standing Incident Management Council chaired by the process owner, meeting monthly with engineering, support, compliance and product representation.
- Give the process owner a documented mandate to change standards, with formal change control for any change to severity, classes, roles, communication timings or compensation.
- Revalidate the severity and class taxonomy quarterly against real declarations, and re-baseline the entire metric set annually.
- Report quarterly to the executive team on the metric set and the top five systemic causes of incidents.
- Feed incident themes into architecture review, release engineering and hiring, so the programme improves the system and not only the response.
- Keep a public backlog of process improvement ideas so any engineer can propose a change, and act on it every quarter.
23. Resilience track for the shared ledger and multi-region blast radius (after 1, 17)
Better incident management cannot save a single shared ledger during a corruption event, so this runs as a funded engineering track beside the process, not inside it.
- Reduce ledger blast radius: read replicas for read paths, connection-pool isolation per service domain, statement timeouts, write throttles and per-tenant limits.
- Build a ledger protection mode: an automatic tripwire that degrades to read-only or queues writes before corruption can spread.
- Rehearse multi-region failover with the ledger as the binding constraint, and set a measured recovery time objective per region.
- Test restore and point-in-time recovery on a schedule, and publish the timings.
- Tighten deploy safety: progressive rollout, automatic rollback on SLO burn, and change-freeze windows around settlement.
- Track a small number of resilience bets with named owners and quarterly milestones, reviewed in the same executive review as the incident metrics.
24. Culture, incentives and continuous improvement (after 20, 22)
Incident management is a cultural change before it is a process change, and the incentives decide whether reporting goes up or down.
- Restate the amnesty at every rollout wave: reporting, near-misses and false declarations are safe; hiding is not.
- Recognise detection improvements, deleted alerts, closed detection gaps and well-written postmortems publicly, and reward the engineers who do them.
- Correct blame language in leadership communication on the spot, including inside the executive review.
- Watch for pager fatigue with real data, and take an engineer off rotation as soon as the intrusion cap breaks.
- Run a quarterly process retro with the whole engineering organisation and publish what changed as a result.
- Refresh the training and tabletop programme annually and immediately after any SEV1.
- Share wins publicly each quarter — faster detection, fewer customer-detected incidents, credits avoided — to hold momentum after the audit.
- Median time to detect falls from 22 minutes to under 8 minutes by month 6, and under 5 minutes for money-moving journeys by month 9.
- Share of customer-impacting incidents first detected by customers falls from 40% to below 10% by month 9.
- Median time to mitigate for SEV1 falls from 3 h 10 min to under 60 minutes by month 9.
- 100% of SEV1 and SEV2 incidents have a named Triage Owner within 5 minutes and a named Incident Commander within 15 minutes, from month 4.
- Zero incidents with unclear command authority lasting more than 15 minutes, measured monthly from month 4.
- Monthly page volume falls from 3,400 to under 600, with a false-positive rate below 15%, by month 5.
- No service exceeds two pages per on-call shift for three consecutive months by month 6.
- 100% of the 180 services have a named owner, a detection contract and at least one symptom-based alert by month 7.
- All 28 teams have a documented rotation, a service ownership map and at least one trained on-call engineer by month 6.
- 100% of on-call shifts are paid under a published policy from month 2, with zero on-call-attributed voluntary attrition by month 6.
- On-call satisfaction scores 7 out of 10 or better in quarterly surveys from month 6.
- The IC roster holds at least 12 certified ICs covering 24x7 with no uncovered week from month 4.
- 100% of SEV0, SEV1 and SEV2 postmortems are published internally within 15 business days from month 5.
- Action items closed within 60 days rise from 17% to above 90%, with a median action age under 30 days, by month 6.
- Status page first update is posted within 30 minutes on at least 95% of SEV1 incidents from month 4.
- Zero missed regulatory notification windows on any incident requiring notification.
- The customer-impact record is complete for 100% of customer-impacting incidents from month 5, with SLA credits reconciled monthly.
- Annual SLA credits paid fall at least 50% from the $1.3M baseline within 12 months.
- Customer-impacting incidents fall at least 40% year over year from the 31-incident baseline.
- Review cadence is sustained: weekly operational review in at least 90% of weeks, 12 of 12 monthly reliability reviews, 4 of 4 quarterly executive reviews.
- Every page has a recorded disposition — fixed, tuned or deleted — within 10 working days, from month 4.
- SOC 2 Type II is passed with zero findings related to incident response.
- 100% of new engineers complete the incident-response onboarding module within 30 days of joining.
Evolution analysis
The analysis in brief
- Outcome: P2 wins, and I agree. The vote landed on P2 (two votes, one dissent for P1), citing the same structural reasons I did: the two-week minimum viable process (S5), the funded ledger-resilience track (S23) and the explicit handling of "pager for other teams' code".
- The final plan beats every round-0 proposal. New gains: an interim process usable in two weeks, the ownership gap closed by name (Triage Owner, Watch state with 30-minute timer, ambiguity rule, two-simultaneous-SEV-1 rule), SOC 2 control mapping and the golden incident file moved to month one, prevention funded as a track, and honest economics (priced opt-out, credit avoidance vs programme cost, pay detached from incident counts).
- Convergence was real on the merits but executed as imitation. P2's round-0 architecture (baseline evidence pack, SLO/synthetic detection, federated ownership, TSC mapping plus dry run) was genuinely strongest, but P1 and P3 rebuilt on it step-for-step and P1 abandoned its hybrid IC-pool model and per-incident bonus without a sentence of argument.
- Criticism existed only as silent non-adoption. No agent ever named a rival's flaw; disagreement showed up as P2 banning incident-count-based pay, nobody copying P3's action-items-in-performance-reviews, and P2 quietly cutting its own "40+ certified ICs" to 12–16 — which P3 then carried into the final round anyway.
- P3 regressed to a table of contents. Round 2 was 22 bare titles: severity thresholds, status-page timings, compensation mechanics, wave schedule and curriculum all deleted, making it unexecutable and correctly ranked last by everyone.
- Biggest problem: nobody audited the arithmetic. ~28 service rotations with primary and secondary is ~56 engineers on call at $600–1,000/week, roughly $1.8–2.9M/year against stated envelopes of $500–800K — above the $1.3M in credits justifying the programme — and the six-engineer rotation floor was never reconciled with ~9 engineers per team.
- Second problem: no quality gate and unchecked internal consistency. P1's 3-minute status-page rule contradicts its own 30-minute metric for 95% of SEV-1s, its alert (<400 vs <600) and MTTM (45 vs 60 min) targets disagree with its steps, and its step 11 is truncated mid-sentence; the SOC 2 Type II observation-window start date was never established.
- Fixes that would help most: (1) a mandatory objections block — two quoted defects per rival plan plus a one-sentence justification for every borrowed mechanism; (2) a numbers-auditor round recomputing on-call cost, rotation headcount and every metric-to-step date; (3) a forced scenario walkthrough (SEV-1 ledger corruption at 02:00 on settlement day, IC unreachable, status page down) plus a pruning round with a hard cap of ~15 steps to counter additive drift.
At a glance
- Analyst's first choice, blind to the vote: proposal 2 · deepseek-flash_refine_2
- The vote: proposal 2 · deepseek-flash_refine_2 — the analyst agrees
- Against the initial proposals: better than every initial proposal
- Rounds: round 1 converging · round 2 converging
- The process: 10 problems observed, 8 suggestions
The final round, in the analyst's words
The final round offers two near-identical heavyweight plans (P1 and P2) built on the same architecture — evidence clock and SOC 2 control mapping in month one, baseline register, severity×class taxonomy, Triage Owner rule, three on-call rotations (Service/Platform/IC), paid on-call with priced opt-out, SLO and synthetic detection, paging contract, capped and verifiable postmortem actions, pilot then four gated waves, dry run, standing council — plus P3, which submitted 22 step titles with no content. P2 is distinguished by a two-week minimum viable process, a funded ledger-resilience track, a cost-of-downtime model and internally consistent metrics; P1 is distinguished by sharper numeric thresholds but carries several self-contradictions and one truncated step. P3 is not executable as written.
Raw prompts and responses of every analysis call
[ROUND 0]
[SYSTEM]
You are an expert reviewer of multi-agent planning processes.
Several LLM agents drafted plans for a task, refined them over a number of rounds while seeing each other's proposals, and finally voted for the best one.
Be exhaustive but precise: name concrete steps, ideas and metrics, never generalities. Judge plans by their fitness for the task as stated, their realism, their completeness, the soundness of their order and dependencies, how measurable their success is and how they handle things going wrong.
You are an impartial evaluator, not a chronicler: assess the proposals and the process on their merits, never rationalise what happened or assume that the outcome was right.
After your analysis, answer in the requested structure.
Every text field you write will be read by a busy person who skims. Make it easy to skim: short sentences and short paragraphs; when you name several things, prefer a list to a paragraph, with sub-items when an item has parts, but keep a single fact as a sentence; lead with the point and then the evidence; name proposals and steps by number (P2, step 4); no preamble, no repetition of the question, no closing summary; bold at most one key phrase per item or paragraph. Text fields accept Markdown: a blank line between paragraphs, "- " for lists, **bold**.
[HUMAN]
Task given to the agents: "A B2B payments platform in New York (260 engineers in 28 teams, 2,100 customers, $4B processed a month, a 99.95% SLA with service credits) runs about 180 services on Kubernetes across two AWS regions, with a shared PostgreSQL cluster for the ledger. In the last 12 months: 31 customer-impacting incidents, median time to detect 22 minutes (customers detected 40% of them first), median time to mitigate 3 h 10 min, $1.3M paid in SLA credits, two incidents in which nobody was sure who was in charge for over an hour, and a CEO email about "outages we hear about from clients". On-call exists in 12 of the 28 teams, unpaid, fed by alerts from six different tools — 3,400 alerts a month, 85% of them noise; there is no severity scale, status-page updates are written by whoever is around, and postmortems happen for some incidents, in various formats, with action items that are rarely tracked (11 of 64 closed). A SOC 2 Type II audit in eight months will test incident response. Engineers push back against "carrying a pager for other teams' code".
Define the incident management process: severity levels and what each one triggers; roles (incident commander, communications, scribe, subject-matter responders) and how they are staffed 24x7 across 28 teams; the on-call structure, rotations, compensation and the rules for alert quality; detection and escalation paths; internal and customer communications (status page, account managers, regulators when required) with their timings; postmortems (when mandatory, format, blameless review, ownership and tracking of actions); the metrics and reviews that show whether it works; and how it is introduced across the teams without waiting for the audit."
This is round 0: every agent wrote its plan independently, without seeing the others.
--- PROPOSAL 1 (agent claudeHaiku4.5_initial_1, anthropic/claude-haiku-4-5) ---
Estimated complexity: high
Success metrics: - Median time to detect reduced from 22 minutes to <8 minutes within 6 months
- Customer-detected incidents drop from 40% to <5% within 6 months
- Median time to mitigate (MTTR) reduced from 3h 10min to <45 minutes for SEV-1 incidents within 6 months
- Annual SLA credits decrease from $1.3M to <$100k within 12 months
- Alert noise reduced from 3,400 per month (85% false positive) to <400 per month (>95% signal) within 3 months
- Zero incidents with command-and-control ambiguity (>1 hour without clear IC) within 2 months
- Postmortem action item completion rate reaches >80% (from 17%) within 4 months
- On-call satisfaction score (survey) reaches >7/10 for on-call engineers within 3 months
- All 28 teams integrated into incident management system with active on-call rotations by week 20
- SOC 2 Type II audit passes incident response controls with no findings 8 months from start
- Incident commander certification: 100% of active ICs trained and drilled within 2 months
- Monthly incident review meeting established and attended by leadership; trends documented
- New incident system integration complete: single alert tool, single dashboard, all 180 services feeding in, <5 min deployment
Steps (21):
1. Define severity levels and decision criteria
Create a four-tier severity framework (SEV-1 through SEV-4) that guides all downstream decisions about response, escalation, and communications.
Each level must specify: customer impact (revenue at risk, customers affected, data loss risk); financial threshold triggering service credits; whether an incident commander is required; response time SLA (e.g., SEV-1 < 5 min notification, SEV-4 < 2 hours); and the go-live decision tree (when to declare and when to resolve).
- SEV-1: Complete service down or critical path broken for >5% of customers; every minute costs money; IC required; 99.99% uptime threatened
- SEV-2: Significant degradation, features unavailable, affecting 1–5% of customers; IC typically required
- SEV-3: Minor impact, limited customer footprint or workaround exists; escalation path but not automatic IC
- SEV-4: Observations or minor issues; alert-driven, no escalation unless pattern emerges
2. Define incident roles and responsibilities (depends on: 1)
Create the organizational roles that operate during an incident: who is in charge, who talks to customers, who writes down what happened, who fixes the system, and how decisions are made under pressure.
Each role must have a single-sentence mission, decision authority, and escalation upward.
- Incident Commander: owns decision-making and timeline; declares severity; resolves conflicts; may or may not be technical
- Deputy IC: shadow to IC and takes over if IC becomes unavailable
- Communications Lead: writes status page, notifies account teams, manages customer perception
- Scribe: records decisions, who did what, key timestamps; not responsible for fixing
- SME Responders: engineers with context on the failing service(s); take IC's direction without debate
3. Design 24x7 on-call rotation structure (depends on: 2)
Build a rotation model that covers all 28 teams with primary and backup on-call engineers every hour across weekdays, evenings, weekends, and holidays; addresses the pager-carrying resistance.
Key design decisions: Is coverage per-team (each team owns its services) or pooled (shared responder pool handles anything)? How many people per rotation? How long are shifts (one week, two weeks)? When can engineers opt out without leaving the team exposed? Which roles are on-call (IC, communications, SME)?
- Recommend: dedicated IC pool (4–6 people in fast rotation) + per-team SME on-call for each team's own services
- Recommend: two-week rotation blocks to reduce handoff friction
- Recommend: one primary, one secondary per slot; secondary handles during primary's escalation
- Provide swaps, blackout dates, and a rule that no engineer is on-call more than 2 weeks per quarter
4. Define on-call compensation and incentives (depends on: 3)
Create a pay model that makes on-call acceptable and rewards engineers who carry the pager; ties compensation to real business risk.
- Base on-call stipend: e.g., $500–1,000 per week while on-call (regardless of incidents)
- Callback pay: 1.5× hourly rate for time spent mitigating incidents during off-hours
- Incident bonus: $50–100 extra per SEV-1 or SEV-2 incident mitigated (recognition)
- Comp time: full business day off after an incident that required >2 hours mitigation during night/weekend
- Annual bonus tie-in: 10–20% bonus multiplier for flawless on-call reviews
- Communicate: position as investment in reliability, not punishment for being online
5. Consolidate alert routing infrastructure
Replace six alert tools with a single ingestion and routing system; stop engineers from being woken by duplicate alerts, and make escalation automated instead of manual.
Evaluate existing tools (likely candidates: PagerDuty, Opsgenie, or Incident.io) or build a lightweight wrapper. The system must: accept alerts from all 180 services; deduplicate and correlate (same outage, different monitoring source); route to correct on-call engineer; expose an API for playbook automation; log every alert for postmortem analysis.
- Choose tool by week 2 of S5
- Migrate alerting endpoints from 6 sources to 1 by week 4
- Set up audit trail and retention
- Ensure mobile app works (on-call engineers need to engage from phone)
6. Build alert quality rules to cut noise (depends on: 1, 5)
Implement rules that automatically suppress the 85% of alerts that are noise (flapping, transient errors, auto-recovered conditions). Target: <400 actionable alerts per month.
Rules to implement: suppress alerts if service auto-recovered within 30 seconds; deduplicate same alert from multiple monitoring sources; suppress alerts for known maintenance windows; group flapping alerts (same service, >5 occurrences in 2 minutes) into a single page to on-call; rate-limit alerts from noisy services (e.g., max 1 alert per 5 minutes per service until silence clears).
- Audit existing 3,400 alerts per month: which are true signals, which are noise
- Tag each alert source with severity level (S1, S2, S3, S4 from S1)
- Create exceptions list: services known to be noisy, require different rules
- Weekly review: alert teams that trigger >50 alerts per week for reduction strategies
7. Implement automated detection and escalation paths (depends on: 1, 5, 6)
Wire the alert system to automatically escalate based on time or severity; removes the need for manual judgment calls during chaos.
Logic: SEV-1 alert arrives → IC notified instantly via phone call + SMS + Slack + mobile; if IC does not acknowledge within 2 minutes, page deputy IC; Communications Lead pinged simultaneously. SEV-2: on-call SME for that service + IC notify via Slack and mobile, escalate to IC's manager if not acknowledged in 10 min. SEV-3/4: on-call SME only, escalate after 30 min.
- Implement in alert routing system (S5)
- Test all paths weekly via synthetic page to on-call
- Track escalation metrics: how many pages reach secondary, how many hit manager
- Adjust timing based on first month of operations
8. Build incident dashboard and status tracking (depends on: 5)
Create a single source of truth during an incident that every responder sees in real time: who is on-call, incident timeline, who said what, current status, next steps.
Dashboard displays: active incidents and their severity; who is the IC and communications lead; timeline of all events (alert fired, IC assigned, customer notified, mitigation started, resolved); Slack channel and mobile notification status; on-call rosters (who is on-call right now for each team); postmortem link as soon as incident closes.
- Integrate with alert tool (S5) to auto-populate incident creation and initial severity
- Push updates to status page and customer account managers automatically
- Log all timeline entries for audit and postmortem completeness
- Mobile-optimized so IC can work from any device
9. Write incident playbooks for each severity (depends on: 1, 2, 3)
Create a one-page (or one-screen) reference for the IC and SMEs during an incident; sequences the steps and removes ambiguity.
Each severity level gets its own playbook: who gets paged (roles, order); first questions to ask (is it real, how big, who knows); what the IC should declare in first message (status page text, account manager notification, regulatory trigger); how long before escalating to executive team; decision rules for going dark vs. continuing to update customers.
- SEV-1 playbook: immediate IC + comms + CTO notification; customer status every 5 minutes
- SEV-2 playbook: IC + comms + tech lead notification; status every 15 minutes
- SEV-3 playbook: on-call SME + comms if customer-visible; status every 30 min or as resolved
- SEV-4 playbook: on-call SME only; update customers only if promised SLA is at risk
- Include decision trees: is this SEV-1 or SEV-2? Is it our code or dependency? Escalate or containment?
10. Define internal communication workflows (depends on: 2, 3)
Specify who informs whom, in what order, via what channel (call, Slack, email) during an incident; prevents gaps like "nobody knew who was in charge for an hour."
Workflow for SEV-1: IC assigned → IC calls CTO/VP Eng and incident channel lead within 1 minute; incident declared in #incidents Slack channel with severity, IC name, service affected; SME on-call for that service joins call automatically; IC pushes updates to #incidents every 5 minutes or when material change occurs. For SEV-2: IC notifies team leads via Slack, updates #incidents every 15 min. Define escalation: if IC is unreachable, deputy IC takes over and announces it.
- Create a phone tree or on-call list accessible to responders
- Set expectations: "If you don't hear from IC in 2 minutes, call them"
- Use a single incident Slack channel per incident (auto-created by incident tool)
- Log all comms in the incident dashboard for postmortem review
11. Design customer communication and status page process (depends on: 1, 2)
Plan when and how to inform customers, account managers, and regulators; ensure 2,100 customers are not learning about outages from Twitter before you tell them.
Rules by severity: SEV-1 detected → status page updated within 3 minutes (even if root cause unknown; post "Investigating"); account managers of affected customers called within 5 minutes; regulatory notification (if payment processing down) queued for approval; customer email within 10 minutes with ETA for next update. SEV-2: status page within 10 min, account managers called within 15 min, email if affecting >10 customers. SEV-3/4: no customer communication unless SLA at risk.
- Empower Communications Lead to update status page without IC approval if delay >3 min
- Prepare templated messages for common scenarios (database failover, data pipeline stuck, service crashed)
- Route regulatory notifications through legal/compliance; don't wait for perfect root cause
- Track customer impact in real time: how many customers affected by severity
12. Establish blameless postmortem process and format (depends on: 1, 2)
Build a systematic way to learn from incidents so the same failure does not happen twice; counter the fear that admitting a mistake leads to being blamed.
Mandatory postmortems: all SEV-1 and SEV-2 incidents, within 48 hours of resolution. Optional but encouraged: SEV-3 if interesting or if >3 of same type in 30 days. Format: what was the user-visible impact and for how long; what was the root cause (not "human error" but the system condition that made error possible); timeline of discovery and response; action items with owner and deadline; blameless tone (focus on process and system design, not individual mistakes).
- Assign a facilitator (not the on-call IC) to run postmortem
- Attendees: IC, comms lead, SMEs involved, team lead, customer success if customer-facing
- Write postmortem in shared doc; make it findable (searchable, linked from incident)
- No discussion of "who screwed up"; only "why did the system allow this to happen"
13. Build action item tracking and accountability (depends on: 12)
Create a system that tracks postmortem action items so they are not forgotten; currently 11 of 64 (17%) are being tracked, leaving 53 unfinished improvements.
System: each postmortem generates action items (e.g., "add monitoring for X," "update runbook for Y," "write test for Z"). Each item gets: clear description, owner (engineer's name), due date (1–4 weeks based on priority), severity (critical = must do before similar incident happens again; important = improve next month; nice-to-have = backlog). Action items live in a dedicated Jira project visible to all teams; owners are accountable (their manager reviews quarterly). Weekly: incident commander reviews open items due that week. Monthly: each team's postmortem items reviewed in their standup.
- Export action items from postmortem document to tracking system automatically
- Require IC to sign off that an action is complete before closing
- Report on completion rate as a metric (target: >80% by month 3)
14. Define incident metrics and KPIs
Establish what "good" looks like; measure so you can improve. Target metrics for 12 months out: mean time to detect 8 minutes (vs. 22 now), customers detect first <5% of incidents (vs. 40%), MTTR 45 minutes for SEV-1 (vs. 190), SLA credits <$100k/year.
Metrics to track: (1) MTTD = time from incident start to first alert/report; disaggregate: external report vs. internal detection. (2) MTTR = time from first report to full mitigation; track by severity and by service. (3) Customer-reported incidents per month (should drop to <2 per month). (4) Alert signal-to-noise ratio (goal: <5% false positive after S6 rules). (5) On-call satisfaction (survey: would you do this again?). (6) Postmortem action completion rate. (7) Incident commander and responder utilization (hours per week per person).
- Dashboard: auto-populated from incident tool, updated daily
- Disaggregate by team and service: which teams have bad MTTR? Which service is most incident-prone?
15. Create review cadence and governance process (depends on: 14)
Establish regular rhythm to inspect the metrics, spot trends, and adjust the process itself; prevent the system from calcifying.
Weekly: incident commander and on-call lead review prior week—number of incidents, any escalations, any communication gaps. Monthly: director-level incident review—trends by service, top causes of incidents, action item status, whether severity classification is working. Quarterly: full leadership review—MTTD, MTTR, customer impact, SLA credit spend, on-call satisfaction score, any systemic changes needed. Annually: audit the entire process for SOC 2 compliance.
- Assign meeting owners: weekly = on-call lead; monthly = director of reliability; quarterly = VP Eng + CFO (SLA cost) + customer success
- Use same data dashboard (S14) for all reviews
- Publish a monthly "incident newsletter" to all engineers: what happened, what we learned, what's improving
16. Prepare SOC 2 Type II audit checklist (depends on: 1, 2, 9, 12, 13, 14, 15)
Document that the incident management system meets the control requirements for a SOC 2 audit; audit happens in 8 months, so this work builds confidence in coverage.
Audit will test: (1) Is there a defined incident response process? (2) Are roles and responsibilities clear? (3) Are incidents logged and tracked? (4) Is root cause analysis performed? (5) Are action items tracked and completed? (6) Is on-call staffing adequate? (7) Are communications timely? (8) Are postmortems documented and blameless? Create a control mapping document that links each SOC 2 requirement to your process (S1–S15). Collect evidence: incident logs, postmortem documents, action item tickets, metrics reports, training records.
- Designate a compliance owner (often a reliability lead or security engineer)
- Run a mock audit at month 6 to identify gaps
- Ensure all postmortems and incidents are retained and searchable for auditor review
17. Develop implementation and rollout plan (depends on: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16)
Create a phased timeline to roll out the incident management system across all 28 teams; avoids big-bang failure and builds credibility.
Recommended structure: Phase 1 (weeks 1–4): build and test infrastructure (S5, S8, alerting); deploy severity levels and roles (S1, S2); pick pilot teams (2–3 high-traffic teams). Phase 2 (weeks 5–12): train pilot teams, run incident drills, refine playbooks based on learning; expand to half of remaining teams. Phase 3 (weeks 13–20): full rollout to all 28 teams; continue drills; track metrics. Phase 4 (weeks 21–28): stabilize, iterate on metrics, prepare for audit.
- Assign a release manager to coordinate across teams
- Create a detailed Gantt chart with swim lanes (infra, process, training, rollout)
- Identify risks: competing priorities, engineers worried about pager burden, tool adoption friction
- Plan stakeholder engagement: weekly updates to eng leadership, monthly town halls for all engineers
18. Build training and documentation (depends on: 17)
Create role-specific education so engineers understand the new system and are confident executing during an incident.
Training tracks: (1) All engineers: 30-min video on severity levels, communication expectations, postmortem format, action items (mandatory); (2) On-call responders: 2-hour workshop on playbooks, alert tool, escalation paths, when to call manager, case studies of real incidents, mobile app walkthrough; (3) Incident commanders: 4-hour IC bootcamp covering leadership under pressure, decision-making, communicating with executives, status page discipline, postmortem facilitation, practiced drills; (4) Communications leads: templates, when to update, how to talk to customers, regulatory notification rules.
- Record videos so async teams can learn on their schedule
- Create runbooks and quick-reference cards for each role (print + digital)
- Pair new on-call engineers with experienced responder for first week
- Require IC certification before anyone joins IC rotation (pass a practical drill)
19. Execute staged rollout across teams (depends on: 18)
Progressively activate the incident management system with feedback loops at each stage; reduces risk of system-wide failure.
Wave 1 (week 6–8): 3–4 pilot teams begin on-call rotations and incident response using new system; capture feedback daily. Wave 2 (week 10–14): 8–10 additional teams, incorporating lessons from Wave 1; ensure diversity of team types (payment processing, monitoring, data pipeline, auth, etc.). Wave 3 (week 15–20): remaining teams; by now, the system is proven and less hand-holding needed.
- Daily retros with Wave 1 teams: what worked, what was confusing, what broke
- Each wave produces a "lessons learned" document that informs the next
- Track adoption metrics: how many incidents reported per team, alert quality, MTTD/MTTR
- Address resistance: engineers who are skeptical of the system, on-call burden, tool friction; assign a "change champion" in each team
20. Run incident response drills and simulations (depends on: 19)
Practice incidents in a controlled setting so responders gain confidence and gaps are found before a real crisis.
Drill 1 (week 8, pilot teams only): simulated database failover; test alert routing, IC assignment, on-call notification, status page update, postmortem. Drill 2 (week 12, all teams): full incident spanning multiple services; test escalation, communications, decision-making. Drill 3 (week 18): chaos drill with intentional confusion—what if IC is unreachable, what if comms lead is on PTO, what if an alert tool goes down? Quarterly thereafter: repeat drills with new on-call engineers.
- Schedule drills at a time when key people are available (not Friday afternoon)
- Run drill with live observers (reliability lead, manager) who provide feedback
- Measure drill success: all steps completed on time, communications clear, postmortem written
- Debrief immediately: what did we learn, what do we change?
21. Establish continuous improvement process (depends on: 15, 20)
Build a feedback loop so the incident management system evolves based on what you learn.
Monthly improvement cycle: in the monthly incident review (S15), leadership identifies one or two friction points from recent incidents (e.g., "alert tool response is slow," "postmortem meetings are boring," "action items keep slipping"). Assign an owner to investigate and propose a fix (change to playbook, tweak to tool config, additional training, process change). Test the fix with pilot team, gather feedback, roll out if successful. Annual review: can we reduce on-call burden, simplify communication, improve MTTD further? Do any teams need special handling (new services, difficult customers)?
- Keep a public backlog of improvement ideas (Jira, shared doc) so any engineer can suggest
- Quarterly all-hands: share wins ("MTTD improved 50%," "on-call satisfaction up," "zero customer-detected incidents") to build momentum
- Link improvements back to metrics (S14): show that changes actually worked
--- PROPOSAL 2 (agent deepseek-flash_initial_2, deepseek/deepseek-flash) ---
Estimated complexity: high
Success metrics: - Median time to detect reduced from 22 minutes to 5 minutes or less by month 6.
- Customer-detected incidents reduced from 40 percent to 10 percent or less of all customer-impacting incidents by month 9.
- Median time to mitigate reduced from 3 h 10 min to 60 minutes or less for SEV1 and SEV2 by month 9.
- Customer-impacting incidents reduced by at least 40 percent year over year from the 31 incident baseline.
- SLA credits paid reduced by at least 50 percent year over year from the $1.3M baseline.
- 100 percent of SEV1 incidents have a named Incident Commander within 5 minutes and a scribe within 15 minutes.
- Zero incidents in which command authority is unclear for more than 15 minutes.
- 28 of 28 teams have a documented on-call rotation, an owning service map and at least one trained on-call engineer by month 6.
- Central Incident Commander roster holds at least 40 certified ICs covering 24x7 with no single point of failure.
- 100 percent of on-call rotations are paid under a published policy by month 5.
- Monthly alert volume reduced from 3,400 to below 700, with a false-positive rate below 20 percent.
- No service exceeds 2 pages per on-call shift, measured monthly for three consecutive months.
- 100 percent of SEV1 and SEV2 postmortems published internally within 15 business days.
- At least 90 percent of postmortem action items closed within 60 days, up from 17 percent (11 of 64).
- Status-page first update published within 30 minutes on at least 95 percent of SEV1 incidents.
- Zero missed regulatory notification windows on any incident requiring notification.
- SOC 2 Type II audit passed with no findings related to incident response.
- Review cadence sustained: weekly operational review in at least 90 percent of weeks, 12 of 12 monthly reliability reviews, 4 of 4 quarterly executive reviews.
- On-call satisfaction at 70 percent or higher on the quarterly survey, with zero on-call-attributed voluntary attrition.
- 100 percent of new engineers complete the incident-response onboarding module within 30 days of joining.
Steps (21):
1. Programme charter, ownership and executive mandate
This step turns the CEO's email into a funded programme with a named owner and explicit authority. Without it, every downstream decision stalls in cross-team negotiation.
- Appoint a single accountable process owner (for example a Director of Incident Management) reporting to the CTO, with a dotted line to the COO for customer and SLA matters.
- Publish a one-page charter covering scope (all customer-impacting and money-moving incidents), decision rights, and the power to override team preferences during an active incident.
- Define the funding envelope: tooling licences, training time, exercise time and on-call compensation, with an indicative annual figure.
- Set the timeline against the SOC 2 date: a working process in four months, evidence accumulating from month two, audit-ready by month seven.
- Stand up a steering group with CTO, VP Engineering, Head of Support, Head of Compliance and one engineering manager per region.
- Agree that incident-process participation is a documented performance expectation for engineering managers, not an optional extra.
2. Baseline measurement and evidence pack (depends on: 1)
You cannot demonstrate improvement without a defensible starting point, and you cannot win the argument about noise without numbers.
- Build a 12-month incident register: date, retro-assigned severity, detection source, time to detect, time to mitigate, customer impact, services involved and SLA credits paid.
- Quantify the alert estate per tool, per team and per service; compute page-to-action ratio, list the 50 noisiest rules and count off-hours interruptions per engineer.
- Survey on-call engineers and managers on burden, fairness and clarity of escalation, targeting a response rate above 70 percent.
- Interview Support, Account Management and Sales about how customers learn of incidents and what they complain about.
- Document exactly where the current process breaks: unclear command in the two known incidents, postmortem action closure at 11 of 64, and ad-hoc status-page authorship.
- Publish the pack internally as the problem statement and retain it as management-review evidence for the SOC 2 audit.
3. Severity taxonomy and trigger matrix (depends on: 1, 2)
Severity is the keystone of the whole process. Every other rule, from paging to communications timing to postmortems, is keyed off it.
- Define four levels plus a special SEV0 for security or regulatory events: SEV1 for total or material loss of a payment path, SEV2 for degradation or single-region loss, SEV3 for limited impact with a workaround, SEV4 for internal-only issues and near-misses.
- Anchor each level in customer and money terms: customers affected, dollars at risk, regions lost, ledger integrity, settlement delay.
- Specify automatic triggers, for example loss of one AWS region, ledger write failures, a missed settlement cut-off, or payment success rate below threshold for five minutes.
- State who may declare each level (any engineer, Support or account manager may declare) and who may only recommend a downgrade (the Incident Commander alone).
- Map each level to SLA credit exposure and to the customer-visible status-page state.
- Include worked examples from the last 12 months so teams recognise their own incidents in the definitions.
- Add a review clause: the taxonomy is re-validated quarterly against real declarations.
4. Incident roles, command structure and decision rights (depends on: 2, 3)
The two incidents where nobody knew who was in charge for over an hour are the direct brief for this step.
- Define roles with one-page role cards: Incident Commander, Deputy IC, Operations Lead, Internal Communications Lead, Customer Communications Lead, Scribe, Subject-Matter Responders and an Executive Sponsor for SEV1 only.
- State the core rule plainly: the IC owns the incident, not the fix, and does not debug.
- Give the IC explicit decision rights: declaring and escalating severity, freezing changes, halting deploys, approving customer messaging and calling additional responders.
- Define minimum viable role coverage per severity: SEV1 staffs every role, SEV3 staffs an IC and a scribe only.
- Define handover discipline: maximum four-hour IC shifts on SEV1, a written handover template, and a Deputy IC nominated within 15 minutes.
- Set behaviour standards for responders: one incident channel, one bridge, no side channels, and explicit asks with a named owner and a time.
- Publish role cards on the internal wiki and link them from every paging notification.
5. Escalation, paging and incident lifecycle policy (depends on: 3, 4)
This step defines the mechanical path from an alert to a declared incident and back to normal service.
- Define lifecycle states with entry and exit criteria: Detected, Triaged, Declared, Mitigated, Resolved, Postmortem, Closed.
- Set acknowledgement targets: page acknowledged within five minutes, triage decision within fifteen, declaration within thirty for anything customer-facing.
- Define escalation ladders per layer (responder, service owner, team manager, IC on-call, VP Engineering) each with an automatic timer.
- Make escalation blameless and automatic: no responder is ever criticised for escalating, and timers fire whether or not a human asks.
- Define change freeze and rollback authority during SEV1 and SEV2, and the single condition that lifts the freeze.
- Enforce one incident, one record: the incident record is the sole source of truth for timeline, roles and communications.
- Require every SEV1 and SEV2 to produce an automatically captured timeline from channel and bridge, never one written from memory afterwards.
6. Detection strategy: SLOs, signals and customer-journey monitoring (depends on: 2, 3)
Customers detected 40 percent of incidents first. That number is the reason this step exists.
- Define SLIs and SLOs for the top 20 customer journeys, including payment initiation, settlement, ledger read and write, API availability and webhook delivery, measured per region.
- Require symptom-based alerting on those SLOs rather than cause-based alerting on infrastructure metrics.
- Add synthetic transaction monitoring from outside AWS in both regions and a third location, on a one-minute cadence for money-moving paths.
- Add ledger and shared-PostgreSQL specific signals: replication lag, connection saturation, write latency, transaction ID exhaustion and checkpoint pressure.
- Open a customer-reported path so Support and account managers can raise an incident directly, and count that path as a detection source in reporting.
- Publish a detection contract per service: every one of the 180 services needs an owner, at least one symptom alert and a documented expected detect time.
- Fund a separate resilience track to reduce shared-cluster blast radius, because better detection will not save a single shared ledger during a corruption event.
7. Alert quality standard and noise-reduction programme (depends on: 2, 3, 6)
3,400 alerts a month with 85 percent noise is the reason engineers resent the pager. Fixing it is the price of admission for everything else.
- Publish alert standards: every page must be symptom-based, actionable, owned, linked to a runbook and mapped to a severity. No runbook, no page.
- Separate paging from ticketing: only customer-impacting or imminent-impact signals may page; everything else becomes a ticket or a dashboard entry.
- Set a noise budget per team and per service, for example no service may exceed two pages per on-call shift, measured monthly.
- Define the default action for a noisy alert: fix, tune or delete within ten working days, with deletion treated as a legitimate and successful outcome.
- Introduce correlation and deduplication at the event pipeline so a single root cause produces one page instead of forty.
- Require expiry dates on every silencing rule and temporary threshold so suppression cannot become permanent blindness.
- Run a 90-day noise sprint with a visible burn-down of the top 100 noisiest rules, each owned by a named manager.
- Report page-to-action ratio per team in the monthly reliability review.
8. On-call architecture and 24x7 coverage model across 28 teams (depends on: 3, 4)
This is the hardest political step. The answer to carrying a pager for another team's code is that every team carries its own, and the platform carries the shared risk.
- Adopt a federated model: every service has exactly one owning team, and that team's primary on-call carries its own pager. No team is paged for code it does not own.
- State the consequence honestly: 16 of 28 teams currently have no on-call. They must build one or formally transfer ownership of their services to a team that will.
- Merge small or low-traffic teams into shared rotations so no rotation has fewer than six engineers, below which coverage is not sustainable.
- Create a separate, centrally staffed Incident Commander on-call roster drawn from trained senior engineers across all teams, covering 24x7.
- Define primary and secondary per rotation, with the secondary engaged only on a no-acknowledge or an explicit request.
- Define coverage across the two AWS regions and New York business hours: one global IC rotation, service on-call aligned to their service's users.
- Define the unresponsive-team path: fifteen minutes without acknowledgement escalates to the team manager, thirty minutes to the IC, who may then direct any available engineer.
- Publish a coverage matrix of all 28 teams showing services, rotation size and gaps, reviewed monthly.
- Make on-call participation an explicit expectation in engineering job levels and hiring criteria.
9. On-call compensation, wellbeing and sustainability policy (depends on: 8)
Unpaid on-call is the most cited reason for resistance. The policy must be settled before rollout, not negotiated during it.
- Move to paid on-call: a per-shift stipend or salary uplift agreed with HR and Finance and benchmarked to the New York market.
- Pay event-based compensation for incident callouts outside business hours, with a minimum call-out block.
- Provide compensatory rest: no engineer works a normal day after a night incident, and the rest day is documented, not granted as a favour.
- Cap intrusion by defining a maximum number of off-hours pages per shift, with a mandatory review triggered whenever it is exceeded.
- Define a voluntary opt-out path for engineers with genuine constraints, balanced by an explicit obligation that someone else is paid to take the shift.
- Include on-call expectation and compensation in offers and job descriptions so the commitment is set before hiring.
- Publish the policy with an effective date before any team is asked to join a new rotation.
- Review the policy every six months against actual page volumes, attrition and survey results.
10. Internal and customer communications policy with timing SLAs (depends on: 3, 4)
Today the status page is written by whoever is around. This step replaces improvisation with a clock and a named owner.
- Set internal cadence: first update within 15 minutes of declaration, then every 30 minutes for SEV1 and 60 minutes for SEV2, regardless of whether there is progress.
- Set customer cadence: status page updated within 30 minutes of a SEV1 and 60 minutes of a SEV2, and a no-new-information update is still mandatory.
- Define the channel hierarchy: status page for everyone, direct email to affected customers on SEV1, named account-manager calls for the top 50 accounts.
- Prepare templates per severity in advance with legal and compliance pre-approval, covering detection, impact, workaround, mitigation and next-update time.
- Define regulatory obligations explicitly: money transmitter and banking regulator notification windows, security breach notification, and who signs off (Compliance, not Engineering).
- Prohibit speculation: customer communications never guess at cause or blame and never commit to a root cause before the postmortem.
- Define the closing communication: resolution notice, expected SLA credit handling and the committed date for a written report.
- Assign a named Customer Communications Lead per incident with a trained deputy on every SEV1.
11. Status page, notification tooling and account-manager playbook (depends on: 10)
Policy without tooling collapses at three in the morning. This step makes publishing a five-minute action.
- Upgrade or replace the status page so components map to customer journeys rather than internal services, with subscriber control per component.
- Integrate the incident tool with the status page so the incident record drives the update and the public timeline.
- Provide one-click templates pre-filled with severity, impact language and next-update time.
- Give account managers a playbook: contact tree, what they may say, what they must not say, and how to escalate a customer question into the incident channel within minutes.
- Define the SLA credit process end to end, covering computation, approval, customer notification and finance treatment, so credits stop being a manual scramble.
- Host the status page outside the production failure domain so it survives a total platform outage.
- Test publishing during game days, including a simulated status-page outage and a simulated loss of the primary region.
12. Postmortem policy, template and blameless review process (depends on: 3, 4)
Only 11 of 64 action items closed means the postmortem ritual is currently a writing exercise. This step rebuilds it around learning and tracking.
- Make postmortems mandatory for every SEV1 and SEV2, every SEV3 with customer impact or a repeat pattern, and any near-miss the IC flags.
- Set deadlines: draft within five business days, blameless review within ten, internal publication within fifteen.
- Adopt a single template: impact, timeline, detection, response, contributing factors, what went well, what went badly, action items.
- Train a pool of blameless facilitators and require a trained facilitator for every SEV1 review.
- Prohibit counterfactual and blame language in the template, and require contributing factors across tooling, process, organisation and human factors.
- Limit action items to a small number of concrete, verifiable, owned changes with dates.
- Publish all postmortems internally by default, with security review only for genuinely sensitive material.
- Produce a customer-facing root cause report variant for SEV1 incidents, especially those affecting regulated or top-tier accounts.
13. Corrective action tracking and reliability backlog governance (depends on: 12)
A postmortem without durable action tracking is a complaint, not a control.
- Create a single reliability backlog in the engineering tracker with a mandatory label, owner, due date and link to the originating incident.
- Define closure criteria that require evidence: a merged change, a tested alert or a verified drill, never a self-reported status change.
- Protect capacity by reserving a fixed percentage of each team's sprint for reliability work, with unspent capacity visible to vice presidents.
- Run a weekly ageing review of open actions and escalate anything overdue by more than 30 days to the VP Engineering.
- Report closure rate and median age monthly, targeting more than 90 percent closed within 60 days.
- Require a repeat incident in the same area to trigger a design review rather than another action item.
14. Incident tooling consolidation and integration (depends on: 3, 5, 7, 11)
Six alerting tools and no single incident record is a structural cause of the 22-minute detection and the three-hour mitigation.
- Select one incident management platform for paging, on-call schedules, escalation policies, incident records and postmortem workflow.
- Consolidate the six alerting sources into a single event pipeline feeding that platform, with deduplication and severity mapping applied at ingest.
- Integrate with platform and ledger observability so responders see dashboards and runbooks inside the incident record.
- Integrate chat and bridge: incident channel auto-created, timeline auto-captured, decisions logged as they happen.
- Define the data model and retention required for SOC 2 evidence: who did what, when, and under whose authority.
- Run a dual-run period alongside the old tools with a defined rollback, then switch off the legacy tools on a published date.
- Budget for licences, migration effort and a two-week hardening period after cutover.
15. Training, certification and exercise programme (depends on: 4, 5, 10, 12)
A process that exists only on a wiki page fails on the first real page. Skills have to be built and tested before they are needed.
- Build a short practical curriculum: how to be on call, how to declare an incident, how to run an incident as IC, how to communicate and how to write a postmortem.
- Require certification before joining the IC on-call roster: a written assessment plus a live simulated incident.
- Train at least two certified ICs per team group so the central roster has depth across all 28 teams.
- Run monthly tabletops on realistic scenarios drawn from the last 12 months, including region loss and ledger corruption.
- Run quarterly game days with deliberately injected failure, including shared-PostgreSQL failover and status-page outage.
- Add a mandatory onboarding module for every engineer joining or transferring in, with completion records suitable for audit.
- Track training completion by team and publish it in the monthly reliability review.
16. Metrics, dashboards and review cadence (depends on: 2, 3)
The programme needs a public scoreboard, or it will quietly rot after the audit.
- Define the outcome metrics: time to detect by source, time to mitigate, percentage of incidents detected by customers (target below ten), incidents by severity and SLA credits paid.
- Define the process metrics: declaration latency, page acknowledgement rate, IC roster coverage, first-update timeliness and update-cadence adherence.
- Define the health metrics: alert volume and noise ratio per team, off-hours pages per engineer, postmortem timeliness, action closure rate and action age.
- Publish live dashboards visible to every engineer, not only to managers, refreshed daily.
- Institute a weekly operational review of 30 minutes going incident by incident, a monthly reliability review of trends and actions, and a quarterly executive review with the CEO's office.
- Baseline every metric against the S2 evidence pack and set 90-day and 12-month targets.
- Require every review to end with decisions and owners, not just numbers.
17. Pilot with volunteer teams (depends on: 5, 7, 9, 11, 12, 13, 14, 15, 16)
Do not roll out to 28 teams untested. Run the entire process end to end with a small cohort first.
- Recruit three to four volunteer teams covering a mix of criticality: one payment-path team, one ledger-adjacent team, one platform team and one low-traffic team.
- Run the complete process in the pilot: new severity scale, roles, escalation, communications, postmortems, action tracking and paid on-call.
- Instrument the pilot against the S16 metrics and compare results with the S2 baseline.
- Hold weekly retrospectives with pilot teams and iterate on the written policies, the tooling and the training.
- Fix the top issues found before any wider rollout and document what changed and why.
- Produce a pilot report with before-and-after numbers to carry into every rollout conversation.
- Set explicit pilot exit criteria: rotation coverage achieved, no unacknowledged pages over a defined period, postmortems delivered on time and actions tracked.
18. Phased rollout to all 28 teams (depends on: 13, 16, 17)
Rollout is a staged migration with readiness gates, not an email announcement.
- Sequence the 28 teams into four waves of roughly seven teams, ordered by customer impact, with three weeks between waves.
- Define a per-team readiness checklist: services mapped and owned, alerts cleaned to standard, runbooks written, rotation staffed, training complete and manager briefed.
- Hold a gate review with the process owner before each team joins, and move unready teams to the next wave with a dated remediation plan.
- Give each wave a named champion and run an internal communications cadence that explains the why using pilot numbers.
- Handle resistance directly by publishing the own-your-code-own-your-pager rule and the paid on-call mechanics before each wave, not after.
- Retire legacy tools, escalation lists and the informal status-page process at the end of each wave on a published cutover date.
- Harvest feedback formally at each wave and push accepted changes back into the policy documents through change control.
19. SOC 2 incident-response control mapping and evidence framework (depends on: 1, 3, 10, 12)
The audit will test incident response. Mapping controls now means evidence accumulates from the first real incident under the new process.
- Map the process to the relevant Trust Services Criteria for incident identification, response, evaluation of incidents and communication of security events.
- Write control statements in auditor language and name a single owner for each control.
- Specify the evidence artifact for each control: incident record, severity classification, paging log, communications log, postmortem, action tracker entry and training record.
- Set evidence retention and storage location so nothing depends on a laptop or on chat history that expires.
- Run an early walkthrough with an experienced compliance partner or the auditor's readiness team to test the design before the audit window.
- Flag controls that are blocked until rollout reaches the relevant teams, and keep a gap register with owners and dates reviewed fortnightly by the steering group.
20. SOC 2 dry run, gap remediation and audit support (depends on: 18, 19)
This step converts a good process into a provable one, a few months before the auditors arrive.
- Run an internal dry run six weeks before the audit window, sampling real incidents from the pilot and early waves against each control's evidence requirements.
- Remediate every gap found, prioritising anything that risks a qualified opinion.
- Brief every engineer who may be interviewed so they can describe the process as they actually practise it.
- Prepare the auditor package: process documentation, sample incident records, training records, on-call schedules and the action tracking register.
- Designate a single audit liaison and a small evidence-request team so requests do not land on 28 teams at once.
- Rehearse the walkthrough with the incident commander and communications roles, since auditors probe realism under pressure.
- Schedule the dry run and the audit alongside rollout waves so no team is hit with both in the same week.
21. Standing governance, process ownership and continuous improvement (depends on: 18, 20)
The classic post-audit failure is that the process freezes and then decays. This step keeps it alive.
- Establish a standing Incident Management Council chaired by the process owner, meeting monthly with engineering, support, compliance and product representation.
- Give the process owner a documented mandate to change standards, and require formal change control for any change to severity, roles, communications timings or compensation.
- Re-validate the severity taxonomy quarterly against real declarations and re-baseline the metric set annually.
- Fund a resilience roadmap separate from incident response: reducing shared-database blast radius, multi-region failover readiness and deploy safety.
- Report quarterly to the executive team on the metric set and on the top five systemic causes of incidents.
- Feed incident themes into architecture review, release engineering and hiring so the programme improves the system, not just the response.
- Refresh the training and tabletop programme annually and after any SEV1.
--- PROPOSAL 3 (agent qwen3.8-flash_initial_3, alibaba/qwen3.8-flash) ---
Estimated complexity: high
Success metrics: - MTTD reduced to under 10 minutes with >90% detection occurring before customer reporting.
- MTTR for Sev1/Sev2 incidents reduced to under 60 minutes.
- SLA credits paid reduced by 50% within 12 months.
- On-call alert volume reduced by 80% (from 3,400 to <600 actionable alerts/month).
- 100% of Sev1/Sev2 incidents have a blameless postmortem completed within 5 days.
- >90% completion rate of postmortem action items within 30 days of assignment.
- Zero ambiguity in incident command: 100% of major incidents have a designated IC logged in the timeline.
- SOC 2 Type II audit passes incident response controls with zero critical findings.
Steps (12):
1. Establish Governance and Executive Sponsorship
Secure explicit CEO and CTO backing to mandate the program, framing it as a revenue protection and compliance necessity rather than an engineering burden. Form a cross-functional Incident Management Task Force.
- Identify a single executive owner (e.g., VP of Engineering or Head of Platform) with authority over budget and staffing.
- Create a charter defining the scope: severity definitions, on-call compensation, and tool consolidation.
- Communicate the strategic priority to all 28 team leads to preempt pushback regarding pager responsibility.
2. Define Severity Matrix and Response SLAs (depends on: 1)
Create a rigorous, binary severity scale to eliminate ambiguity in escalation and communication speed. Align internal response times with the 99.95% SLA.
- **Sev1 (Critical)**: System-wide outage, data corruption, or >5% transaction failure rate. Immediate executive page.
- **Sev2 (Major)**: Single customer segment impacted or >1% transaction failure rate. High-priority team response.
- **Sev3 (Minor)**: Degraded performance or isolated bug. Standard business-hours follow-up.
- **Sev4 (Cosmetic)**: No functional impact. Backlog item.
- Define maximum detection and mitigation times for each level.
3. Restructure On-Call Organization and Compensation (depends on: 1)
Transform the current unpaid, fragmented on-call system into a professional, compensated, and centralized rotation to ensure 24x7 coverage across all 28 teams. Address the 'pushback' by making it mandatory and paid.
- Consolidate the 12 current on-call teams into a unified, rotating pool covering all 28 teams.
- Implement a fair compensation model: base stipend + hourly rate for active hours + incident bonus for Sev1/2.
- Define a strict 'Buddy System' or 'Follow-the-Sun' rotation if geographically feasible, otherwise ensure adequate team size to rotate fairly.
- Mandate that every team must participate in the rotation; no exemptions.
4. Define Incident Roles and RACI Model (depends on: 2)
Eliminate the 'who is in charge' ambiguity by pre-defining roles and establishing a clear chain of command for every incident. Roles are assigned per incident, not permanently held.
- **Incident Commander (IC)**: Single point of authority. Owns the timeline, delegates tasks, and approves communications. Must be senior.
- **Scribe**: Maintains the real-time timeline log of actions, timestamps, and decisions.
- **Communications Lead**: Owns external/status page updates and internal Slack broadcasts. Shields the IC from interruptions.
- **Technical Leads (SMEs)**: Deployed by IC to troubleshoot specific services. They do *not* coordinate with each other; they report to IC.
- Create a 'Rapid Response Team' (RRT) for high-severity incidents involving the shared PostgreSQL cluster.
5. Consolidate Monitoring and Enforce Alert Hygiene (depends on: 3)
Reduce the 3,400 monthly alerts (85% noise) to a signal-to-noise ratio that allows engineers to trust the pager. Standardize tooling.
- Consolidate the six disparate alerting tools into a single platform (e.g., PagerDuty, Opsgenie, or Datadog).
- Implement 'Actionable Alerting': If an alert pages, there must be a documented, immediate runbook action. No 'informational' pages.
- Set a hard cap on alert volume per engineer/night. Breaching the cap triggers a mandatory alert quality review.
- Deprecate legacy tools that cannot map to the severity matrix defined in S2.
6. Build Detection and Escalation Automation (depends on: 5)
Reduce the 22-minute mean time to detect (MTTD) by automating the path from metric breach to page. Ensure customers never find an issue before the internal pager does.
- Implement synthetic transactions for core payment flows to detect failures before real users do.
- Configure escalation policies: If not acknowledged in 5 minutes, auto-escalate to the team lead, then to the Incident Commander pool.
- Map specific Kubernetes and PostgreSQL metrics to the Severity Matrix (S2) to trigger automated pages.
- Create a 'Major Incident Channel' in Slack that is automatically opened and pinned when a Sev1/2 is triggered.
7. Standardize Communication Templates and Timing (depends on: 4, 6)
Codify the rules for external and internal communications to stop the 'whoever is around' approach. Link these timings to the Severity Matrix.
- **Status Page Updates**:
- Sev1: First update within 5 minutes of detection. Updates every 15 minutes.
- Sev2: First update within 30 minutes. Updates every 60 minutes.
- **Customer Direct Mail**: For Sev1 impacting specific accounts, Account Managers send pre-approved templates within 1 hour of mitigation.
- **Regulator Notification**: Define legal triggers for NY financial regulations.
- Create a 'Comms War Room' Slack channel restricted to the IC and Comms Lead.
8. Formalize the Postmortem Process (depends on: 4, 7)
Transform postmortems from optional/various formats into a mandatory, blameless, and tracked process. This directly addresses the SOC 2 audit requirement for continuous improvement.
- Mandate a postmortem for every Sev1 and Sev2 incident within 5 business days.
- Enforce a 'Blameless' charter: Focus on systemic/process failures, not individual errors.
- Adopt a standardized template: Timeline, Root Cause Analysis (5 Whys or Fishbone), Impact Analysis, and Action Items.
- Action Items must have: Owner, Due Date, and Priority. They are tracked in Jira/Linear, not the doc.
9. Implement Action Item Tracking and Governance (depends on: 8)
Solve the '11 of 64 closed' problem by integrating postmortem actions into the core engineering workflow and executive reviews. Ensure fixes are shipped.
- Integrate Jira with the Postmortem tool (e.g., OpsLevel, Rootly, or Notion API). Auto-create tickets for every AI.
- Block 'Sev1 Closure' until all high-priority Action Items are closed or explicitly deferred with executive approval.
- Weekly Engineering Leadership review of open postmortem actions older than 30 days.
- Tie action item completion rates to team performance reviews.
10. Deploy Training and Runbook Library (depends on: 6, 7, 8)
Address the skill gap and the 'pushback' by equipping engineers with the tools and knowledge to succeed on-call. Reduce anxiety through clarity.
- Audit all 180 services. Require every service to have a 'Page-Worthy' runbook.
- Run 'Game Day' (Chaos Engineering) exercises quarterly to practice role assignments (IC, Scribe, Comms) in a simulated Sev1.
- Conduct mandatory training for the new on-call rotation and compensation model.
- Provide a 'Playbook' for the Incident Commander role, including scripts for communication.
11. Phased Rollout and Change Management (depends on: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10)
Introduce the new process in phases to manage cultural resistance and technical debt, ensuring adoption before the SOC 2 audit. Do not 'big bang' the change.
- **Month 1-2**: Launch Severity Matrix and Consolidated Tooling (S2, S5). Pilot on 5 high-traffic teams.
- **Month 3-4**: Roll out new On-Call Compensation and Roles (S3, S4) to all 28 teams. Enforce new Communication Templates.
- **Month 5-6**: Enforce strict Postmortem and Action Item Tracking (S8, S9).
- **Month 7**: Final audit readiness check. Simulate SOC 2 auditor questions on incident response.
12. Define Metrics for Success and Continuous Review (depends on: 2, 6, 8, 9)
Establish a dashboard to prove the new system works, linking directly to the $1.3M savings in SLA credits. Monitor the health of the incident process itself.
- Track MTTD (Target: <5 min), MTTR (Target: <30 min for Sev2), and Detection Ratio (Target: >95% internal detection).
- Monitor 'Alert Noise' (target: <5% of alerts result in a non-actionable page).
- Track 'Action Item Completion Rate' (Target: 100% on-time closure for Sev1/2).
- Calculate 'SLA Credit Avoidance' vs. On-Call Program Cost to prove ROI to leadership.
Your answer has these parts:
- "round_summary": one or two sentences on the round as a whole.
- "shared": a short list (four items at most) of what most proposals have in common: approaches, steps, priorities.
- "differences": a short list (four items at most) of what separates them, naming the proposals and the steps concerned.
- "proposals": one entry per proposal, each with "proposal" (its number), "summary" (a very concise summary of what the agent proposes: three or four sentences at most) and "approach" (the angle it takes, in a few words).
[ROUND 1]
[SYSTEM]
You are an expert reviewer of multi-agent planning processes.
Several LLM agents drafted plans for a task, refined them over a number of rounds while seeing each other's proposals, and finally voted for the best one.
Be exhaustive but precise: name concrete steps, ideas and metrics, never generalities. Judge plans by their fitness for the task as stated, their realism, their completeness, the soundness of their order and dependencies, how measurable their success is and how they handle things going wrong.
You are an impartial evaluator, not a chronicler: assess the proposals and the process on their merits, never rationalise what happened or assume that the outcome was right.
After your analysis, answer in the requested structure.
Every text field you write will be read by a busy person who skims. Make it easy to skim: short sentences and short paragraphs; when you name several things, prefer a list to a paragraph, with sub-items when an item has parts, but keep a single fact as a sentence; lead with the point and then the evidence; name proposals and steps by number (P2, step 4); no preamble, no repetition of the question, no closing summary; bold at most one key phrase per item or paragraph. Text fields accept Markdown: a blank line between paragraphs, "- " for lists, **bold**.
[HUMAN]
Task given to the agents: "A B2B payments platform in New York (260 engineers in 28 teams, 2,100 customers, $4B processed a month, a 99.95% SLA with service credits) runs about 180 services on Kubernetes across two AWS regions, with a shared PostgreSQL cluster for the ledger. In the last 12 months: 31 customer-impacting incidents, median time to detect 22 minutes (customers detected 40% of them first), median time to mitigate 3 h 10 min, $1.3M paid in SLA credits, two incidents in which nobody was sure who was in charge for over an hour, and a CEO email about "outages we hear about from clients". On-call exists in 12 of the 28 teams, unpaid, fed by alerts from six different tools — 3,400 alerts a month, 85% of them noise; there is no severity scale, status-page updates are written by whoever is around, and postmortems happen for some incidents, in various formats, with action items that are rarely tracked (11 of 64 closed). A SOC 2 Type II audit in eight months will test incident response. Engineers push back against "carrying a pager for other teams' code".
Define the incident management process: severity levels and what each one triggers; roles (incident commander, communications, scribe, subject-matter responders) and how they are staffed 24x7 across 28 teams; the on-call structure, rotations, compensation and the rules for alert quality; detection and escalation paths; internal and customer communications (status page, account managers, regulators when required) with their timings; postmortems (when mandatory, format, blameless review, ownership and tracking of actions); the metrics and reviews that show whether it works; and how it is introduced across the teams without waiting for the audit."
This is round 1, a refinement round: every agent received ALL the proposals of round 0 and wrote a new plan, improving on them or taking a different approach. By convention, the previous version of proposal N is proposal N of round 0, written by the same model.
PROPOSALS OF ROUND 0 (the previous versions):
--- PROPOSAL 1 (agent claudeHaiku4.5_initial_1, anthropic/claude-haiku-4-5) ---
Estimated complexity: high
Success metrics: - Median time to detect reduced from 22 minutes to <8 minutes within 6 months
- Customer-detected incidents drop from 40% to <5% within 6 months
- Median time to mitigate (MTTR) reduced from 3h 10min to <45 minutes for SEV-1 incidents within 6 months
- Annual SLA credits decrease from $1.3M to <$100k within 12 months
- Alert noise reduced from 3,400 per month (85% false positive) to <400 per month (>95% signal) within 3 months
- Zero incidents with command-and-control ambiguity (>1 hour without clear IC) within 2 months
- Postmortem action item completion rate reaches >80% (from 17%) within 4 months
- On-call satisfaction score (survey) reaches >7/10 for on-call engineers within 3 months
- All 28 teams integrated into incident management system with active on-call rotations by week 20
- SOC 2 Type II audit passes incident response controls with no findings 8 months from start
- Incident commander certification: 100% of active ICs trained and drilled within 2 months
- Monthly incident review meeting established and attended by leadership; trends documented
- New incident system integration complete: single alert tool, single dashboard, all 180 services feeding in, <5 min deployment
Steps (21):
1. Define severity levels and decision criteria
Create a four-tier severity framework (SEV-1 through SEV-4) that guides all downstream decisions about response, escalation, and communications.
Each level must specify: customer impact (revenue at risk, customers affected, data loss risk); financial threshold triggering service credits; whether an incident commander is required; response time SLA (e.g., SEV-1 < 5 min notification, SEV-4 < 2 hours); and the go-live decision tree (when to declare and when to resolve).
- SEV-1: Complete service down or critical path broken for >5% of customers; every minute costs money; IC required; 99.99% uptime threatened
- SEV-2: Significant degradation, features unavailable, affecting 1–5% of customers; IC typically required
- SEV-3: Minor impact, limited customer footprint or workaround exists; escalation path but not automatic IC
- SEV-4: Observations or minor issues; alert-driven, no escalation unless pattern emerges
2. Define incident roles and responsibilities (depends on: 1)
Create the organizational roles that operate during an incident: who is in charge, who talks to customers, who writes down what happened, who fixes the system, and how decisions are made under pressure.
Each role must have a single-sentence mission, decision authority, and escalation upward.
- Incident Commander: owns decision-making and timeline; declares severity; resolves conflicts; may or may not be technical
- Deputy IC: shadow to IC and takes over if IC becomes unavailable
- Communications Lead: writes status page, notifies account teams, manages customer perception
- Scribe: records decisions, who did what, key timestamps; not responsible for fixing
- SME Responders: engineers with context on the failing service(s); take IC's direction without debate
3. Design 24x7 on-call rotation structure (depends on: 2)
Build a rotation model that covers all 28 teams with primary and backup on-call engineers every hour across weekdays, evenings, weekends, and holidays; addresses the pager-carrying resistance.
Key design decisions: Is coverage per-team (each team owns its services) or pooled (shared responder pool handles anything)? How many people per rotation? How long are shifts (one week, two weeks)? When can engineers opt out without leaving the team exposed? Which roles are on-call (IC, communications, SME)?
- Recommend: dedicated IC pool (4–6 people in fast rotation) + per-team SME on-call for each team's own services
- Recommend: two-week rotation blocks to reduce handoff friction
- Recommend: one primary, one secondary per slot; secondary handles during primary's escalation
- Provide swaps, blackout dates, and a rule that no engineer is on-call more than 2 weeks per quarter
4. Define on-call compensation and incentives (depends on: 3)
Create a pay model that makes on-call acceptable and rewards engineers who carry the pager; ties compensation to real business risk.
- Base on-call stipend: e.g., $500–1,000 per week while on-call (regardless of incidents)
- Callback pay: 1.5× hourly rate for time spent mitigating incidents during off-hours
- Incident bonus: $50–100 extra per SEV-1 or SEV-2 incident mitigated (recognition)
- Comp time: full business day off after an incident that required >2 hours mitigation during night/weekend
- Annual bonus tie-in: 10–20% bonus multiplier for flawless on-call reviews
- Communicate: position as investment in reliability, not punishment for being online
5. Consolidate alert routing infrastructure
Replace six alert tools with a single ingestion and routing system; stop engineers from being woken by duplicate alerts, and make escalation automated instead of manual.
Evaluate existing tools (likely candidates: PagerDuty, Opsgenie, or Incident.io) or build a lightweight wrapper. The system must: accept alerts from all 180 services; deduplicate and correlate (same outage, different monitoring source); route to correct on-call engineer; expose an API for playbook automation; log every alert for postmortem analysis.
- Choose tool by week 2 of S5
- Migrate alerting endpoints from 6 sources to 1 by week 4
- Set up audit trail and retention
- Ensure mobile app works (on-call engineers need to engage from phone)
6. Build alert quality rules to cut noise (depends on: 1, 5)
Implement rules that automatically suppress the 85% of alerts that are noise (flapping, transient errors, auto-recovered conditions). Target: <400 actionable alerts per month.
Rules to implement: suppress alerts if service auto-recovered within 30 seconds; deduplicate same alert from multiple monitoring sources; suppress alerts for known maintenance windows; group flapping alerts (same service, >5 occurrences in 2 minutes) into a single page to on-call; rate-limit alerts from noisy services (e.g., max 1 alert per 5 minutes per service until silence clears).
- Audit existing 3,400 alerts per month: which are true signals, which are noise
- Tag each alert source with severity level (S1, S2, S3, S4 from S1)
- Create exceptions list: services known to be noisy, require different rules
- Weekly review: alert teams that trigger >50 alerts per week for reduction strategies
7. Implement automated detection and escalation paths (depends on: 1, 5, 6)
Wire the alert system to automatically escalate based on time or severity; removes the need for manual judgment calls during chaos.
Logic: SEV-1 alert arrives → IC notified instantly via phone call + SMS + Slack + mobile; if IC does not acknowledge within 2 minutes, page deputy IC; Communications Lead pinged simultaneously. SEV-2: on-call SME for that service + IC notify via Slack and mobile, escalate to IC's manager if not acknowledged in 10 min. SEV-3/4: on-call SME only, escalate after 30 min.
- Implement in alert routing system (S5)
- Test all paths weekly via synthetic page to on-call
- Track escalation metrics: how many pages reach secondary, how many hit manager
- Adjust timing based on first month of operations
8. Build incident dashboard and status tracking (depends on: 5)
Create a single source of truth during an incident that every responder sees in real time: who is on-call, incident timeline, who said what, current status, next steps.
Dashboard displays: active incidents and their severity; who is the IC and communications lead; timeline of all events (alert fired, IC assigned, customer notified, mitigation started, resolved); Slack channel and mobile notification status; on-call rosters (who is on-call right now for each team); postmortem link as soon as incident closes.
- Integrate with alert tool (S5) to auto-populate incident creation and initial severity
- Push updates to status page and customer account managers automatically
- Log all timeline entries for audit and postmortem completeness
- Mobile-optimized so IC can work from any device
9. Write incident playbooks for each severity (depends on: 1, 2, 3)
Create a one-page (or one-screen) reference for the IC and SMEs during an incident; sequences the steps and removes ambiguity.
Each severity level gets its own playbook: who gets paged (roles, order); first questions to ask (is it real, how big, who knows); what the IC should declare in first message (status page text, account manager notification, regulatory trigger); how long before escalating to executive team; decision rules for going dark vs. continuing to update customers.
- SEV-1 playbook: immediate IC + comms + CTO notification; customer status every 5 minutes
- SEV-2 playbook: IC + comms + tech lead notification; status every 15 minutes
- SEV-3 playbook: on-call SME + comms if customer-visible; status every 30 min or as resolved
- SEV-4 playbook: on-call SME only; update customers only if promised SLA is at risk
- Include decision trees: is this SEV-1 or SEV-2? Is it our code or dependency? Escalate or containment?
10. Define internal communication workflows (depends on: 2, 3)
Specify who informs whom, in what order, via what channel (call, Slack, email) during an incident; prevents gaps like "nobody knew who was in charge for an hour."
Workflow for SEV-1: IC assigned → IC calls CTO/VP Eng and incident channel lead within 1 minute; incident declared in #incidents Slack channel with severity, IC name, service affected; SME on-call for that service joins call automatically; IC pushes updates to #incidents every 5 minutes or when material change occurs. For SEV-2: IC notifies team leads via Slack, updates #incidents every 15 min. Define escalation: if IC is unreachable, deputy IC takes over and announces it.
- Create a phone tree or on-call list accessible to responders
- Set expectations: "If you don't hear from IC in 2 minutes, call them"
- Use a single incident Slack channel per incident (auto-created by incident tool)
- Log all comms in the incident dashboard for postmortem review
11. Design customer communication and status page process (depends on: 1, 2)
Plan when and how to inform customers, account managers, and regulators; ensure 2,100 customers are not learning about outages from Twitter before you tell them.
Rules by severity: SEV-1 detected → status page updated within 3 minutes (even if root cause unknown; post "Investigating"); account managers of affected customers called within 5 minutes; regulatory notification (if payment processing down) queued for approval; customer email within 10 minutes with ETA for next update. SEV-2: status page within 10 min, account managers called within 15 min, email if affecting >10 customers. SEV-3/4: no customer communication unless SLA at risk.
- Empower Communications Lead to update status page without IC approval if delay >3 min
- Prepare templated messages for common scenarios (database failover, data pipeline stuck, service crashed)
- Route regulatory notifications through legal/compliance; don't wait for perfect root cause
- Track customer impact in real time: how many customers affected by severity
12. Establish blameless postmortem process and format (depends on: 1, 2)
Build a systematic way to learn from incidents so the same failure does not happen twice; counter the fear that admitting a mistake leads to being blamed.
Mandatory postmortems: all SEV-1 and SEV-2 incidents, within 48 hours of resolution. Optional but encouraged: SEV-3 if interesting or if >3 of same type in 30 days. Format: what was the user-visible impact and for how long; what was the root cause (not "human error" but the system condition that made error possible); timeline of discovery and response; action items with owner and deadline; blameless tone (focus on process and system design, not individual mistakes).
- Assign a facilitator (not the on-call IC) to run postmortem
- Attendees: IC, comms lead, SMEs involved, team lead, customer success if customer-facing
- Write postmortem in shared doc; make it findable (searchable, linked from incident)
- No discussion of "who screwed up"; only "why did the system allow this to happen"
13. Build action item tracking and accountability (depends on: 12)
Create a system that tracks postmortem action items so they are not forgotten; currently 11 of 64 (17%) are being tracked, leaving 53 unfinished improvements.
System: each postmortem generates action items (e.g., "add monitoring for X," "update runbook for Y," "write test for Z"). Each item gets: clear description, owner (engineer's name), due date (1–4 weeks based on priority), severity (critical = must do before similar incident happens again; important = improve next month; nice-to-have = backlog). Action items live in a dedicated Jira project visible to all teams; owners are accountable (their manager reviews quarterly). Weekly: incident commander reviews open items due that week. Monthly: each team's postmortem items reviewed in their standup.
- Export action items from postmortem document to tracking system automatically
- Require IC to sign off that an action is complete before closing
- Report on completion rate as a metric (target: >80% by month 3)
14. Define incident metrics and KPIs
Establish what "good" looks like; measure so you can improve. Target metrics for 12 months out: mean time to detect 8 minutes (vs. 22 now), customers detect first <5% of incidents (vs. 40%), MTTR 45 minutes for SEV-1 (vs. 190), SLA credits <$100k/year.
Metrics to track: (1) MTTD = time from incident start to first alert/report; disaggregate: external report vs. internal detection. (2) MTTR = time from first report to full mitigation; track by severity and by service. (3) Customer-reported incidents per month (should drop to <2 per month). (4) Alert signal-to-noise ratio (goal: <5% false positive after S6 rules). (5) On-call satisfaction (survey: would you do this again?). (6) Postmortem action completion rate. (7) Incident commander and responder utilization (hours per week per person).
- Dashboard: auto-populated from incident tool, updated daily
- Disaggregate by team and service: which teams have bad MTTR? Which service is most incident-prone?
15. Create review cadence and governance process (depends on: 14)
Establish regular rhythm to inspect the metrics, spot trends, and adjust the process itself; prevent the system from calcifying.
Weekly: incident commander and on-call lead review prior week—number of incidents, any escalations, any communication gaps. Monthly: director-level incident review—trends by service, top causes of incidents, action item status, whether severity classification is working. Quarterly: full leadership review—MTTD, MTTR, customer impact, SLA credit spend, on-call satisfaction score, any systemic changes needed. Annually: audit the entire process for SOC 2 compliance.
- Assign meeting owners: weekly = on-call lead; monthly = director of reliability; quarterly = VP Eng + CFO (SLA cost) + customer success
- Use same data dashboard (S14) for all reviews
- Publish a monthly "incident newsletter" to all engineers: what happened, what we learned, what's improving
16. Prepare SOC 2 Type II audit checklist (depends on: 1, 2, 9, 12, 13, 14, 15)
Document that the incident management system meets the control requirements for a SOC 2 audit; audit happens in 8 months, so this work builds confidence in coverage.
Audit will test: (1) Is there a defined incident response process? (2) Are roles and responsibilities clear? (3) Are incidents logged and tracked? (4) Is root cause analysis performed? (5) Are action items tracked and completed? (6) Is on-call staffing adequate? (7) Are communications timely? (8) Are postmortems documented and blameless? Create a control mapping document that links each SOC 2 requirement to your process (S1–S15). Collect evidence: incident logs, postmortem documents, action item tickets, metrics reports, training records.
- Designate a compliance owner (often a reliability lead or security engineer)
- Run a mock audit at month 6 to identify gaps
- Ensure all postmortems and incidents are retained and searchable for auditor review
17. Develop implementation and rollout plan (depends on: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16)
Create a phased timeline to roll out the incident management system across all 28 teams; avoids big-bang failure and builds credibility.
Recommended structure: Phase 1 (weeks 1–4): build and test infrastructure (S5, S8, alerting); deploy severity levels and roles (S1, S2); pick pilot teams (2–3 high-traffic teams). Phase 2 (weeks 5–12): train pilot teams, run incident drills, refine playbooks based on learning; expand to half of remaining teams. Phase 3 (weeks 13–20): full rollout to all 28 teams; continue drills; track metrics. Phase 4 (weeks 21–28): stabilize, iterate on metrics, prepare for audit.
- Assign a release manager to coordinate across teams
- Create a detailed Gantt chart with swim lanes (infra, process, training, rollout)
- Identify risks: competing priorities, engineers worried about pager burden, tool adoption friction
- Plan stakeholder engagement: weekly updates to eng leadership, monthly town halls for all engineers
18. Build training and documentation (depends on: 17)
Create role-specific education so engineers understand the new system and are confident executing during an incident.
Training tracks: (1) All engineers: 30-min video on severity levels, communication expectations, postmortem format, action items (mandatory); (2) On-call responders: 2-hour workshop on playbooks, alert tool, escalation paths, when to call manager, case studies of real incidents, mobile app walkthrough; (3) Incident commanders: 4-hour IC bootcamp covering leadership under pressure, decision-making, communicating with executives, status page discipline, postmortem facilitation, practiced drills; (4) Communications leads: templates, when to update, how to talk to customers, regulatory notification rules.
- Record videos so async teams can learn on their schedule
- Create runbooks and quick-reference cards for each role (print + digital)
- Pair new on-call engineers with experienced responder for first week
- Require IC certification before anyone joins IC rotation (pass a practical drill)
19. Execute staged rollout across teams (depends on: 18)
Progressively activate the incident management system with feedback loops at each stage; reduces risk of system-wide failure.
Wave 1 (week 6–8): 3–4 pilot teams begin on-call rotations and incident response using new system; capture feedback daily. Wave 2 (week 10–14): 8–10 additional teams, incorporating lessons from Wave 1; ensure diversity of team types (payment processing, monitoring, data pipeline, auth, etc.). Wave 3 (week 15–20): remaining teams; by now, the system is proven and less hand-holding needed.
- Daily retros with Wave 1 teams: what worked, what was confusing, what broke
- Each wave produces a "lessons learned" document that informs the next
- Track adoption metrics: how many incidents reported per team, alert quality, MTTD/MTTR
- Address resistance: engineers who are skeptical of the system, on-call burden, tool friction; assign a "change champion" in each team
20. Run incident response drills and simulations (depends on: 19)
Practice incidents in a controlled setting so responders gain confidence and gaps are found before a real crisis.
Drill 1 (week 8, pilot teams only): simulated database failover; test alert routing, IC assignment, on-call notification, status page update, postmortem. Drill 2 (week 12, all teams): full incident spanning multiple services; test escalation, communications, decision-making. Drill 3 (week 18): chaos drill with intentional confusion—what if IC is unreachable, what if comms lead is on PTO, what if an alert tool goes down? Quarterly thereafter: repeat drills with new on-call engineers.
- Schedule drills at a time when key people are available (not Friday afternoon)
- Run drill with live observers (reliability lead, manager) who provide feedback
- Measure drill success: all steps completed on time, communications clear, postmortem written
- Debrief immediately: what did we learn, what do we change?
21. Establish continuous improvement process (depends on: 15, 20)
Build a feedback loop so the incident management system evolves based on what you learn.
Monthly improvement cycle: in the monthly incident review (S15), leadership identifies one or two friction points from recent incidents (e.g., "alert tool response is slow," "postmortem meetings are boring," "action items keep slipping"). Assign an owner to investigate and propose a fix (change to playbook, tweak to tool config, additional training, process change). Test the fix with pilot team, gather feedback, roll out if successful. Annual review: can we reduce on-call burden, simplify communication, improve MTTD further? Do any teams need special handling (new services, difficult customers)?
- Keep a public backlog of improvement ideas (Jira, shared doc) so any engineer can suggest
- Quarterly all-hands: share wins ("MTTD improved 50%," "on-call satisfaction up," "zero customer-detected incidents") to build momentum
- Link improvements back to metrics (S14): show that changes actually worked
--- PROPOSAL 2 (agent deepseek-flash_initial_2, deepseek/deepseek-flash) ---
Estimated complexity: high
Success metrics: - Median time to detect reduced from 22 minutes to 5 minutes or less by month 6.
- Customer-detected incidents reduced from 40 percent to 10 percent or less of all customer-impacting incidents by month 9.
- Median time to mitigate reduced from 3 h 10 min to 60 minutes or less for SEV1 and SEV2 by month 9.
- Customer-impacting incidents reduced by at least 40 percent year over year from the 31 incident baseline.
- SLA credits paid reduced by at least 50 percent year over year from the $1.3M baseline.
- 100 percent of SEV1 incidents have a named Incident Commander within 5 minutes and a scribe within 15 minutes.
- Zero incidents in which command authority is unclear for more than 15 minutes.
- 28 of 28 teams have a documented on-call rotation, an owning service map and at least one trained on-call engineer by month 6.
- Central Incident Commander roster holds at least 40 certified ICs covering 24x7 with no single point of failure.
- 100 percent of on-call rotations are paid under a published policy by month 5.
- Monthly alert volume reduced from 3,400 to below 700, with a false-positive rate below 20 percent.
- No service exceeds 2 pages per on-call shift, measured monthly for three consecutive months.
- 100 percent of SEV1 and SEV2 postmortems published internally within 15 business days.
- At least 90 percent of postmortem action items closed within 60 days, up from 17 percent (11 of 64).
- Status-page first update published within 30 minutes on at least 95 percent of SEV1 incidents.
- Zero missed regulatory notification windows on any incident requiring notification.
- SOC 2 Type II audit passed with no findings related to incident response.
- Review cadence sustained: weekly operational review in at least 90 percent of weeks, 12 of 12 monthly reliability reviews, 4 of 4 quarterly executive reviews.
- On-call satisfaction at 70 percent or higher on the quarterly survey, with zero on-call-attributed voluntary attrition.
- 100 percent of new engineers complete the incident-response onboarding module within 30 days of joining.
Steps (21):
1. Programme charter, ownership and executive mandate
This step turns the CEO's email into a funded programme with a named owner and explicit authority. Without it, every downstream decision stalls in cross-team negotiation.
- Appoint a single accountable process owner (for example a Director of Incident Management) reporting to the CTO, with a dotted line to the COO for customer and SLA matters.
- Publish a one-page charter covering scope (all customer-impacting and money-moving incidents), decision rights, and the power to override team preferences during an active incident.
- Define the funding envelope: tooling licences, training time, exercise time and on-call compensation, with an indicative annual figure.
- Set the timeline against the SOC 2 date: a working process in four months, evidence accumulating from month two, audit-ready by month seven.
- Stand up a steering group with CTO, VP Engineering, Head of Support, Head of Compliance and one engineering manager per region.
- Agree that incident-process participation is a documented performance expectation for engineering managers, not an optional extra.
2. Baseline measurement and evidence pack (depends on: 1)
You cannot demonstrate improvement without a defensible starting point, and you cannot win the argument about noise without numbers.
- Build a 12-month incident register: date, retro-assigned severity, detection source, time to detect, time to mitigate, customer impact, services involved and SLA credits paid.
- Quantify the alert estate per tool, per team and per service; compute page-to-action ratio, list the 50 noisiest rules and count off-hours interruptions per engineer.
- Survey on-call engineers and managers on burden, fairness and clarity of escalation, targeting a response rate above 70 percent.
- Interview Support, Account Management and Sales about how customers learn of incidents and what they complain about.
- Document exactly where the current process breaks: unclear command in the two known incidents, postmortem action closure at 11 of 64, and ad-hoc status-page authorship.
- Publish the pack internally as the problem statement and retain it as management-review evidence for the SOC 2 audit.
3. Severity taxonomy and trigger matrix (depends on: 1, 2)
Severity is the keystone of the whole process. Every other rule, from paging to communications timing to postmortems, is keyed off it.
- Define four levels plus a special SEV0 for security or regulatory events: SEV1 for total or material loss of a payment path, SEV2 for degradation or single-region loss, SEV3 for limited impact with a workaround, SEV4 for internal-only issues and near-misses.
- Anchor each level in customer and money terms: customers affected, dollars at risk, regions lost, ledger integrity, settlement delay.
- Specify automatic triggers, for example loss of one AWS region, ledger write failures, a missed settlement cut-off, or payment success rate below threshold for five minutes.
- State who may declare each level (any engineer, Support or account manager may declare) and who may only recommend a downgrade (the Incident Commander alone).
- Map each level to SLA credit exposure and to the customer-visible status-page state.
- Include worked examples from the last 12 months so teams recognise their own incidents in the definitions.
- Add a review clause: the taxonomy is re-validated quarterly against real declarations.
4. Incident roles, command structure and decision rights (depends on: 2, 3)
The two incidents where nobody knew who was in charge for over an hour are the direct brief for this step.
- Define roles with one-page role cards: Incident Commander, Deputy IC, Operations Lead, Internal Communications Lead, Customer Communications Lead, Scribe, Subject-Matter Responders and an Executive Sponsor for SEV1 only.
- State the core rule plainly: the IC owns the incident, not the fix, and does not debug.
- Give the IC explicit decision rights: declaring and escalating severity, freezing changes, halting deploys, approving customer messaging and calling additional responders.
- Define minimum viable role coverage per severity: SEV1 staffs every role, SEV3 staffs an IC and a scribe only.
- Define handover discipline: maximum four-hour IC shifts on SEV1, a written handover template, and a Deputy IC nominated within 15 minutes.
- Set behaviour standards for responders: one incident channel, one bridge, no side channels, and explicit asks with a named owner and a time.
- Publish role cards on the internal wiki and link them from every paging notification.
5. Escalation, paging and incident lifecycle policy (depends on: 3, 4)
This step defines the mechanical path from an alert to a declared incident and back to normal service.
- Define lifecycle states with entry and exit criteria: Detected, Triaged, Declared, Mitigated, Resolved, Postmortem, Closed.
- Set acknowledgement targets: page acknowledged within five minutes, triage decision within fifteen, declaration within thirty for anything customer-facing.
- Define escalation ladders per layer (responder, service owner, team manager, IC on-call, VP Engineering) each with an automatic timer.
- Make escalation blameless and automatic: no responder is ever criticised for escalating, and timers fire whether or not a human asks.
- Define change freeze and rollback authority during SEV1 and SEV2, and the single condition that lifts the freeze.
- Enforce one incident, one record: the incident record is the sole source of truth for timeline, roles and communications.
- Require every SEV1 and SEV2 to produce an automatically captured timeline from channel and bridge, never one written from memory afterwards.
6. Detection strategy: SLOs, signals and customer-journey monitoring (depends on: 2, 3)
Customers detected 40 percent of incidents first. That number is the reason this step exists.
- Define SLIs and SLOs for the top 20 customer journeys, including payment initiation, settlement, ledger read and write, API availability and webhook delivery, measured per region.
- Require symptom-based alerting on those SLOs rather than cause-based alerting on infrastructure metrics.
- Add synthetic transaction monitoring from outside AWS in both regions and a third location, on a one-minute cadence for money-moving paths.
- Add ledger and shared-PostgreSQL specific signals: replication lag, connection saturation, write latency, transaction ID exhaustion and checkpoint pressure.
- Open a customer-reported path so Support and account managers can raise an incident directly, and count that path as a detection source in reporting.
- Publish a detection contract per service: every one of the 180 services needs an owner, at least one symptom alert and a documented expected detect time.
- Fund a separate resilience track to reduce shared-cluster blast radius, because better detection will not save a single shared ledger during a corruption event.
7. Alert quality standard and noise-reduction programme (depends on: 2, 3, 6)
3,400 alerts a month with 85 percent noise is the reason engineers resent the pager. Fixing it is the price of admission for everything else.
- Publish alert standards: every page must be symptom-based, actionable, owned, linked to a runbook and mapped to a severity. No runbook, no page.
- Separate paging from ticketing: only customer-impacting or imminent-impact signals may page; everything else becomes a ticket or a dashboard entry.
- Set a noise budget per team and per service, for example no service may exceed two pages per on-call shift, measured monthly.
- Define the default action for a noisy alert: fix, tune or delete within ten working days, with deletion treated as a legitimate and successful outcome.
- Introduce correlation and deduplication at the event pipeline so a single root cause produces one page instead of forty.
- Require expiry dates on every silencing rule and temporary threshold so suppression cannot become permanent blindness.
- Run a 90-day noise sprint with a visible burn-down of the top 100 noisiest rules, each owned by a named manager.
- Report page-to-action ratio per team in the monthly reliability review.
8. On-call architecture and 24x7 coverage model across 28 teams (depends on: 3, 4)
This is the hardest political step. The answer to carrying a pager for another team's code is that every team carries its own, and the platform carries the shared risk.
- Adopt a federated model: every service has exactly one owning team, and that team's primary on-call carries its own pager. No team is paged for code it does not own.
- State the consequence honestly: 16 of 28 teams currently have no on-call. They must build one or formally transfer ownership of their services to a team that will.
- Merge small or low-traffic teams into shared rotations so no rotation has fewer than six engineers, below which coverage is not sustainable.
- Create a separate, centrally staffed Incident Commander on-call roster drawn from trained senior engineers across all teams, covering 24x7.
- Define primary and secondary per rotation, with the secondary engaged only on a no-acknowledge or an explicit request.
- Define coverage across the two AWS regions and New York business hours: one global IC rotation, service on-call aligned to their service's users.
- Define the unresponsive-team path: fifteen minutes without acknowledgement escalates to the team manager, thirty minutes to the IC, who may then direct any available engineer.
- Publish a coverage matrix of all 28 teams showing services, rotation size and gaps, reviewed monthly.
- Make on-call participation an explicit expectation in engineering job levels and hiring criteria.
9. On-call compensation, wellbeing and sustainability policy (depends on: 8)
Unpaid on-call is the most cited reason for resistance. The policy must be settled before rollout, not negotiated during it.
- Move to paid on-call: a per-shift stipend or salary uplift agreed with HR and Finance and benchmarked to the New York market.
- Pay event-based compensation for incident callouts outside business hours, with a minimum call-out block.
- Provide compensatory rest: no engineer works a normal day after a night incident, and the rest day is documented, not granted as a favour.
- Cap intrusion by defining a maximum number of off-hours pages per shift, with a mandatory review triggered whenever it is exceeded.
- Define a voluntary opt-out path for engineers with genuine constraints, balanced by an explicit obligation that someone else is paid to take the shift.
- Include on-call expectation and compensation in offers and job descriptions so the commitment is set before hiring.
- Publish the policy with an effective date before any team is asked to join a new rotation.
- Review the policy every six months against actual page volumes, attrition and survey results.
10. Internal and customer communications policy with timing SLAs (depends on: 3, 4)
Today the status page is written by whoever is around. This step replaces improvisation with a clock and a named owner.
- Set internal cadence: first update within 15 minutes of declaration, then every 30 minutes for SEV1 and 60 minutes for SEV2, regardless of whether there is progress.
- Set customer cadence: status page updated within 30 minutes of a SEV1 and 60 minutes of a SEV2, and a no-new-information update is still mandatory.
- Define the channel hierarchy: status page for everyone, direct email to affected customers on SEV1, named account-manager calls for the top 50 accounts.
- Prepare templates per severity in advance with legal and compliance pre-approval, covering detection, impact, workaround, mitigation and next-update time.
- Define regulatory obligations explicitly: money transmitter and banking regulator notification windows, security breach notification, and who signs off (Compliance, not Engineering).
- Prohibit speculation: customer communications never guess at cause or blame and never commit to a root cause before the postmortem.
- Define the closing communication: resolution notice, expected SLA credit handling and the committed date for a written report.
- Assign a named Customer Communications Lead per incident with a trained deputy on every SEV1.
11. Status page, notification tooling and account-manager playbook (depends on: 10)
Policy without tooling collapses at three in the morning. This step makes publishing a five-minute action.
- Upgrade or replace the status page so components map to customer journeys rather than internal services, with subscriber control per component.
- Integrate the incident tool with the status page so the incident record drives the update and the public timeline.
- Provide one-click templates pre-filled with severity, impact language and next-update time.
- Give account managers a playbook: contact tree, what they may say, what they must not say, and how to escalate a customer question into the incident channel within minutes.
- Define the SLA credit process end to end, covering computation, approval, customer notification and finance treatment, so credits stop being a manual scramble.
- Host the status page outside the production failure domain so it survives a total platform outage.
- Test publishing during game days, including a simulated status-page outage and a simulated loss of the primary region.
12. Postmortem policy, template and blameless review process (depends on: 3, 4)
Only 11 of 64 action items closed means the postmortem ritual is currently a writing exercise. This step rebuilds it around learning and tracking.
- Make postmortems mandatory for every SEV1 and SEV2, every SEV3 with customer impact or a repeat pattern, and any near-miss the IC flags.
- Set deadlines: draft within five business days, blameless review within ten, internal publication within fifteen.
- Adopt a single template: impact, timeline, detection, response, contributing factors, what went well, what went badly, action items.
- Train a pool of blameless facilitators and require a trained facilitator for every SEV1 review.
- Prohibit counterfactual and blame language in the template, and require contributing factors across tooling, process, organisation and human factors.
- Limit action items to a small number of concrete, verifiable, owned changes with dates.
- Publish all postmortems internally by default, with security review only for genuinely sensitive material.
- Produce a customer-facing root cause report variant for SEV1 incidents, especially those affecting regulated or top-tier accounts.
13. Corrective action tracking and reliability backlog governance (depends on: 12)
A postmortem without durable action tracking is a complaint, not a control.
- Create a single reliability backlog in the engineering tracker with a mandatory label, owner, due date and link to the originating incident.
- Define closure criteria that require evidence: a merged change, a tested alert or a verified drill, never a self-reported status change.
- Protect capacity by reserving a fixed percentage of each team's sprint for reliability work, with unspent capacity visible to vice presidents.
- Run a weekly ageing review of open actions and escalate anything overdue by more than 30 days to the VP Engineering.
- Report closure rate and median age monthly, targeting more than 90 percent closed within 60 days.
- Require a repeat incident in the same area to trigger a design review rather than another action item.
14. Incident tooling consolidation and integration (depends on: 3, 5, 7, 11)
Six alerting tools and no single incident record is a structural cause of the 22-minute detection and the three-hour mitigation.
- Select one incident management platform for paging, on-call schedules, escalation policies, incident records and postmortem workflow.
- Consolidate the six alerting sources into a single event pipeline feeding that platform, with deduplication and severity mapping applied at ingest.
- Integrate with platform and ledger observability so responders see dashboards and runbooks inside the incident record.
- Integrate chat and bridge: incident channel auto-created, timeline auto-captured, decisions logged as they happen.
- Define the data model and retention required for SOC 2 evidence: who did what, when, and under whose authority.
- Run a dual-run period alongside the old tools with a defined rollback, then switch off the legacy tools on a published date.
- Budget for licences, migration effort and a two-week hardening period after cutover.
15. Training, certification and exercise programme (depends on: 4, 5, 10, 12)
A process that exists only on a wiki page fails on the first real page. Skills have to be built and tested before they are needed.
- Build a short practical curriculum: how to be on call, how to declare an incident, how to run an incident as IC, how to communicate and how to write a postmortem.
- Require certification before joining the IC on-call roster: a written assessment plus a live simulated incident.
- Train at least two certified ICs per team group so the central roster has depth across all 28 teams.
- Run monthly tabletops on realistic scenarios drawn from the last 12 months, including region loss and ledger corruption.
- Run quarterly game days with deliberately injected failure, including shared-PostgreSQL failover and status-page outage.
- Add a mandatory onboarding module for every engineer joining or transferring in, with completion records suitable for audit.
- Track training completion by team and publish it in the monthly reliability review.
16. Metrics, dashboards and review cadence (depends on: 2, 3)
The programme needs a public scoreboard, or it will quietly rot after the audit.
- Define the outcome metrics: time to detect by source, time to mitigate, percentage of incidents detected by customers (target below ten), incidents by severity and SLA credits paid.
- Define the process metrics: declaration latency, page acknowledgement rate, IC roster coverage, first-update timeliness and update-cadence adherence.
- Define the health metrics: alert volume and noise ratio per team, off-hours pages per engineer, postmortem timeliness, action closure rate and action age.
- Publish live dashboards visible to every engineer, not only to managers, refreshed daily.
- Institute a weekly operational review of 30 minutes going incident by incident, a monthly reliability review of trends and actions, and a quarterly executive review with the CEO's office.
- Baseline every metric against the S2 evidence pack and set 90-day and 12-month targets.
- Require every review to end with decisions and owners, not just numbers.
17. Pilot with volunteer teams (depends on: 5, 7, 9, 11, 12, 13, 14, 15, 16)
Do not roll out to 28 teams untested. Run the entire process end to end with a small cohort first.
- Recruit three to four volunteer teams covering a mix of criticality: one payment-path team, one ledger-adjacent team, one platform team and one low-traffic team.
- Run the complete process in the pilot: new severity scale, roles, escalation, communications, postmortems, action tracking and paid on-call.
- Instrument the pilot against the S16 metrics and compare results with the S2 baseline.
- Hold weekly retrospectives with pilot teams and iterate on the written policies, the tooling and the training.
- Fix the top issues found before any wider rollout and document what changed and why.
- Produce a pilot report with before-and-after numbers to carry into every rollout conversation.
- Set explicit pilot exit criteria: rotation coverage achieved, no unacknowledged pages over a defined period, postmortems delivered on time and actions tracked.
18. Phased rollout to all 28 teams (depends on: 13, 16, 17)
Rollout is a staged migration with readiness gates, not an email announcement.
- Sequence the 28 teams into four waves of roughly seven teams, ordered by customer impact, with three weeks between waves.
- Define a per-team readiness checklist: services mapped and owned, alerts cleaned to standard, runbooks written, rotation staffed, training complete and manager briefed.
- Hold a gate review with the process owner before each team joins, and move unready teams to the next wave with a dated remediation plan.
- Give each wave a named champion and run an internal communications cadence that explains the why using pilot numbers.
- Handle resistance directly by publishing the own-your-code-own-your-pager rule and the paid on-call mechanics before each wave, not after.
- Retire legacy tools, escalation lists and the informal status-page process at the end of each wave on a published cutover date.
- Harvest feedback formally at each wave and push accepted changes back into the policy documents through change control.
19. SOC 2 incident-response control mapping and evidence framework (depends on: 1, 3, 10, 12)
The audit will test incident response. Mapping controls now means evidence accumulates from the first real incident under the new process.
- Map the process to the relevant Trust Services Criteria for incident identification, response, evaluation of incidents and communication of security events.
- Write control statements in auditor language and name a single owner for each control.
- Specify the evidence artifact for each control: incident record, severity classification, paging log, communications log, postmortem, action tracker entry and training record.
- Set evidence retention and storage location so nothing depends on a laptop or on chat history that expires.
- Run an early walkthrough with an experienced compliance partner or the auditor's readiness team to test the design before the audit window.
- Flag controls that are blocked until rollout reaches the relevant teams, and keep a gap register with owners and dates reviewed fortnightly by the steering group.
20. SOC 2 dry run, gap remediation and audit support (depends on: 18, 19)
This step converts a good process into a provable one, a few months before the auditors arrive.
- Run an internal dry run six weeks before the audit window, sampling real incidents from the pilot and early waves against each control's evidence requirements.
- Remediate every gap found, prioritising anything that risks a qualified opinion.
- Brief every engineer who may be interviewed so they can describe the process as they actually practise it.
- Prepare the auditor package: process documentation, sample incident records, training records, on-call schedules and the action tracking register.
- Designate a single audit liaison and a small evidence-request team so requests do not land on 28 teams at once.
- Rehearse the walkthrough with the incident commander and communications roles, since auditors probe realism under pressure.
- Schedule the dry run and the audit alongside rollout waves so no team is hit with both in the same week.
21. Standing governance, process ownership and continuous improvement (depends on: 18, 20)
The classic post-audit failure is that the process freezes and then decays. This step keeps it alive.
- Establish a standing Incident Management Council chaired by the process owner, meeting monthly with engineering, support, compliance and product representation.
- Give the process owner a documented mandate to change standards, and require formal change control for any change to severity, roles, communications timings or compensation.
- Re-validate the severity taxonomy quarterly against real declarations and re-baseline the metric set annually.
- Fund a resilience roadmap separate from incident response: reducing shared-database blast radius, multi-region failover readiness and deploy safety.
- Report quarterly to the executive team on the metric set and on the top five systemic causes of incidents.
- Feed incident themes into architecture review, release engineering and hiring so the programme improves the system, not just the response.
- Refresh the training and tabletop programme annually and after any SEV1.
--- PROPOSAL 3 (agent qwen3.8-flash_initial_3, alibaba/qwen3.8-flash) ---
Estimated complexity: high
Success metrics: - MTTD reduced to under 10 minutes with >90% detection occurring before customer reporting.
- MTTR for Sev1/Sev2 incidents reduced to under 60 minutes.
- SLA credits paid reduced by 50% within 12 months.
- On-call alert volume reduced by 80% (from 3,400 to <600 actionable alerts/month).
- 100% of Sev1/Sev2 incidents have a blameless postmortem completed within 5 days.
- >90% completion rate of postmortem action items within 30 days of assignment.
- Zero ambiguity in incident command: 100% of major incidents have a designated IC logged in the timeline.
- SOC 2 Type II audit passes incident response controls with zero critical findings.
Steps (12):
1. Establish Governance and Executive Sponsorship
Secure explicit CEO and CTO backing to mandate the program, framing it as a revenue protection and compliance necessity rather than an engineering burden. Form a cross-functional Incident Management Task Force.
- Identify a single executive owner (e.g., VP of Engineering or Head of Platform) with authority over budget and staffing.
- Create a charter defining the scope: severity definitions, on-call compensation, and tool consolidation.
- Communicate the strategic priority to all 28 team leads to preempt pushback regarding pager responsibility.
2. Define Severity Matrix and Response SLAs (depends on: 1)
Create a rigorous, binary severity scale to eliminate ambiguity in escalation and communication speed. Align internal response times with the 99.95% SLA.
- **Sev1 (Critical)**: System-wide outage, data corruption, or >5% transaction failure rate. Immediate executive page.
- **Sev2 (Major)**: Single customer segment impacted or >1% transaction failure rate. High-priority team response.
- **Sev3 (Minor)**: Degraded performance or isolated bug. Standard business-hours follow-up.
- **Sev4 (Cosmetic)**: No functional impact. Backlog item.
- Define maximum detection and mitigation times for each level.
3. Restructure On-Call Organization and Compensation (depends on: 1)
Transform the current unpaid, fragmented on-call system into a professional, compensated, and centralized rotation to ensure 24x7 coverage across all 28 teams. Address the 'pushback' by making it mandatory and paid.
- Consolidate the 12 current on-call teams into a unified, rotating pool covering all 28 teams.
- Implement a fair compensation model: base stipend + hourly rate for active hours + incident bonus for Sev1/2.
- Define a strict 'Buddy System' or 'Follow-the-Sun' rotation if geographically feasible, otherwise ensure adequate team size to rotate fairly.
- Mandate that every team must participate in the rotation; no exemptions.
4. Define Incident Roles and RACI Model (depends on: 2)
Eliminate the 'who is in charge' ambiguity by pre-defining roles and establishing a clear chain of command for every incident. Roles are assigned per incident, not permanently held.
- **Incident Commander (IC)**: Single point of authority. Owns the timeline, delegates tasks, and approves communications. Must be senior.
- **Scribe**: Maintains the real-time timeline log of actions, timestamps, and decisions.
- **Communications Lead**: Owns external/status page updates and internal Slack broadcasts. Shields the IC from interruptions.
- **Technical Leads (SMEs)**: Deployed by IC to troubleshoot specific services. They do *not* coordinate with each other; they report to IC.
- Create a 'Rapid Response Team' (RRT) for high-severity incidents involving the shared PostgreSQL cluster.
5. Consolidate Monitoring and Enforce Alert Hygiene (depends on: 3)
Reduce the 3,400 monthly alerts (85% noise) to a signal-to-noise ratio that allows engineers to trust the pager. Standardize tooling.
- Consolidate the six disparate alerting tools into a single platform (e.g., PagerDuty, Opsgenie, or Datadog).
- Implement 'Actionable Alerting': If an alert pages, there must be a documented, immediate runbook action. No 'informational' pages.
- Set a hard cap on alert volume per engineer/night. Breaching the cap triggers a mandatory alert quality review.
- Deprecate legacy tools that cannot map to the severity matrix defined in S2.
6. Build Detection and Escalation Automation (depends on: 5)
Reduce the 22-minute mean time to detect (MTTD) by automating the path from metric breach to page. Ensure customers never find an issue before the internal pager does.
- Implement synthetic transactions for core payment flows to detect failures before real users do.
- Configure escalation policies: If not acknowledged in 5 minutes, auto-escalate to the team lead, then to the Incident Commander pool.
- Map specific Kubernetes and PostgreSQL metrics to the Severity Matrix (S2) to trigger automated pages.
- Create a 'Major Incident Channel' in Slack that is automatically opened and pinned when a Sev1/2 is triggered.
7. Standardize Communication Templates and Timing (depends on: 4, 6)
Codify the rules for external and internal communications to stop the 'whoever is around' approach. Link these timings to the Severity Matrix.
- **Status Page Updates**:
- Sev1: First update within 5 minutes of detection. Updates every 15 minutes.
- Sev2: First update within 30 minutes. Updates every 60 minutes.
- **Customer Direct Mail**: For Sev1 impacting specific accounts, Account Managers send pre-approved templates within 1 hour of mitigation.
- **Regulator Notification**: Define legal triggers for NY financial regulations.
- Create a 'Comms War Room' Slack channel restricted to the IC and Comms Lead.
8. Formalize the Postmortem Process (depends on: 4, 7)
Transform postmortems from optional/various formats into a mandatory, blameless, and tracked process. This directly addresses the SOC 2 audit requirement for continuous improvement.
- Mandate a postmortem for every Sev1 and Sev2 incident within 5 business days.
- Enforce a 'Blameless' charter: Focus on systemic/process failures, not individual errors.
- Adopt a standardized template: Timeline, Root Cause Analysis (5 Whys or Fishbone), Impact Analysis, and Action Items.
- Action Items must have: Owner, Due Date, and Priority. They are tracked in Jira/Linear, not the doc.
9. Implement Action Item Tracking and Governance (depends on: 8)
Solve the '11 of 64 closed' problem by integrating postmortem actions into the core engineering workflow and executive reviews. Ensure fixes are shipped.
- Integrate Jira with the Postmortem tool (e.g., OpsLevel, Rootly, or Notion API). Auto-create tickets for every AI.
- Block 'Sev1 Closure' until all high-priority Action Items are closed or explicitly deferred with executive approval.
- Weekly Engineering Leadership review of open postmortem actions older than 30 days.
- Tie action item completion rates to team performance reviews.
10. Deploy Training and Runbook Library (depends on: 6, 7, 8)
Address the skill gap and the 'pushback' by equipping engineers with the tools and knowledge to succeed on-call. Reduce anxiety through clarity.
- Audit all 180 services. Require every service to have a 'Page-Worthy' runbook.
- Run 'Game Day' (Chaos Engineering) exercises quarterly to practice role assignments (IC, Scribe, Comms) in a simulated Sev1.
- Conduct mandatory training for the new on-call rotation and compensation model.
- Provide a 'Playbook' for the Incident Commander role, including scripts for communication.
11. Phased Rollout and Change Management (depends on: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10)
Introduce the new process in phases to manage cultural resistance and technical debt, ensuring adoption before the SOC 2 audit. Do not 'big bang' the change.
- **Month 1-2**: Launch Severity Matrix and Consolidated Tooling (S2, S5). Pilot on 5 high-traffic teams.
- **Month 3-4**: Roll out new On-Call Compensation and Roles (S3, S4) to all 28 teams. Enforce new Communication Templates.
- **Month 5-6**: Enforce strict Postmortem and Action Item Tracking (S8, S9).
- **Month 7**: Final audit readiness check. Simulate SOC 2 auditor questions on incident response.
12. Define Metrics for Success and Continuous Review (depends on: 2, 6, 8, 9)
Establish a dashboard to prove the new system works, linking directly to the $1.3M savings in SLA credits. Monitor the health of the incident process itself.
- Track MTTD (Target: <5 min), MTTR (Target: <30 min for Sev2), and Detection Ratio (Target: >95% internal detection).
- Monitor 'Alert Noise' (target: <5% of alerts result in a non-actionable page).
- Track 'Action Item Completion Rate' (Target: 100% on-time closure for Sev1/2).
- Calculate 'SLA Credit Avoidance' vs. On-Call Program Cost to prove ROI to leadership.
Step-level differences computed by the tool:
Proposal 1 vs the previous-round proposal it resembles most (deepseek-flash_initial_2): 20 steps kept, added ['Executive mandate and governance structure', 'Alert consolidation and event pipeline', 'Playbooks and communication templates by severity'], removed ['Programme charter, ownership and executive mandate']
Proposal 2 vs the previous-round proposal it resembles most (deepseek-flash_initial_2): 6 steps kept, added ['Charter, mandate and the evidence clock', 'Baseline evidence and problem statement', 'Severity times class taxonomy', 'Roles, command and the never-without-an-owner rule', 'Declaration, lifecycle and escalation policy', 'Three on-call rotations across 28 teams', 'Compensation, rest and the price of opting out', 'Paging contract and alert quality', 'Communications: internal, customer and regulator', 'Customer-impact ledger and SLA credit automation', 'Postmortem policy with three artifact levels', 'Action items: capped, verifiable, with a repeat-incident rule', 'Pilot with three to four teams, using real incidents', 'Phased rollout sequenced by cost of failure', 'Audit dry run and evidence review'], removed ['Programme charter, ownership and executive mandate', 'Baseline measurement and evidence pack', 'Severity taxonomy and trigger matrix', 'Incident roles, command structure and decision rights', 'Escalation, paging and incident lifecycle policy', 'Alert quality standard and noise-reduction programme', 'On-call architecture and 24x7 coverage model across 28 teams', 'On-call compensation, wellbeing and sustainability policy', 'Internal and customer communications policy with timing SLAs', 'Status page, notification tooling and account-manager playbook', 'Postmortem policy, template and blameless review process', 'Corrective action tracking and reliability backlog governance', 'Pilot with volunteer teams', 'Phased rollout to all 28 teams', 'SOC 2 dry run, gap remediation and audit support']
Proposal 3 vs the previous-round proposal it resembles most (deepseek-flash_initial_2): 8 steps kept, added ['On-Call Architecture & Compensation Policy', 'Detection Strategy & SLOs', 'Communication Protocols & Templates', 'Postmortem Framework (Blameless)', 'Action Item Tracking & Governance', 'SOC 2 Control Mapping', 'Pilot Implementation (Wave 1)', 'Full Rollout Strategy (Waves 2-4)', 'Simulations & Game Days', 'Culture & Change Management', 'SOC 2 Dry Run & Evidence Prep', 'Continuous Improvement Loop'], removed ['Escalation, paging and incident lifecycle policy', 'Detection strategy: SLOs, signals and customer-journey monitoring', 'On-call architecture and 24x7 coverage model across 28 teams', 'On-call compensation, wellbeing and sustainability policy', 'Internal and customer communications policy with timing SLAs', 'Status page, notification tooling and account-manager playbook', 'Postmortem policy, template and blameless review process', 'Corrective action tracking and reliability backlog governance', 'Pilot with volunteer teams', 'Phased rollout to all 28 teams', 'SOC 2 incident-response control mapping and evidence framework', 'SOC 2 dry run, gap remediation and audit support', 'Standing governance, process ownership and continuous improvement']
Origin of the steps of the new proposals, matched by title by the tool (evidence for "taken"; ideas can also travel without a matching title):
Proposal 1: 1 of its 23 steps match its own previous version, 2 are new; steps 2, 3, 4, 5, 7, 8, 9, 10, 11, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23 resemble steps 2, 3, 4, 14, 6, 7, 5, 8, 9, 11, 12, 13, 16, 15, 17, 18, 19, 20, 21 of proposal 2; step 12 resembles step 7 of proposal 3
Proposal 2: 6 of its 21 steps match its own previous version, 15 are new
Proposal 3: 2 of its 20 steps match its own previous version, 9 are new; step 20 resembles step 21 of proposal 1; steps 1, 2, 3, 4, 6, 7, 13, 17 resemble steps 1, 2, 3, 4, 14, 7, 15, 16 of proposal 2
PROPOSALS OF ROUND 1 (to assess):
--- PROPOSAL 1 (agent claudeHaiku4.5_refine_1, anthropic/claude-haiku-4-5) ---
Estimated complexity: high
Success metrics: - Median time to detect reduced from 22 minutes to under 5 minutes by month 6, with >90% internal detection (vs. 40% customer-detected now).
- Median time to mitigate reduced from 3h 10min to under 60 minutes for SEV1 and SEV2 by month 9.
- Customer-impacting incidents detected by customers drop from 40% to <5% of all customer-impacting incidents.
- Alert volume reduced from 3,400 per month to <600 per month; signal-to-noise ratio improves from 15:85 to >95:5.
- SLA credits paid reduced from $1.3M to <$100K annually by month 12.
- Zero incidents with command-and-control ambiguity lasting >15 minutes; all SEV1/2 incidents have named IC logged in timeline within 5 minutes.
- Postmortem action item completion rate reaches >80% (from 11 of 64, or 17%) by month 4.
- 100% of SEV1 and SEV2 postmortems published within 15 business days by month 5.
- All 28 teams integrated into incident management system with active on-call rotations by week 20; no team unresponsive to pages for >30 minutes.
- On-call satisfaction score reaches >7/10 on survey; zero on-call-attributed voluntary attrition by month 6.
- Incident commander roster: 40+ certified ICs covering 24x7 with no single point of failure by month 4.
- Status-page first update published within 30 minutes on ≥95% of SEV1 incidents by month 3.
- SOC 2 Type II audit passes incident response controls with zero findings by month 8.
- Weekly incident review cadence sustained in ≥90% of weeks; monthly reliability reviews 12 of 12; quarterly executive reviews 4 of 4.
- 100% of new engineers complete incident-response onboarding within 30 days of joining.
Steps (23):
1. Executive mandate and governance structure
Turn the CEO email into a funded, authorized program with clear ownership and decision rights. Without executive backing, every downstream decision stalls in negotiation.
- Appoint a single program owner (e.g., Director of Incident Management) reporting to the CTO and COO.
- Publish a one-page charter covering scope (all customer-impacting incidents), authority to override team preferences during incidents, and funding for tooling, training, and on-call compensation.
- Establish a standing Incident Management Steering Group with CTO, VP Engineering, VP Support, Head of Compliance, and one engineering manager per region meeting monthly.
- Secure budget envelope: tool licenses, training time, incident-response infrastructure, and on-call compensation (estimated $400–600K annually).
- Make incident-response participation a documented performance expectation for all engineering managers and team leads.
2. Baseline measurement and evidence pack (depends on: 1)
You cannot prove improvement without defensible baseline numbers, and you cannot win arguments about noise or impact without data.
- Build a 12-month incident register: date, detection source, impact scope, time to detect, time to mitigate, SLA credits paid, and services involved.
- Audit the current alert estate: count alerts per tool, per team, per service; compute page-to-action ratio; identify top 50 noisiest rules and off-hours interruptions per engineer.
- Survey on-call engineers and managers on burden, fairness, and escalation clarity (target >70% response rate).
- Interview Support and Account Management: how do customers discover incidents, what do they complain about?
- Document the two command-ambiguity incidents: exactly when unclear who was in charge and why, how long it lasted.
- Publish this pack internally as the problem statement and retain all artifacts for SOC 2 audit evidence.
3. Severity taxonomy and trigger matrix (depends on: 2)
Severity is the keystone. Every other rule—paging, communications, postmortems, compensation—keys off it. Define four levels plus a special SEV0 for security/regulatory events.
- **SEV1 (Critical)**: Complete service outage, data corruption, or >5% payment-path failure rate for >5 min. Every minute costs money. IC required; 99.99% uptime threatened.
- **SEV2 (Major)**: Significant degradation, single region loss, or 1–5% transaction failure. IC typically required; service credit exposure.
- **SEV3 (Minor)**: Limited customer impact with workaround available, or internal issues affecting operations. On-call SME + escalation if SLA at risk.
- **SEV4 (Cosmetic)**: Observations, non-impacting bugs, alerts. Alert-driven, no escalation unless pattern emerges.
- Specify automatic triggers: region loss, ledger write failures, missed settlement window, payment success rate thresholds.
- Define who may declare (any engineer, Support, account manager) and who may downgrade (IC only).
- Include worked examples from the last 12 months so teams recognize their incidents in the definitions.
- Review and re-validate quarterly against actual declarations.
4. Incident roles, command structure, and decision rights (depends on: 3)
The two incidents with >1 hour of command ambiguity prove this step is non-negotiable. Define clear roles with explicit decision authority.
- **Incident Commander**: Owns the incident timeline, not the fix. Declares severity, decides escalation, approves communications, freezes changes, calls responders. Non-technical ICs are acceptable.
- **Deputy IC**: Shadows IC; takes over if IC unavailable. Nominated within 5 minutes of incident declaration.
- **Communications Lead**: Owns internal Slack updates and status-page messaging. Shields IC from interruptions.
- **Scribe**: Records real-time timeline: decisions, who did what, key timestamps. Not responsible for fixing.
- **Subject-Matter Responders**: Engineers with service context. Take IC direction without debate. Report only to IC.
- **Operations Lead** (SEV1 only): Coordinates across multiple responders, manages incident bridge.
- Write one-page role cards with mission, decision authority, and escalation upward. Publish on wiki and link from every paging notification.
- Define minimum viable coverage per severity: SEV1 staffs all roles; SEV2 staffs IC, Comms, Scribe; SEV3 staffs IC + Scribe.
- Establish handover discipline: maximum 4-hour IC shifts on SEV1, written handover template required.
5. Incident tooling consolidation and integration (depends on: 1, 3)
Six alert tools and ad-hoc incident records are structural causes of the 22-minute detection and 3+ hour mitigation. Consolidate to a single incident platform that is the source of truth.
- Select an incident management platform (PagerDuty, Opsgenie, Incident.io, etc.) that supports paging, schedules, escalation, incident records, and postmortem workflow.
- Requirement: the platform must integrate with observability tools, auto-create and pin incident channels in Slack, auto-capture timeline from chat, and support API-driven playbook automation.
- Plan a dual-run period alongside legacy tools with a published cutover date; define rollback criteria.
- Integrate incident record with the 180 services' monitoring and dashboards so responders see everything in one place.
- Define data retention and audit trail to satisfy SOC 2 evidence requirements: who did what, when, under whose authority.
- Budget for licenses, migration effort, and two-week hardening period post-cutover.
6. Alert consolidation and event pipeline (depends on: 5)
Replace six alert sources with a single ingestion point. Deduplicate and route alerts with minimal manual judgment, removing a major source of detection delay.
- Consolidate alert endpoints from six tools into a single event pipeline; this often sits in front of the incident platform (S5).
- Implement deduplication and correlation so a single outage triggering alerts from five monitoring tools produces one page, not five.
- Map every alert to a severity level from S3 (SEV1, SEV2, SEV3, SEV4) at ingestion.
- Log every alert for postmortem analysis and trending.
- Ensure the platform's mobile app works reliably; on-call responders need to engage from any device.
7. Detection strategy: SLOs, signals, and customer-journey monitoring (depends on: 3, 6)
Customers detected 40% of incidents first—a detection gap that must be closed. Build symptom-based alerting that detects outages before customers do.
- Define SLIs and SLOs for the top 20 customer journeys: payment initiation, settlement, ledger read/write, API availability, webhook delivery, measured per region.
- Require symptom-based alerting on SLOs, not cause-based infrastructure metrics (e.g., alert on "payment success rate <99%" not "database CPU >80%").
- Add synthetic transaction monitoring from outside AWS in both regions and a third location, one-minute cadence, for all money-moving paths.
- Add ledger-critical signals: PostgreSQL replication lag, connection saturation, write latency, transaction ID exhaustion, checkpoint pressure.
- Create a detection contract per service: owner identified, at least one symptom alert defined, expected detect time documented.
- Open a customer-reported incident path: Support and account managers can declare an incident directly, counted as a detection source in metrics.
- Fund a separate resilience roadmap to reduce shared-database blast radius, because detection improvements do not protect against ledger corruption.
8. Alert quality standards and noise-reduction program (depends on: 3, 6, 7)
3,400 monthly alerts with 85% noise is the reason engineers resent the pager. Cutting noise is the price of admission for on-call buy-in.
- Publish alert standards: every page must be symptom-based, have an immediate runbook action, be owned by a team, and map to a severity level. No runbook, no page.
- Separate paging from ticketing: only customer-impacting or imminent-impact signals trigger pages. Everything else becomes a dashboard entry or a ticket.
- Set a noise budget per team and service: no service may exceed two pages per on-call shift per month. Breaching triggers a mandatory alert quality review.
- Define the default action for a noisy alert: fix the root cause, tune the threshold, or delete it within ten working days. Deletion is a legitimate successful outcome.
- Implement automatic suppression rules: silence alerts if service auto-recovered within 30 seconds; suppress known maintenance windows; group flapping alerts (>5 in 2 min) into one page; rate-limit noisy services (max 1 alert per 5 min until condition clears).
- Require expiry dates on all silencing rules so suppression cannot become permanent blindness.
- Run a 90-day noise sprint with a visible burn-down of the top 100 noisiest rules, each owned by a named manager.
9. Escalation policies and incident lifecycle (depends on: 3, 4, 5, 6)
Define the mechanical path from alert to incident declaration to resolution. Escalation must be automatic and blameless.
- Define incident lifecycle states with clear entry/exit criteria: Detected → Triaged → Declared → Mitigated → Resolved → Postmortem → Closed.
- Set acknowledgement targets: page acknowledged in 5 minutes; triage decision (is this real?) in 15 minutes; severity declaration (is this customer-impacting?) in 30 minutes.
- Build escalation ladders: if responder does not acknowledge in 5 min, page escalates to service owner, then team manager, then IC on-call. Escalation is automatic, not manual.
- Implement escalation for SEV1: IC paged via phone call + SMS + Slack + mobile; if not acknowledged in 2 min, Deputy IC paged simultaneously; Communications Lead pinged at same time.
- Define change freeze during SEV1 and SEV2: no deployments except to fix the incident. Freeze lifts only when mitigation is confirmed.
- Enforce one incident, one record: the incident record is the sole source of truth. Auto-capture timeline from Slack and bridge; never write timeline from memory later.
- Test all escalation paths weekly via synthetic page to on-call; adjust timings based on first month of operations.
10. On-call architecture and 24x7 coverage model (depends on: 3, 4, 9)
The answer to "carrying a pager for another team's code" is that every team carries its own, and the platform carries shared risk. Design a sustainable model.
- Adopt a federated model: every service has one owning team; that team's on-call carries its service's pager. No team is paged for code it does not own.
- State the consequence clearly: 16 of 28 teams currently have no on-call. They must either build one or formally transfer service ownership to a team that will.
- Create a separate, centrally staffed Incident Commander on-call roster drawn from trained senior engineers across all teams, covering 24x7 with no single point of failure.
- Merge small or low-traffic teams into shared rotations so no rotation has fewer than six engineers (below which coverage is unsustainable).
- Define primary and secondary per slot: secondary engages only on no-acknowledge or explicit request from IC.
- Align coverage across two AWS regions and New York business hours: one global IC rotation; service on-call aligned to service users' time zones.
- Define unresponsive-team escalation: 15 min without acknowledgement escalates to team manager; 30 min escalates to IC, who may direct any available engineer.
- Publish a coverage matrix of all 28 teams showing services, rotation size, and gaps, reviewed monthly.
11. On-call compensation, wellbeing, and sustainability policy (depends on: 10)
Unpaid on-call is the most cited reason for resistance. Settle compensation before rollout, not during it. Make it sustainable.
- Introduce paid on-call: a weekly stipend ($500–1,000) while on-call, regardless of incident volume, benchmarked to New York market rates.
- Pay event-based compensation: 1.5× hourly rate for time spent mitigating incidents during off-hours (minimum one-hour block per callout).
- Provide compensatory rest: no engineer works a normal business day after a night incident requiring >2 hours mitigation. Rest day is documented, not granted as a favor.
- Cap intrusion: define maximum off-hours pages per week (e.g., no more than three per shift). Mandatory review and escalation if exceeded.
- Offer a voluntary opt-out path for engineers with genuine constraints, balanced by explicit obligation that someone else is paid to cover.
- Include on-call expectation and compensation in job descriptions and hiring conversations so commitment is understood before joining.
- Publish the policy with an effective date before any team is asked to join a rotation.
- Review the policy every six months against actual page volumes, attrition rates, and survey feedback.
12. Playbooks and communication templates by severity (depends on: 3, 4)
Playbooks remove ambiguity and decision fatigue during incidents. Templates ensure consistent, compliant messaging.
- Create a one-page (or one-screen) playbook for each severity level: who gets paged (roles, order); first questions (is it real, how big, who knows); what IC declares first (status page text, account manager notification, regulatory trigger); escalation timeline.
- **SEV1 playbook**: Immediate IC + Comms + CTO notification; customer status every 5 minutes; sample message templates.
- **SEV2 playbook**: IC + Comms + tech lead notification; status every 15 minutes; decision tree for escalation to executive team.
- **SEV3 playbook**: On-call SME + Comms if customer-visible; status every 30 minutes or as resolved.
- **SEV4 playbook**: On-call SME only; update customers only if promised SLA is at risk.
- Include decision trees: is this SEV1 or SEV2? Is it our code or dependency? Escalate or containment?
- Prepare customer-communication templates pre-approved by legal and compliance: sample language for detection, impact, workaround, mitigation phases.
- Publish playbooks on wiki and embed links in incident management platform.
13. Internal and customer communications workflows (depends on: 4, 9, 12)
Specify who informs whom, in what order, via what channel. Prevent gaps like "nobody knew who was in charge for an hour."
- **Internal cadence**: First update to #incidents Slack channel within 3 minutes of declaration (even if "investigating"). Updates every 5 minutes (SEV1), 15 minutes (SEV2), 30 minutes (SEV3) or when material change occurs.
- IC calls CTO/VP Eng and incident channel lead within 1 minute of declaration (SEV1); incident declared in Slack with severity, IC name, and service affected.
- SME on-call for the failing service joins incident bridge automatically; escalation call includes them within 5 minutes.
- Designate a single Customer Communications Lead per incident (pre-identified on-call roster) who owns external messaging exclusively. Shields IC from customer contact.
- **Customer communication**: Status page updated within 3 minutes (SEV1) or 10 minutes (SEV2) even if root cause unknown; post "Investigating" with next-update ETA. Account managers of affected top-tier customers called within 5 minutes (SEV1) with templated language.
- Create a phone tree or escalation list accessible to responders; set expectation: "If you don't hear from IC in 2 minutes, call them."
- Use a single incident Slack channel per incident (auto-created by incident tool); log all communications for postmortem review.
- Define regulatory notification path: compliance must approve before sending, but do not wait for root cause; flag incidents triggering payment-processing regulations to legal immediately.
14. Status page, customer notifications, and account-manager playbook (depends on: 5, 12, 13)
Policy without tooling collapses at 3 AM. Make publishing a five-minute action.
- Upgrade or replace status page so components map to customer journeys ("payments", "settlements", "ledger") not internal services. Allow customers to subscribe per component.
- Integrate incident tool (S5) with status page so incident record drives updates and public timeline auto-populates.
- Provide one-click templates pre-filled with severity, impact language, and next-update time; reduce typing and errors.
- Create account-manager playbook: contact tree for top 50 customers, what they may say (facts only), what they must not say (speculation, blame, false ETAs), escalation path if customer escalates.
- Define SLA credit process end to end: impact detection → credit calculation (based on duration × severity) → approval → customer notification → finance treatment. Automate where possible.
- Host status page outside production failure domain on separate infrastructure so it survives total platform outage.
- Test status-page reliability during game days (S17), including simulated status-page outage and total region loss.
15. Postmortem policy, blameless process, and facilitation (depends on: 3, 4)
Only 11 of 64 action items closed means postmortems are currently a writing exercise. Rebuild around learning and tracking.
- Make postmortems mandatory: all SEV1 and SEV2, all SEV3 with customer impact or repeat pattern, any near-miss the IC flags.
- Set deadlines: draft within 5 business days, blameless review within 10 days, internal publication within 15 days.
- Adopt a single standardized template: impact and duration, timeline (detection through resolution), root cause (not "human error" but system condition that enabled error), contributing factors, what went well, what went badly, action items.
- Train a pool of blameless facilitators; require a trained facilitator (not the IC) for every SEV1 and SEV2 review.
- Prohibit counterfactual and blame language in postmortems; require contributing factors addressing tooling, process, organization, and human factors.
- Limit action items to small number of concrete, verifiable, owned changes with dates.
- Publish all postmortems internally by default, with security review only for genuinely sensitive material (e.g., unpatched vulnerability details).
- For SEV1 incidents affecting regulated customers, produce a variant customer-facing root cause report.
16. Action item tracking, reliability backlog, and completion governance (depends on: 15)
A postmortem without durable action tracking is a complaint. Solve the 11-of-64 problem.
- Create a single reliability backlog in the engineering tracker (Jira, Linear, etc.) with mandatory label, owner, due date, and link to originating incident.
- Define closure criteria: evidence required (merged code change, tested alert, verified drill) not self-reported status.
- Protect capacity: reserve a fixed percentage of each team's sprint (e.g., 10%) for reliability work; track unspent capacity and report to VP Engineering.
- Run a weekly ageing review of open actions; escalate anything overdue by >30 days to team lead and VP Engineering.
- Require IC or postmortem facilitator to sign off on action completion.
- Report completion rate and median action age in monthly incident review (target: >90% closed within 60 days).
- If the same service repeats an incident in the same area, trigger a design review rather than another action item; break the cycle.
17. Metrics, dashboards, and review cadence (depends on: 2, 3)
Measure to prove the system works. Publish dashboards so everyone sees the scoreboard.
- Define outcome metrics: time to detect (by source, target <5 min internally detected); time to mitigate SEV1/SEV2 (target <60 min); customer-detected incidents per month (target <2); SLA credits paid (target <$100K/year by month 12).
- Define process metrics: declaration latency, page acknowledgement rate, IC roster coverage (no single point of failure), first-update timeliness (% within SLA), update-cadence adherence.
- Define health metrics: alert volume and noise ratio per team, off-hours pages per engineer per month, postmortem timeliness, action closure rate and median action age.
- Build live dashboards visible to every engineer (not just managers); auto-populate from incident tool, update daily.
- Institute review cadence: weekly operational review (30 min, prior week incidents), monthly reliability review (trends, top causes, action status), quarterly executive review (CEO's office, SLA cost, systemic changes).
- Baseline every metric against S2 evidence pack; set 90-day and 12-month targets.
- End every review with decisions and owners, not just numbers.
18. Training, certification, and exercise program (depends on: 4, 12, 13, 14, 15)
A process that exists only on a wiki fails on the first real page. Build skills before deployment.
- Build curriculum: how to be on-call, how to declare an incident, how to run incidents as IC, how to communicate, how to write blameless postmortems.
- Create role-specific tracks: (1) All engineers: 30-min video on severity levels, communication expectations, postmortem format, action items (mandatory); (2) On-call responders: 2-hour workshop on playbooks (S12), alert tool (S5), escalation paths (S9), when to call manager, case studies, mobile app walkthrough; (3) Incident commanders: 4-hour IC bootcamp on leadership under pressure, decision-making, communicating with executives, status-page discipline, postmortem facilitation, practiced drills; (4) Communications leads: 2-hour training on templates, update timings, how to talk to customers, regulatory rules.
- Require certification before joining IC on-call roster: written assessment plus live simulated incident (pass/fail).
- Record videos so async teams can learn on their schedule; create runbooks and quick-reference cards (print + digital); pair new on-call engineers with experienced responder for first week.
- Run monthly tabletop exercises on realistic scenarios from the prior 12 months: region loss, ledger corruption, cascading failures.
- Run quarterly game days with intentional failure injection (database failover, status-page outage, alert tool downtime); include all on-call roles.
19. Pilot with volunteer teams (depends on: 5, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 18)
Do not roll out untested to 28 teams. Run the entire process end to end with a small cohort first.
- Recruit 3–4 volunteer teams covering a mix of criticality: one payment-path team, one ledger-adjacent team, one platform team, one low-traffic team.
- Run complete process in pilot: new severity scale (S3), consolidated tooling (S5, S6), roles (S4), escalation (S9), communications (S13, S14), postmortems (S15), action tracking (S16), paid on-call (S11), training (S18), metrics (S17).
- Instrument pilot against S17 metrics; compare results with S2 baseline.
- Hold weekly retrospectives with pilot teams; iterate on written policies, tooling, training based on feedback.
- Fix top issues found before wider rollout; document what changed and why.
- Produce pilot report with before/after numbers (MTTD, MTTR, alert noise, action completion rate) to carry into rollout conversations.
- Set explicit pilot exit criteria: rotation coverage achieved, zero unacknowledged pages over 2 weeks, all postmortems delivered on time, >80% of action items tracked.
20. Phased rollout to all 28 teams (depends on: 16, 17, 19)
Rollout is a staged migration with readiness gates, not an email announcement.
- Sequence the 28 teams into four waves of ~7 teams each, ordered by customer-impact criticality and readiness; space waves three weeks apart.
- Define per-team readiness checklist: services mapped and owned, alerts cleaned to standard (S8), runbooks written, rotation staffed, training complete, manager briefed.
- Hold gate review with program owner before each team joins; move unready teams to next wave with dated remediation plan.
- Assign named champion per wave; run internal communications cadence explaining why using pilot numbers from S19.
- Handle resistance directly: publish the "own-your-code, own-your-pager" rule and paid on-call mechanics before each wave, not after.
- Retire legacy tools, informal escalation lists, and ad-hoc status-page process at end of each wave on published cutover date.
- Harvest feedback formally at each wave and push accepted process changes through change control.
21. SOC 2 control mapping and evidence framework (depends on: 1, 3, 13, 15)
The audit will test incident response. Mapping controls now means evidence accumulates from the first real incident under the new process.
- Map the process to Trust Services Criteria for incident identification, response, evaluation, and communication of security incidents.
- Write control statements in auditor language; name a single owner per control.
- Specify the evidence artifact for each control: incident record, severity classification, paging log, communications log, postmortem, action tracker entry, training record.
- Define evidence retention and storage location (not on laptops, not on chat history that expires); plan for audit access.
- Identify which controls are blocked until certain rollout waves complete; keep a gap register with owners and review fortnightly with steering group.
- Run early walkthrough with an experienced compliance partner or pre-audit readiness team to test control design before the audit window.
22. SOC 2 dry run, gap remediation, and audit support (depends on: 20, 21)
Convert a good process into a provable one, a few months before auditors arrive. Prove the system works at scale.
- Schedule a dry run six weeks before audit window, sampling real incidents from completed waves against each control's evidence requirements.
- Remediate every gap found; prioritize anything risking a qualified opinion.
- Brief every engineer who may be interviewed so they can describe the process as practiced.
- Prepare auditor package: process documentation, sample incident records, training records, on-call schedules, action tracking register, alert quality metrics.
- Designate a single audit liaison and small evidence-request team so requests do not scatter across teams.
- Rehearse IC and Communications Lead roles under interview conditions; auditors probe realism under pressure.
- Ensure all postmortems, incidents, and evidence are retained, searchable, and accessible to auditors for the required audit period.
23. Standing governance, process ownership, and continuous improvement (depends on: 20, 22)
The classic post-audit failure is the process freezing and decaying. Lock in continuous improvement as a permanent structure.
- Establish a standing Incident Management Council chaired by the program owner, meeting monthly with engineering, support, compliance, and product representation.
- Give program owner documented mandate to change standards; require formal change control for any change to severity, roles, communications timings, or compensation.
- Re-validate severity taxonomy quarterly against real declarations; re-baseline metrics annually.
- Feed incident themes into architecture review and hiring so the program improves the system, not just the response.
- Report quarterly to executive team on metric set (MTTD, MTTR, SLA credits, customer-detected %) and top five systemic causes of incidents.
- Fund a resilience roadmap separate from incident response: reducing shared-database blast radius, multi-region failover, deploy safety.
- Refresh training and tabletop program annually and after any SEV1.
- Keep public backlog of improvement ideas (Jira, shared doc) so any engineer can suggest; publish monthly "incident newsletter" to all engineers with wins and learnings.
--- PROPOSAL 2 (agent deepseek-flash_refine_2, deepseek/deepseek-flash) ---
Estimated complexity: high
Success metrics: - Median time to detect falls from 22 minutes to under 8 minutes by month 6, and under 5 minutes for money-moving journeys by month 9.
- Share of customer-impacting incidents first detected by customers falls from 40% to under 10% by month 9.
- Median time to mitigate for SEV-1 falls from 3 h 10 min to under 60 minutes by month 9.
- 100% of SEV-1 and SEV-2 incidents have a named Triage Owner within 5 minutes and a named Incident Commander within 15 minutes, from month 4.
- Zero incidents in which command authority is unclear for more than 15 minutes, measured monthly from month 4.
- Monthly page volume falls from 3,400 to under 600, with a false-positive rate under 15%, by month 5.
- No service exceeds two pages per on-call shift for three consecutive months, measured by month 6.
- All 28 teams have a documented rotation, a service ownership map and at least one trained on-call engineer by month 6.
- Incident Commander roster holds at least 12 certified ICs covering 24x7 with no uncovered week, from month 4.
- 100% of on-call shifts are paid under a published policy from month 2.
- 100% of SEV-0, SEV-1 and SEV-2 postmortems are published internally within 15 business days, from month 5.
- Postmortem action items closed within 60 days rise from 17% to over 90%, with median age under 30 days, by month 6.
- At least one incident review or near-miss report is filed per team per quarter, from month 6.
- Status page first update is published within 30 minutes on at least 95% of SEV-1 incidents, from month 4.
- Zero missed regulatory notification windows on any incident requiring notification.
- The customer-impact ledger is complete for 100% of customer-impacting incidents from month 5, with SLA credits reconciled monthly.
- Annual SLA credits paid fall at least 50% from the $1.3M baseline within 12 months.
- Customer-impacting incidents fall at least 40% year over year from the 31-incident baseline.
- On-call satisfaction is at or above 7 out of 10, with zero voluntary attrition attributed to on-call, measured quarterly from month 6.
- Review cadence is sustained: weekly operational review in at least 90% of weeks, 12 of 12 monthly reliability reviews, 4 of 4 quarterly executive reviews.
- SOC 2 Type II is passed with zero findings related to incident response.
- 100% of new engineers complete the incident-response onboarding module within 30 days of joining.
- 100% of the 180 services have a named owner, a detection contract and at least one symptom-based alert held to the paging contract by month 7.
Steps (21):
1. Charter, mandate and the evidence clock
This step turns the CEO's email into a funded programme with one accountable owner and explicit authority, and it starts the SOC 2 clock on day one.
- Appoint a single accountable process owner — a Director of Incident Management reporting to the CTO, with a dotted line to the COO for customer and SLA matters.
- Publish a one-page charter: scope (every customer-impacting, money-moving, security and data-integrity event across all 28 teams and both regions), decision rights during an active incident, and the power to freeze deploys and override team preferences.
- Fix the funding envelope up front: tooling licences, training and drill time, and on-call compensation, with an indicative annual figure and the expected return in avoided SLA credits.
- **Start the evidence clock immediately.** A SOC 2 Type II report covers a period of operating effectiveness, so evidence cannot be backfilled; an audit in eight months means the operating period begins in the first weeks of this programme.
- Stand up a steering group of CTO, VP Engineering, Head of Support, Head of Compliance and two engineering managers, meeting fortnightly.
- Make incident-process participation a documented performance expectation for every engineering manager, not an optional extra.
- Agree the timeline explicitly: a working process in four months, evidence accumulating from month two, audit-ready by month seven.
2. Baseline evidence and problem statement (depends on: 1)
You cannot demonstrate improvement without a defensible starting point, and you cannot win the argument about noise without numbers.
- Build a 12-month incident register and re-classify all 31 customer-impacting incidents: date, severity, class, detection source, time to detect, time to mitigate, customers affected, services involved and SLA credits paid.
- Run an alert census per tool, per team and per service: total volume, page-to-action ratio, off-hours interruptions per engineer, and the 50 noisiest rules with a named owner.
- Build a silent-failure register: incidents in which no internal alert fired at all. This is the number that explains the 40% customer-detected rate.
- Survey every on-call engineer and manager on burden, fairness, escalation clarity and pay expectations, targeting above 70% response.
- Interview Support, Account Management and Sales about how customers learn of incidents, what they complain about, and what they are promised.
- Reconstruct the two command-ambiguity incidents minute by minute, from alert to mitigation, to find exactly where ownership lapsed.
- Publish the pack internally as the problem statement and retain it as management-review evidence for the audit.
3. Control mapping and evidence architecture (depends on: 1)
Most programmes leave compliance to the end. This one maps controls in the first month, because the mapping determines what the process must capture from day one.
- Map the process to the relevant Trust Services Criteria for incident identification, evaluation, response, containment, recovery and communication of events (CC7.1 to CC7.5, CC2.2, CC2.3, CC4.1, CC3.x).
- Write each control as a plain-language statement with one named owner and its evidence artifact: incident record, severity classification, paging log, communications log, postmortem, action tracker entry, training record.
- Define the **golden incident file**: one auto-assembled, single-click export per incident containing timeline, roles, severity, communications, postmortem, actions and closure evidence.
- Set retention, storage location and immutability so no control depends on a laptop, a private Slack channel or chat history that expires.
- Keep a gap register with owners and dates, reviewed fortnightly by the steering group.
- Run an early design walkthrough with the auditor's readiness team inside the first 90 days, to test the design before building on it.
4. Severity times class taxonomy (depends on: 2)
Severity alone is not enough. Severity decides how loud the response is; class decides which playbook runs and what must never be traded away.
- Define severity by impact in customer and money terms: customers affected, dollars at risk, regions lost, ledger integrity, settlement delay, data loss and SLA exposure.
- Use SEV-0 for security, privacy and regulatory events; SEV-1 for total or material loss of a payment path; SEV-2 for degradation or single-region loss; SEV-3 for limited impact with a workaround; SEV-4 for internal-only issues and near-misses.
- Define response classes: Availability, Performance, Data integrity and Ledger, Security and Privacy, Third-party dependency, Settlement and Financial, and Process failure.
- **Class can raise a response, never lower it.** A SEV-2 data-integrity incident gets SEV-1 posture, because integrity failures are not recoverable by moving faster.
- Specify automatic triggers: loss of one AWS region, ledger write failures, replication lag above threshold, a missed settlement cut-off, payment success rate below threshold for five minutes.
- State who may declare: any engineer, Support agent or account manager. State who may downgrade: the Incident Commander alone.
- Map every level to its SLA credit exposure and to the customer-visible status page state.
- Include worked examples from the last 12 months so each of the 28 teams recognises its own incidents, and re-validate the taxonomy quarterly against real declarations.
5. Roles, command and the never-without-an-owner rule (depends on: 4)
The two incidents where nobody was in charge for over an hour did not fail at declaration. They failed in the gap before it, when an alert had fired and no one owned it.
- Create one-page role cards: Triage Owner, Incident Commander, Deputy IC, Operations Lead, Internal Communications Lead, Customer Communications Lead, Scribe, Subject-Matter Responders and Executive Sponsor for SEV-1 only.
- Introduce the **Triage Owner** rule: from the moment a page is acknowledged, that person owns the incident until an IC takes over or the incident is stood down. There is never an unowned minute between first page and close.
- State the core rule plainly: the IC owns the incident, not the fix, and does not debug. An IC who starts troubleshooting has abandoned command.
- Give the IC explicit decision rights: declare and escalate severity, freeze changes, halt deploys, approve customer messaging, and pull in any engineer in the company.
- Define minimum viable staffing per severity: SEV-1 staffs every role; SEV-2 staffs IC, scribe, comms and responders; SEV-3 staffs an IC and a scribe only.
- Set handover discipline: four-hour maximum IC shifts on SEV-1, a written handover template, and a deputy named within 15 minutes of declaration.
- Set behaviour standards for responders: one incident channel, one bridge, no side channels, and asks phrased with a named owner and a time.
- Publish the role cards on the internal wiki and link them from every paging notification.
6. Declaration, lifecycle and escalation policy (depends on: 5)
This step defines the mechanical path from an alert to a declared incident and back to normal service, and it removes judgment calls from the worst moments.
- Define lifecycle states with entry and exit criteria: Detected, Triaged, Declared, Mitigated, Resolved, Postmortem, Closed — plus a Watch state with a hard 30-minute timer, after which the incident is either declared or stood down.
- **Make declaring free.** A declaration that turns out to be a false alarm is closed as a false declaration, with no blame and no follow-up, and it is tracked as a metric so the cost of caution stays visible.
- Apply the ambiguity rule: if two responders disagree about whether this is an incident, it is an incident.
- Set acknowledgement and declaration targets: page acknowledged within five minutes, triage decision within fifteen, declaration within thirty for anything customer-facing.
- Build automatic escalation ladders with timers at every layer: responder, service owner, team manager, IC on-call, VP Engineering. Escalation never requires a human decision and is never criticised.
- Define the unresponsive-team path: fifteen minutes without acknowledgement escalates to the team manager, thirty minutes to the IC, who may then direct any available engineer.
- Define change freeze and rollback authority during SEV-1 and SEV-2, and the single named condition that lifts the freeze.
- Enforce one incident, one record, with the timeline auto-captured from the channel and bridge rather than written from memory afterwards.
- Add the two-simultaneous-incidents rule: a second IC is designated, and no IC ever runs two incidents at once.
7. Detection strategy: journeys, synthetic signals and customer-report intake (depends on: 4)
Customers detected 40% of incidents first. That number is the reason this step exists, and it is fixed by measuring customer journeys rather than infrastructure.
- Define SLIs and SLOs for the top 20 customer journeys — payment initiation, authorisation, settlement, ledger read and write, API availability, webhook delivery and payout — measured per region.
- Require symptom-based alerting on those SLOs instead of cause-based alerting on infrastructure metrics.
- Add synthetic transaction monitoring from outside AWS in both regions plus a third location, on a one-minute cadence for money-moving paths.
- Add ledger and shared-PostgreSQL specific signals: replication lag, connection saturation, write latency, lock waits, transaction ID exhaustion, checkpoint pressure and bloat.
- Open a customer-report intake: Support and account managers can raise an incident directly, every customer report creates an incident record, and the customer-report path is counted as a detection source in all reporting.
- Apply the detection-gap rule: whenever a customer reports an incident before internal monitoring, a detection-gap ticket is opened automatically and owned by the responsible team.
- Publish a detection contract per service: every one of the 180 services needs an owner, at least one symptom alert, and a documented expected detect time.
- Measure current detect time per journey, set targets, and run a detection drill per team: break something in staging and see whether it pages before a human notices.
8. Three on-call rotations across 28 teams (depends on: 5)
The objection is that engineers will not carry a pager for another team's code. The answer is not to argue with it, but to build three rotations so the objection becomes structurally impossible.
- Run a **Service On-Call** rotation per team, covering only that team's own services. No engineer is ever paged for code their team does not own.
- Run a **Platform Duty** rotation for genuinely shared infrastructure: the shared PostgreSQL cluster, Kubernetes, networking, CI/CD and observability. This is nobody's product code, so it gets its own paid rotation, staffed from platform teams plus volunteers from other teams.
- Run a central **Incident Commander** roster of 12 to 16 certified senior engineers drawn from across all 28 teams, covering 24x7 on one-week shifts with a primary and a secondary.
- State the consequence honestly: 16 of 28 teams have no on-call today. Each must either build a rotation or formally transfer ownership of its services to a team that will, with the transfer dated and recorded.
- Merge small or low-traffic teams into shared rotations so no rotation has fewer than six engineers, the level below which coverage stops being sustainable.
- Cap load in the scheduling tool: no engineer is on-call more than two weeks per quarter, enforced by configuration rather than negotiation.
- Publish a coverage matrix for all 28 teams showing services, owners, rotation size and gaps, reviewed monthly.
- Write on-call participation into engineering job levels and hiring criteria, so the commitment is set before someone joins rather than after.
9. Compensation, rest and the price of opting out (depends on: 8)
Unpaid on-call is the most cited reason for resistance. The policy must be settled before rollout, not negotiated during it.
- Move immediately to paid on-call: a per-shift stipend benchmarked to the New York market, agreed with HR and Finance, and published with an effective date before any team is asked to join a new rotation.
- Pay event-based compensation for out-of-hours callouts, with a minimum call-out block and a 1.5x rate for time actually spent mitigating.
- Provide documented compensatory rest: no engineer works a normal day after a night incident, and the rest day is policy rather than a favour granted by a manager.
- Cap intrusion with a maximum number of off-hours pages per shift, and trigger a mandatory review whenever the cap is exceeded.
- Allow opt-out but **put a price on it**: an engineer may step out of a rotation, and their team must buy coverage from the paid pool at a published internal rate. This turns a cultural argument into a visible budget decision.
- Publish an explicit amnesty: incident records, near-miss reports and false declarations are never used in performance reviews. Only failure to report is a performance issue.
- Protect against gaming: pay attaches to shifts and callouts, never to incident counts.
- Review the policy every six months against real page volumes, attrition and survey results.
10. Paging contract and alert quality (depends on: 2, 7)
3,400 alerts a month at 85% noise is why engineers resent the pager. Fixing that is the price of admission for everything else in this plan.
- Publish the paging contract: every page must be symptom-based, actionable, owned, mapped to a severity and a class, and linked to a runbook. **No runbook, no page**, enforced by a CI check on the alert definition.
- Separate paging from ticketing: only customer-impacting or imminent-impact signals may wake a human; everything else becomes a ticket or a dashboard entry.
- Set a page budget per team and per service: a maximum number of pages per on-call shift. Breaching it auto-opens a remediation ticket with the engineering manager as owner.
- Put new alerts on two-week probation: a new rule runs as a ticket only and becomes a pager only after it has proved actionable, so teams stop being woken by untested rules.
- Define the default action for a noisy alert: fix, tune or delete within ten working days, with deletion treated as a legitimate and celebrated outcome and the deleted-alert count published.
- Correlate and deduplicate at the ingest pipeline so one root cause produces one page instead of forty.
- Require an expiry date on every silencing rule and temporary threshold, so suppression cannot quietly become permanent blindness.
- Run a 90-day noise sprint with a public burn-down of the top 100 noisiest rules, each owned by a named manager, and report page-to-action ratio per team monthly.
11. Incident tooling consolidation (depends on: 5, 10)
Six alerting tools and no single incident record is a structural cause of the 22-minute detection and the three-hour mitigation.
- Select one incident platform for paging, schedules, escalation policies, incident records, postmortem workflow and status-page integration, and choose it within four weeks.
- Consolidate the six alert sources into a single event pipeline feeding that platform, with deduplication, correlation and severity and class mapping applied at ingest.
- Integrate observability, runbooks, chat and voice bridges so responders see dashboards and runbooks inside the incident record and the timeline is captured automatically.
- Implement the golden incident file export defined in S3, so audit evidence is a single click rather than a reconstruction.
- Run a dual-run period alongside the legacy tools with a defined rollback, then switch the legacy tools off on a published date.
- Host the status page outside the production failure domain so it survives a total platform outage, and test that during a game day.
- Make the platform usable from a phone: an Incident Commander must be able to run a SEV-1 from a mobile device at 3am.
12. Communications: internal, customer and regulator (depends on: 4, 5)
Today the status page is written by whoever is around. This step replaces improvisation with a clock, a named owner and a pre-cleared template.
- Set the internal cadence: first update within 15 minutes of declaration, then every 30 minutes for SEV-1 and hourly for SEV-2, whether or not there is progress.
- Never let the status page be how an employee learns of an incident: internal communication leads, external follows.
- Set the customer cadence: status page updated within 30 minutes of a SEV-1 and 60 minutes of a SEV-2, with a no-new-information update still mandatory.
- Define the channel hierarchy: status page for all 2,100 customers, direct email to affected customers on SEV-1, and named account-manager calls for the top 50 accounts.
- Prepare templates per severity and class in advance, pre-approved by Legal and Compliance, each with the next-update time built in.
- Forbid speculation: customer communications never guess at cause, never assign blame, and never commit to a root cause before the postmortem.
- Define the closing communication: resolution notice, how SLA credits will be handled, and the committed date for the written report.
- Build a **regulator clock matrix**: for each event type, which regulator, which window, who signs off, and the shortest clock that drives the first action. Cover NYDFS Part 500, money-transmitter and banking notification, security breach notification, card-network rules, and public-company disclosure where applicable.
- Route every regulatory notification through Compliance, never Engineering, and pre-clear the templates.
- Publish a customer-facing root-cause report for SEV-1 incidents, especially for regulated and top-tier accounts.
- Assign a named Customer Communications Lead plus a trained deputy on every SEV-1.
13. Customer-impact ledger and SLA credit automation (depends on: 12)
The $1.3M in credits is a symptom of having no single record of customer impact. This step creates one, and makes it do four jobs at once.
- Maintain one durable customer-impact ledger per incident: which customers, which journeys, from when to when, and the estimated credit.
- Use that single record for customer communications, SLA credit computation, regulatory reporting and the annual review, so nobody reconciles two versions of the same outage.
- Automate the credit path end to end — computation, approval, customer notification and finance treatment — so credits stop being a manual scramble three weeks later.
- Reconcile accrued against paid credits monthly and report the result in the executive review.
- Track credit avoidance against programme cost, so the funding case is a number rather than an argument.
14. Postmortem policy with three artifact levels (depends on: 5)
Postmortems currently happen for some incidents, in various formats. This step makes them mandatory where they matter and light where they do not.
- Make postmortems mandatory for every SEV-0, SEV-1 and SEV-2, every SEV-3 with customer impact or a repeat pattern, every near-miss the IC flags, and every incident in which the process itself failed.
- Use three levels so the ritual matches the weight: a lightweight async incident review for SEV-3 and SEV-4, a standard facilitated postmortem for SEV-2, and a full review with an executive sponsor for SEV-1.
- Set deadlines: draft within five business days, blameless review within ten, internal publication within fifteen.
- Adopt one template: impact, timeline, detection, response, contributing factors across tooling, process, organisation and human factors, what went well, what went badly, and actions.
- Train a pool of blameless facilitators and require a trained facilitator for every SEV-1 review.
- Prohibit counterfactual and blame language in the template, and specifically ban the phrase human error as a root cause — the question is always what made the error possible.
- Publish all postmortems internally by default, with security review only for genuinely sensitive material, and produce a customer-facing root-cause variant for SEV-1.
- Treat postmortems as the learning product of the process, not as a compliance artifact.
15. Action items: capped, verifiable, with a repeat-incident rule (depends on: 14)
Eleven of 64 action items closed is not a tracking problem. It is a generation problem: the process produces more actions than the organisation can absorb.
- **Cap each postmortem at three action items.** Anything beyond three goes into a ranked reliability backlog, not into the postmortem.
- Require each action to have a named human owner, a due date, and a definition of done that is an artifact: a merged change, a new alert that fired in a drill, or a test that fails without the fix.
- Prohibit self-reported closure. Closure requires the artifact, signed off by the process owner or the Incident Commander.
- Keep one reliability backlog in the engineering tracker with a mandatory label and a link back to the originating incident.
- Reserve a fixed percentage of every team's sprint capacity for reliability work, with unspent capacity visible to the vice presidents.
- Apply the repeat-incident rule: a second incident with the same contributing factor triggers a design review owned by an architect, not another action item.
- Run a weekly ageing review of open actions and escalate anything more than 30 days overdue to the VP of Engineering.
- Target more than 90% of actions closed within 60 days and a median age under 30 days, reported monthly by team.
16. Training, certification and drills (depends on: 5, 6, 12, 14)
A process that exists only on a wiki page fails on the first real page. Skills have to be built and tested before they are needed.
- Build a short practical curriculum: how to be on call, how to triage and declare, how to run an incident as IC, how to communicate, and how to write a postmortem.
- Require certification before joining the IC roster: a written assessment plus a live simulated incident.
- Certify at least two ICs per team group, so the central roster has depth across all 28 teams and no holiday week is uncovered.
- Run monthly tabletops drawn from the last 12 months of real incidents, including region loss, ledger write failure and a missed settlement cut-off.
- Run quarterly game days with deliberately injected failure, including shared-PostgreSQL failover, status-page outage and alerting-pipeline outage.
- **Drill the process's own failure modes**, not just technical ones: IC unreachable, comms lead on PTO, two simultaneous SEV-1s, a paging storm, and a false alarm that burns an hour.
- Audit the incident process for single points of failure: who is the only person who can do each critical task, and what happens in their holiday week.
- Add a mandatory onboarding module for every engineer joining or transferring in, with completion records kept in an audit-ready form.
- Track training completion by team and publish it in the monthly reliability review.
17. Metrics and review cadence (depends on: 2, 4, 15)
Establish what good looks like, and measure it in a way that makes people report more incidents rather than fewer.
- Define outcome metrics: time to detect by source, time to mitigate by severity and class, percentage of incidents detected by customers, incidents by severity and class, SLA credits paid, and customer-impact minutes.
- Define process metrics, because the process itself has SLOs: IC assigned within five minutes, page acknowledgement rate, first-update timeliness, update-cadence adherence, postmortem timeliness, action closure rate, action median age and IC roster coverage.
- Define health metrics: alert volume and noise ratio per team, off-hours pages per engineer, on-call satisfaction, training completion, and percentage of services with a detection contract.
- **Never publish incident count as a team metric.** It rewards hiding incidents. Publish reporting metrics instead — near-misses filed, detection gaps closed, false declarations made — alongside the outcome metrics.
- Publish live dashboards visible to every engineer, refreshed daily, with each metric baselined against the S2 evidence pack and given a 90-day and a 12-month target.
- Run a 30-minute weekly operational review incident by incident, a monthly reliability review of trends and actions, and a quarterly executive review with the CEO's office covering customer impact, credits and the top five systemic causes.
- Hold a quarterly review of the process itself: what in the process wasted time, what confused responders, and what should be deleted.
- Require every review to end with decisions and named owners, never with numbers alone.
18. Pilot with three to four teams, using real incidents (depends on: 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17)
Do not roll out to 28 teams untested, and do not rely on drills alone when real incidents are available and better.
- Recruit three to four volunteer teams spanning criticality: a payment-path team, a ledger-adjacent team, the platform team that owns the shared PostgreSQL cluster, and a low-traffic team.
- Run the entire process end to end in the pilot: the new severity and class scale, roles, escalation, communications, postmortems, action tracking and paid on-call.
- Treat real incidents during the pilot as the primary training material. Hold a retro within 48 hours of each pilot incident, run by the process owner, while the friction is fresh.
- Instrument the pilot against the S2 baseline and publish before-and-after numbers.
- Hold weekly retrospectives with the pilot teams and change the written policies, the tooling and the training in response, documenting what changed and why.
- Set explicit exit criteria: rotation coverage achieved, zero unacknowledged pages over a defined period, postmortems delivered on time, and actions tracked to closure.
- Produce a pilot report that every later rollout conversation starts from.
19. Phased rollout sequenced by cost of failure (depends on: 15, 18)
Rollout is a staged migration with readiness gates, not an email announcement. The sequencing matters more than the schedule.
- Sequence the 28 teams into four waves of roughly seven, ordered by **incident density and customer-journey ownership**, not by size or alphabet: the teams whose incidents cost the most SLA credits go first, because their improvement is the visible proof.
- Leave three weeks between waves and define a per-team readiness checklist: services mapped and owned, alerts cleaned to the paging contract, runbooks written, rotation staffed to at least six engineers, training complete, manager briefed, and compensation in effect.
- Hold a gate review with the process owner before each team joins, and move an unready team to the next wave with a dated remediation plan rather than granting an exception.
- Give each wave a named champion and run an internal communications cadence that explains the why using the pilot's numbers.
- Handle resistance directly and early: publish the own-your-code-own-your-pager rule and the paid on-call mechanics before each wave, never after.
- Hold a retro within 48 hours of each wave's first incident under the new process, and feed accepted changes back through change control.
- Retire legacy tools, escalation lists and the informal status-page process at the end of each wave on a published cutover date, so two processes never run in parallel for long.
- Sequence the waves so no team faces both a rollout and the audit dry run in the same week.
20. Audit dry run and evidence review (depends on: 3, 19)
This step converts a good process into a provable one, a few months before the auditors arrive.
- Run an internal dry run six weeks before the audit window, sampling real incidents from the pilot and the early waves against each control's evidence requirements.
- Remediate every gap found, prioritising anything that risks a qualified opinion, and re-test the remediated control against the same sample.
- Brief every engineer who may be interviewed so they can describe the process as they actually practise it, not as it is written.
- Prepare the auditor package: process documentation, sample incident records, paging logs, communication logs, training records, on-call schedules and the action tracking register.
- Designate one audit liaison and a small evidence-request team, so requests do not land on 28 teams simultaneously.
- Rehearse the walkthrough with an incident commander and a communications lead, because auditors probe realism under pressure, not documents.
- Keep the audit liaison and the process owner as the same person where possible, so accountability for the control is also accountability for the evidence.
21. Standing governance, resilience roadmap and continuous improvement (depends on: 19, 20)
The classic post-audit failure is that the process freezes and then decays. This step keeps it alive and separates incident response from the engineering that prevents incidents.
- Establish a standing Incident Management Council chaired by the process owner, meeting monthly with engineering, support, compliance and product representation.
- Give the process owner a documented mandate to change standards, and require formal change control for any change to severity, classes, roles, communication timings or compensation.
- Re-validate the severity and class taxonomy quarterly against real declarations, and re-baseline the whole metric set annually.
- Fund a **resilience roadmap separate from incident response**: reduce the shared PostgreSQL blast radius, improve multi-region failover readiness, and tighten deploy safety. Better incident management cannot save a single shared ledger during a corruption event.
- Report quarterly to the executive team on the metric set and the top five systemic causes of incidents, and feed those themes into architecture review, release engineering and hiring.
- Keep a public backlog of process improvement ideas so any engineer can propose one, and change the process from it every quarter.
- Share wins publicly each quarter — faster detection, fewer customer-detected incidents, credits avoided — to sustain momentum after the audit.
- Refresh the training and tabletop programme annually and immediately after any SEV-1.
--- PROPOSAL 3 (agent qwen3.8-flash_refine_3, alibaba/qwen3.8-flash) ---
Estimated complexity: high
Success metrics: - Median Time to Detect (MTTD) < 10 minutes.
- Median Time to Mitigate (MTTR) < 45 minutes for SEV-1.
- Customer-detected incidents < 5% of total incidents.
- Monthly alert volume < 400 actionable alerts (90% reduction in noise).
- SLA credit payouts < $100k annually.
- Postmortem action item completion rate > 90% within 30 days.
- 100% of SEV-1 incidents have a designated IC and Scribe.
- On-call engineer satisfaction score > 4.0/5.0.
- Zero critical findings in SOC 2 Type II audit regarding incident response.
Steps (20):
1. Program Charter & Executive Mandate
Secure explicit authority to mandate the incident management program, transforming it from an engineering initiative into a business-critical operational standard.
* Appoint a Director of Incident Management (DIM) reporting to the CTO, with a dotted line to the Head of Compliance for audit readiness.
* Publish a charter defining scope: all customer-impacting incidents, payment path failures, and data integrity issues.
* Establish a steering committee (CTO, VP Eng, Head of Support, Legal) to resolve cross-team conflicts and approve tooling budgets.
* Define the 'Own Your Code, Own Your Pager' principle as a non-negotiable engineering value, directly addressing the 28-team resistance.
2. Baseline Audit & Evidence Framework (depends on: 1)
Create a defensible starting point for metrics and a structured evidence repository for the SOC 2 Type II audit.
* Consolidate the last 12 months of incident data into a single register, flagging the 31 customer-impacting incidents.
* Map current alert sources to specific services to quantify the 3,400 monthly alerts and identify the top 50 noisiest rules.
* Conduct a survey of on-call engineers to measure current sentiment and identify specific pain points in the unpaid rotations.
* Designate a secure, immutable storage location for all future incident artifacts (timelines, logs, postmortems) to satisfy audit requirements.
3. Severity Taxonomy & Trigger Matrix (depends on: 1)
Define unambiguous severity levels tied to specific triggers, communication SLAs, and executive notifications to eliminate classification debates.
* **SEV-1 (Critical):** System-wide outage, data loss risk, or >5% revenue-impacting transactions. Triggers immediate IC, VP Eng, and legal notification.
* **SEV-2 (Major):** Significant degradation or region-wide latency. Triggers IC and Team Lead.
* **SEV-3 (Minor):** Localized issue with workaround. Triggers on-call engineer.
* **SEV-4 (Internal):** Low-priority noise. Triggers ticket only.
* Map each severity to specific SLA credit exposures to align technical response with financial risk.
4. Incident Roles & Command Structure (depends on: 3)
Standardize the Incident Command System (ICS) to prevent the 'ambiguity of ownership' that caused previous hour-long delays.
* **Incident Commander (IC):** Single point of decision authority; does not debug. Required for SEV-1/2.
* **Scribe:** Maintains the real-time timeline log; mandatory for SEV-1.
* **Comms Lead:** Owns status page and internal broadcasts; shields IC from external noise.
* **SMEs:** Technical responders focused solely on diagnosis/mitigation under IC direction.
* Create role cards with clear 'Do' and 'Don't' responsibilities for quick reference during chaos.
5. On-Call Architecture & Compensation Policy (depends on: 4)
Establish a sustainable, paid 24x7 on-call model that respects ownership boundaries and compensates for disruption.
* Adopt a federated model: each of the 28 teams maintains its own SME rotation for owned services.
* Create a central IC rotation staffed by senior engineers across teams to ensure command coverage.
* Implement a paid policy: base stipend per shift, enhanced pay for call-outs, and mandatory recovery time after SEV-1 night shifts.
* Define the 'Unowned Service' rule: services without a dedicated team rotation must be transferred to an owning team or decommissioned.
6. Tooling Consolidation & Integration (depends on: 2, 5)
Replace the six fragmented alerting tools with a single Incident Management Platform to ensure a unified source of truth.
* Select a platform (e.g., PagerDuty + Statuspage + Slack/Teams) capable of handling 180 services.
* Integrate Kubernetes and PostgreSQL monitoring to feed alerts into the central platform.
* Automate the creation of incident channels and timelines upon alert acknowledgment.
* Ensure the status page is decoupled from the production environment to remain available during outages.
7. Alert Quality & Noise Reduction Program (depends on: 6)
Implement strict alert hygiene rules to reduce the 3,400 monthly alerts and restore trust in the pager.
* **Rule:** All paging alerts must be symptom-based (customer impact) and have a linked runbook.
* **Rule:** Implement deduplication and rate-limiting at the ingestion layer.
* **Sprint:** Conduct a 90-day noise sprint to fix, tune, or delete the top 100 noisiest rules.
* **Metric:** Enforce a 'Page-to-Action' ratio target; if an alert does not result in human action, it must be demoted to ticket-only.
8. Detection Strategy & SLOs (depends on: 3, 6)
Shift from infrastructure metrics to Service Level Objectives (SLOs) to detect customer-impacting issues faster than clients can.
* Define SLIs for critical payment paths: Transaction Success Rate, API Latency, Settlement Lag.
* Implement synthetic transaction monitoring from external vantage points in both AWS regions.
* Establish error budgets; breach of error budget triggers automatic review and potential feature freezes.
* Integrate customer-reported signals (support tickets, social media mentions) into the detection workflow as low-priority alerts.
9. Communication Protocols & Templates (depends on: 4, 6)
Define who talks, when, and what they say to stop ad-hoc status updates and protect brand reputation.
* **Status Page:** SEV-1 initial update within 15 mins; subsequent updates every 30 mins. Use legal-approved templates.
* **Internal:** IC broadcasts to #exec-leadership for SEV-1 every hour.
* **Regulatory:** Define triggers for NY financial regulator notifications; assign Legal/Compliance responsibility for filing.
* **Client Success:** Provide Account Managers with a 'Known Issues' playbook and real-time access to status page APIs for high-value accounts.
10. Postmortem Framework (Blameless) (depends on: 9)
Standardize the learning process to ensure consistent quality and adherence to the 'Blameless' culture.
* **Mandatory:** All SEV-1 and SEV-2 incidents require a postmortem within 5 business days.
* **Format:** Fixed template including Impact, Timeline, Root Cause, Contributing Factors, and Action Items.
* **Review:** Facilitated by a neutral party (not the IC) to enforce blameless language and prevent finger-pointing.
* **Publication:** All postmortems published internally on the Wiki with full searchability.
11. Action Item Tracking & Governance (depends on: 10)
Close the loop on incident resolution by enforcing rigorous tracking of corrective actions to fix the 17% completion rate.
* Automatically create Jira tickets for every action item identified in the postmortem.
* **Enforcement:** SEV-1 incident cannot be marked 'Closed' until high-priority actions are completed or deferred with VP approval.
* **Review:** Weekly review of overdue actions in the Engineering Leadership standup.
* **Metric:** Track 'Mean Time to Remediation' for action items as a key health indicator.
12. SOC 2 Control Mapping (depends on: 2, 10)
Proactively map the new incident processes to SOC 2 Trust Services Criteria to ensure audit readiness.
* Map S4 (Roles), S9 (Comms), and S10 (Postmortems) to Security and Availability criteria.
* Define 'Evidence of Operation' for each control (e.g., automated timeline logs, signed-off postmortems).
* Identify gaps between current state and audit requirements; assign remediation tasks to the DIM.
* Establish a quarterly internal compliance review to test control effectiveness before the Type II audit.
13. Training & Certification Curriculum (depends on: 4, 6)
Equip all engineers with the skills to operate within the new framework, reducing anxiety and improving response quality.
* **Universal Training:** 1-hour module on severity levels and tools for all engineers.
* **IC Certification:** Mandatory workshop and simulation for engineers joining the central IC rotation.
* **Runbook Review:** Each team must update and validate runbooks for their top 3 critical alerts.
* **Onboarding:** Include incident response basics in the engineering onboarding checklist.
14. Pilot Implementation (Wave 1) (depends on: 7, 8, 9, 13)
Deploy the new process to a controlled subset of high-traffic teams to validate assumptions before broad rollout.
* Select 3 teams: Payments Core, Ledger/API, and one Infrastructure team.
* Run the full cycle for 6 weeks: Alerts, IC handover, Comms, Postmortems.
* Collect feedback daily and adjust playbooks and tooling configurations based on real-world friction.
* Measure success using the baseline metrics from S2 (MTTD, Noise, Satisfaction).
15. Full Rollout Strategy (Waves 2-4) (depends on: 14)
Gradually migrate the remaining 25 teams using lessons learned from the pilot to manage change resistance.
* **Wave 2 (Month 3):** Deploy to 8 remaining critical customer-facing teams.
* **Wave 3 (Month 4):** Deploy to internal platform and data teams.
* **Wave 4 (Month 5):** Deploy to remaining low-traffic teams and legacy services.
* Maintain a 'Help Desk' for new on-call engineers to troubleshoot process questions during their first shift.
16. Simulations & Game Days (depends on: 15)
Test the resilience of the process and the tools under controlled failure conditions.
* **Tabletop Exercises:** Quarterly scenario walks-throughs for leadership and ICs (e.g., 'What if the primary region fails?').
* **Chaos Engineering:** Inject failures in non-production or canary environments to test alert accuracy and runbook validity.
* **Communication Drills:** Simulate SEV-1 to test the speed of status page updates and internal notification paths.
* Document findings in postmortems and create action items for identified weaknesses.
17. Metrics Dashboard & Executive Review (depends on: 2, 15)
Establish a continuous feedback loop to monitor the health of the incident management system.
* Build a real-time dashboard tracking MTTD, MTTR, Alert Volume, and SLA Credits.
* **Weekly:** Operational review of new incidents and action items with the DIM and Team Leads.
* **Monthly:** Executive review with CTO/VP Eng focusing on trends, noise ratios, and on-call fatigue.
* Publish 'Reliability Reports' to all engineers to celebrate wins and share learnings.
18. Culture & Change Management (depends on: 5, 15)
Address the cultural shift required to move from 'hero mode' to 'process mode' and mitigate pushback.
* Highlight success stories where the new process reduced toil or prevented customer churn.
* Enforce the 'Blameless' mantra by publicly correcting any leadership communication that implies blame.
* Recognize and reward effective ICs and engineers who improve runbooks or alert quality.
* Address 'Pager Fatigue' proactively by monitoring call-out volume and rotating engineers off if thresholds are breached.
19. SOC 2 Dry Run & Evidence Prep (depends on: 12, 17)
Conduct a mock audit six months out to identify gaps in evidence retention or process execution.
* Sample 10 recent incidents and verify all required artifacts (timeline, logs, postmortem, actions) are present and accessible.
* Interview on-call engineers to ensure they can describe the process and their roles without hesitation.
* Remediate any 'Control Failures' identified during the dry run.
* Prepare the 'Audit Readiness' package for the external auditors.
20. Continuous Improvement Loop (depends on: 17, 19)
Institutionalize the evolution of the incident process to prevent stagnation.
* Review the Severity Taxonomy quarterly to ensure it reflects current system behavior and customer expectations.
* Update Tooling integrations annually based on vendor roadmaps and new AWS/K8s features.
* Revise Compensation Policy annually based on market data and internal fairness reviews.
* Set annual 'Reliability Goals' for the engineering organization to drive long-term focus.
Your answer has these parts:
- "round_summary": one or two sentences on how the round went as a whole.
- "converging": true if the proposals of this round are more similar to each other than those of the previous round, false otherwise.
- "differences": a short list (four items at most) of what still separates the proposals of this round, naming the proposals and the steps concerned.
- "influences": a short list (five items at most) of the flows of ideas of the round: who took what from whom, and what nobody took, naming proposals and steps ("P1 and P3 took P2's CDC sync (steps 2 and 13)").
- "proposals": one entry per proposal of round 1, each with:
- "proposal": its number,
- "assessment": "improved", "worsened", "mixed" or "unchanged" with respect to its previous version ("no_previous_version" if that agent produced nothing in the previous round),
- "what_changed": how it improved or worsened and why: two or three short sentences, or a list,
- "improvements": a list of concrete gains (specific steps, metrics, structure),
- "regressions": a list of concrete losses (dropped steps, vaguer metrics, broken dependencies...),
- "taken": the ideas this proposal visibly adopted from the OTHER proposals of round 0 (not from its own previous version): one entry per idea with "from_proposal" (the number of the proposal it came from), "steps" (the numbers of the steps of that proposal where the idea lives, as listed above; empty if it is not tied to specific steps), "what" (the idea, one sentence) and "why" (how it was used or adapted, one sentence),
- "rejected": the ideas of the OTHER proposals of round 0 that this proposal visibly declined: an explicit contradiction, or a prominent idea it saw and left out while taking the opposite approach. Same fields; "why" gives the evidence (what the proposal does instead). Do not list mere omissions without evidence; an empty list is a valid answer.
[ROUND 2]
[SYSTEM]
You are an expert reviewer of multi-agent planning processes.
Several LLM agents drafted plans for a task, refined them over a number of rounds while seeing each other's proposals, and finally voted for the best one.
Be exhaustive but precise: name concrete steps, ideas and metrics, never generalities. Judge plans by their fitness for the task as stated, their realism, their completeness, the soundness of their order and dependencies, how measurable their success is and how they handle things going wrong.
You are an impartial evaluator, not a chronicler: assess the proposals and the process on their merits, never rationalise what happened or assume that the outcome was right.
After your analysis, answer in the requested structure.
Every text field you write will be read by a busy person who skims. Make it easy to skim: short sentences and short paragraphs; when you name several things, prefer a list to a paragraph, with sub-items when an item has parts, but keep a single fact as a sentence; lead with the point and then the evidence; name proposals and steps by number (P2, step 4); no preamble, no repetition of the question, no closing summary; bold at most one key phrase per item or paragraph. Text fields accept Markdown: a blank line between paragraphs, "- " for lists, **bold**.
[HUMAN]
Task given to the agents: "A B2B payments platform in New York (260 engineers in 28 teams, 2,100 customers, $4B processed a month, a 99.95% SLA with service credits) runs about 180 services on Kubernetes across two AWS regions, with a shared PostgreSQL cluster for the ledger. In the last 12 months: 31 customer-impacting incidents, median time to detect 22 minutes (customers detected 40% of them first), median time to mitigate 3 h 10 min, $1.3M paid in SLA credits, two incidents in which nobody was sure who was in charge for over an hour, and a CEO email about "outages we hear about from clients". On-call exists in 12 of the 28 teams, unpaid, fed by alerts from six different tools — 3,400 alerts a month, 85% of them noise; there is no severity scale, status-page updates are written by whoever is around, and postmortems happen for some incidents, in various formats, with action items that are rarely tracked (11 of 64 closed). A SOC 2 Type II audit in eight months will test incident response. Engineers push back against "carrying a pager for other teams' code".
Define the incident management process: severity levels and what each one triggers; roles (incident commander, communications, scribe, subject-matter responders) and how they are staffed 24x7 across 28 teams; the on-call structure, rotations, compensation and the rules for alert quality; detection and escalation paths; internal and customer communications (status page, account managers, regulators when required) with their timings; postmortems (when mandatory, format, blameless review, ownership and tracking of actions); the metrics and reviews that show whether it works; and how it is introduced across the teams without waiting for the audit."
This is round 2, a refinement round: every agent received ALL the proposals of round 1 and wrote a new plan, improving on them or taking a different approach. By convention, the previous version of proposal N is proposal N of round 1, written by the same model.
PROPOSALS OF ROUND 1 (the previous versions):
--- PROPOSAL 1 (agent claudeHaiku4.5_refine_1, anthropic/claude-haiku-4-5) ---
Estimated complexity: high
Success metrics: - Median time to detect reduced from 22 minutes to under 5 minutes by month 6, with >90% internal detection (vs. 40% customer-detected now).
- Median time to mitigate reduced from 3h 10min to under 60 minutes for SEV1 and SEV2 by month 9.
- Customer-impacting incidents detected by customers drop from 40% to <5% of all customer-impacting incidents.
- Alert volume reduced from 3,400 per month to <600 per month; signal-to-noise ratio improves from 15:85 to >95:5.
- SLA credits paid reduced from $1.3M to <$100K annually by month 12.
- Zero incidents with command-and-control ambiguity lasting >15 minutes; all SEV1/2 incidents have named IC logged in timeline within 5 minutes.
- Postmortem action item completion rate reaches >80% (from 11 of 64, or 17%) by month 4.
- 100% of SEV1 and SEV2 postmortems published within 15 business days by month 5.
- All 28 teams integrated into incident management system with active on-call rotations by week 20; no team unresponsive to pages for >30 minutes.
- On-call satisfaction score reaches >7/10 on survey; zero on-call-attributed voluntary attrition by month 6.
- Incident commander roster: 40+ certified ICs covering 24x7 with no single point of failure by month 4.
- Status-page first update published within 30 minutes on ≥95% of SEV1 incidents by month 3.
- SOC 2 Type II audit passes incident response controls with zero findings by month 8.
- Weekly incident review cadence sustained in ≥90% of weeks; monthly reliability reviews 12 of 12; quarterly executive reviews 4 of 4.
- 100% of new engineers complete incident-response onboarding within 30 days of joining.
Steps (23):
1. Executive mandate and governance structure
Turn the CEO email into a funded, authorized program with clear ownership and decision rights. Without executive backing, every downstream decision stalls in negotiation.
- Appoint a single program owner (e.g., Director of Incident Management) reporting to the CTO and COO.
- Publish a one-page charter covering scope (all customer-impacting incidents), authority to override team preferences during incidents, and funding for tooling, training, and on-call compensation.
- Establish a standing Incident Management Steering Group with CTO, VP Engineering, VP Support, Head of Compliance, and one engineering manager per region meeting monthly.
- Secure budget envelope: tool licenses, training time, incident-response infrastructure, and on-call compensation (estimated $400–600K annually).
- Make incident-response participation a documented performance expectation for all engineering managers and team leads.
2. Baseline measurement and evidence pack (depends on: 1)
You cannot prove improvement without defensible baseline numbers, and you cannot win arguments about noise or impact without data.
- Build a 12-month incident register: date, detection source, impact scope, time to detect, time to mitigate, SLA credits paid, and services involved.
- Audit the current alert estate: count alerts per tool, per team, per service; compute page-to-action ratio; identify top 50 noisiest rules and off-hours interruptions per engineer.
- Survey on-call engineers and managers on burden, fairness, and escalation clarity (target >70% response rate).
- Interview Support and Account Management: how do customers discover incidents, what do they complain about?
- Document the two command-ambiguity incidents: exactly when unclear who was in charge and why, how long it lasted.
- Publish this pack internally as the problem statement and retain all artifacts for SOC 2 audit evidence.
3. Severity taxonomy and trigger matrix (depends on: 2)
Severity is the keystone. Every other rule—paging, communications, postmortems, compensation—keys off it. Define four levels plus a special SEV0 for security/regulatory events.
- **SEV1 (Critical)**: Complete service outage, data corruption, or >5% payment-path failure rate for >5 min. Every minute costs money. IC required; 99.99% uptime threatened.
- **SEV2 (Major)**: Significant degradation, single region loss, or 1–5% transaction failure. IC typically required; service credit exposure.
- **SEV3 (Minor)**: Limited customer impact with workaround available, or internal issues affecting operations. On-call SME + escalation if SLA at risk.
- **SEV4 (Cosmetic)**: Observations, non-impacting bugs, alerts. Alert-driven, no escalation unless pattern emerges.
- Specify automatic triggers: region loss, ledger write failures, missed settlement window, payment success rate thresholds.
- Define who may declare (any engineer, Support, account manager) and who may downgrade (IC only).
- Include worked examples from the last 12 months so teams recognize their incidents in the definitions.
- Review and re-validate quarterly against actual declarations.
4. Incident roles, command structure, and decision rights (depends on: 3)
The two incidents with >1 hour of command ambiguity prove this step is non-negotiable. Define clear roles with explicit decision authority.
- **Incident Commander**: Owns the incident timeline, not the fix. Declares severity, decides escalation, approves communications, freezes changes, calls responders. Non-technical ICs are acceptable.
- **Deputy IC**: Shadows IC; takes over if IC unavailable. Nominated within 5 minutes of incident declaration.
- **Communications Lead**: Owns internal Slack updates and status-page messaging. Shields IC from interruptions.
- **Scribe**: Records real-time timeline: decisions, who did what, key timestamps. Not responsible for fixing.
- **Subject-Matter Responders**: Engineers with service context. Take IC direction without debate. Report only to IC.
- **Operations Lead** (SEV1 only): Coordinates across multiple responders, manages incident bridge.
- Write one-page role cards with mission, decision authority, and escalation upward. Publish on wiki and link from every paging notification.
- Define minimum viable coverage per severity: SEV1 staffs all roles; SEV2 staffs IC, Comms, Scribe; SEV3 staffs IC + Scribe.
- Establish handover discipline: maximum 4-hour IC shifts on SEV1, written handover template required.
5. Incident tooling consolidation and integration (depends on: 1, 3)
Six alert tools and ad-hoc incident records are structural causes of the 22-minute detection and 3+ hour mitigation. Consolidate to a single incident platform that is the source of truth.
- Select an incident management platform (PagerDuty, Opsgenie, Incident.io, etc.) that supports paging, schedules, escalation, incident records, and postmortem workflow.
- Requirement: the platform must integrate with observability tools, auto-create and pin incident channels in Slack, auto-capture timeline from chat, and support API-driven playbook automation.
- Plan a dual-run period alongside legacy tools with a published cutover date; define rollback criteria.
- Integrate incident record with the 180 services' monitoring and dashboards so responders see everything in one place.
- Define data retention and audit trail to satisfy SOC 2 evidence requirements: who did what, when, under whose authority.
- Budget for licenses, migration effort, and two-week hardening period post-cutover.
6. Alert consolidation and event pipeline (depends on: 5)
Replace six alert sources with a single ingestion point. Deduplicate and route alerts with minimal manual judgment, removing a major source of detection delay.
- Consolidate alert endpoints from six tools into a single event pipeline; this often sits in front of the incident platform (S5).
- Implement deduplication and correlation so a single outage triggering alerts from five monitoring tools produces one page, not five.
- Map every alert to a severity level from S3 (SEV1, SEV2, SEV3, SEV4) at ingestion.
- Log every alert for postmortem analysis and trending.
- Ensure the platform's mobile app works reliably; on-call responders need to engage from any device.
7. Detection strategy: SLOs, signals, and customer-journey monitoring (depends on: 3, 6)
Customers detected 40% of incidents first—a detection gap that must be closed. Build symptom-based alerting that detects outages before customers do.
- Define SLIs and SLOs for the top 20 customer journeys: payment initiation, settlement, ledger read/write, API availability, webhook delivery, measured per region.
- Require symptom-based alerting on SLOs, not cause-based infrastructure metrics (e.g., alert on "payment success rate <99%" not "database CPU >80%").
- Add synthetic transaction monitoring from outside AWS in both regions and a third location, one-minute cadence, for all money-moving paths.
- Add ledger-critical signals: PostgreSQL replication lag, connection saturation, write latency, transaction ID exhaustion, checkpoint pressure.
- Create a detection contract per service: owner identified, at least one symptom alert defined, expected detect time documented.
- Open a customer-reported incident path: Support and account managers can declare an incident directly, counted as a detection source in metrics.
- Fund a separate resilience roadmap to reduce shared-database blast radius, because detection improvements do not protect against ledger corruption.
8. Alert quality standards and noise-reduction program (depends on: 3, 6, 7)
3,400 monthly alerts with 85% noise is the reason engineers resent the pager. Cutting noise is the price of admission for on-call buy-in.
- Publish alert standards: every page must be symptom-based, have an immediate runbook action, be owned by a team, and map to a severity level. No runbook, no page.
- Separate paging from ticketing: only customer-impacting or imminent-impact signals trigger pages. Everything else becomes a dashboard entry or a ticket.
- Set a noise budget per team and service: no service may exceed two pages per on-call shift per month. Breaching triggers a mandatory alert quality review.
- Define the default action for a noisy alert: fix the root cause, tune the threshold, or delete it within ten working days. Deletion is a legitimate successful outcome.
- Implement automatic suppression rules: silence alerts if service auto-recovered within 30 seconds; suppress known maintenance windows; group flapping alerts (>5 in 2 min) into one page; rate-limit noisy services (max 1 alert per 5 min until condition clears).
- Require expiry dates on all silencing rules so suppression cannot become permanent blindness.
- Run a 90-day noise sprint with a visible burn-down of the top 100 noisiest rules, each owned by a named manager.
9. Escalation policies and incident lifecycle (depends on: 3, 4, 5, 6)
Define the mechanical path from alert to incident declaration to resolution. Escalation must be automatic and blameless.
- Define incident lifecycle states with clear entry/exit criteria: Detected → Triaged → Declared → Mitigated → Resolved → Postmortem → Closed.
- Set acknowledgement targets: page acknowledged in 5 minutes; triage decision (is this real?) in 15 minutes; severity declaration (is this customer-impacting?) in 30 minutes.
- Build escalation ladders: if responder does not acknowledge in 5 min, page escalates to service owner, then team manager, then IC on-call. Escalation is automatic, not manual.
- Implement escalation for SEV1: IC paged via phone call + SMS + Slack + mobile; if not acknowledged in 2 min, Deputy IC paged simultaneously; Communications Lead pinged at same time.
- Define change freeze during SEV1 and SEV2: no deployments except to fix the incident. Freeze lifts only when mitigation is confirmed.
- Enforce one incident, one record: the incident record is the sole source of truth. Auto-capture timeline from Slack and bridge; never write timeline from memory later.
- Test all escalation paths weekly via synthetic page to on-call; adjust timings based on first month of operations.
10. On-call architecture and 24x7 coverage model (depends on: 3, 4, 9)
The answer to "carrying a pager for another team's code" is that every team carries its own, and the platform carries shared risk. Design a sustainable model.
- Adopt a federated model: every service has one owning team; that team's on-call carries its service's pager. No team is paged for code it does not own.
- State the consequence clearly: 16 of 28 teams currently have no on-call. They must either build one or formally transfer service ownership to a team that will.
- Create a separate, centrally staffed Incident Commander on-call roster drawn from trained senior engineers across all teams, covering 24x7 with no single point of failure.
- Merge small or low-traffic teams into shared rotations so no rotation has fewer than six engineers (below which coverage is unsustainable).
- Define primary and secondary per slot: secondary engages only on no-acknowledge or explicit request from IC.
- Align coverage across two AWS regions and New York business hours: one global IC rotation; service on-call aligned to service users' time zones.
- Define unresponsive-team escalation: 15 min without acknowledgement escalates to team manager; 30 min escalates to IC, who may direct any available engineer.
- Publish a coverage matrix of all 28 teams showing services, rotation size, and gaps, reviewed monthly.
11. On-call compensation, wellbeing, and sustainability policy (depends on: 10)
Unpaid on-call is the most cited reason for resistance. Settle compensation before rollout, not during it. Make it sustainable.
- Introduce paid on-call: a weekly stipend ($500–1,000) while on-call, regardless of incident volume, benchmarked to New York market rates.
- Pay event-based compensation: 1.5× hourly rate for time spent mitigating incidents during off-hours (minimum one-hour block per callout).
- Provide compensatory rest: no engineer works a normal business day after a night incident requiring >2 hours mitigation. Rest day is documented, not granted as a favor.
- Cap intrusion: define maximum off-hours pages per week (e.g., no more than three per shift). Mandatory review and escalation if exceeded.
- Offer a voluntary opt-out path for engineers with genuine constraints, balanced by explicit obligation that someone else is paid to cover.
- Include on-call expectation and compensation in job descriptions and hiring conversations so commitment is understood before joining.
- Publish the policy with an effective date before any team is asked to join a rotation.
- Review the policy every six months against actual page volumes, attrition rates, and survey feedback.
12. Playbooks and communication templates by severity (depends on: 3, 4)
Playbooks remove ambiguity and decision fatigue during incidents. Templates ensure consistent, compliant messaging.
- Create a one-page (or one-screen) playbook for each severity level: who gets paged (roles, order); first questions (is it real, how big, who knows); what IC declares first (status page text, account manager notification, regulatory trigger); escalation timeline.
- **SEV1 playbook**: Immediate IC + Comms + CTO notification; customer status every 5 minutes; sample message templates.
- **SEV2 playbook**: IC + Comms + tech lead notification; status every 15 minutes; decision tree for escalation to executive team.
- **SEV3 playbook**: On-call SME + Comms if customer-visible; status every 30 minutes or as resolved.
- **SEV4 playbook**: On-call SME only; update customers only if promised SLA is at risk.
- Include decision trees: is this SEV1 or SEV2? Is it our code or dependency? Escalate or containment?
- Prepare customer-communication templates pre-approved by legal and compliance: sample language for detection, impact, workaround, mitigation phases.
- Publish playbooks on wiki and embed links in incident management platform.
13. Internal and customer communications workflows (depends on: 4, 9, 12)
Specify who informs whom, in what order, via what channel. Prevent gaps like "nobody knew who was in charge for an hour."
- **Internal cadence**: First update to #incidents Slack channel within 3 minutes of declaration (even if "investigating"). Updates every 5 minutes (SEV1), 15 minutes (SEV2), 30 minutes (SEV3) or when material change occurs.
- IC calls CTO/VP Eng and incident channel lead within 1 minute of declaration (SEV1); incident declared in Slack with severity, IC name, and service affected.
- SME on-call for the failing service joins incident bridge automatically; escalation call includes them within 5 minutes.
- Designate a single Customer Communications Lead per incident (pre-identified on-call roster) who owns external messaging exclusively. Shields IC from customer contact.
- **Customer communication**: Status page updated within 3 minutes (SEV1) or 10 minutes (SEV2) even if root cause unknown; post "Investigating" with next-update ETA. Account managers of affected top-tier customers called within 5 minutes (SEV1) with templated language.
- Create a phone tree or escalation list accessible to responders; set expectation: "If you don't hear from IC in 2 minutes, call them."
- Use a single incident Slack channel per incident (auto-created by incident tool); log all communications for postmortem review.
- Define regulatory notification path: compliance must approve before sending, but do not wait for root cause; flag incidents triggering payment-processing regulations to legal immediately.
14. Status page, customer notifications, and account-manager playbook (depends on: 5, 12, 13)
Policy without tooling collapses at 3 AM. Make publishing a five-minute action.
- Upgrade or replace status page so components map to customer journeys ("payments", "settlements", "ledger") not internal services. Allow customers to subscribe per component.
- Integrate incident tool (S5) with status page so incident record drives updates and public timeline auto-populates.
- Provide one-click templates pre-filled with severity, impact language, and next-update time; reduce typing and errors.
- Create account-manager playbook: contact tree for top 50 customers, what they may say (facts only), what they must not say (speculation, blame, false ETAs), escalation path if customer escalates.
- Define SLA credit process end to end: impact detection → credit calculation (based on duration × severity) → approval → customer notification → finance treatment. Automate where possible.
- Host status page outside production failure domain on separate infrastructure so it survives total platform outage.
- Test status-page reliability during game days (S17), including simulated status-page outage and total region loss.
15. Postmortem policy, blameless process, and facilitation (depends on: 3, 4)
Only 11 of 64 action items closed means postmortems are currently a writing exercise. Rebuild around learning and tracking.
- Make postmortems mandatory: all SEV1 and SEV2, all SEV3 with customer impact or repeat pattern, any near-miss the IC flags.
- Set deadlines: draft within 5 business days, blameless review within 10 days, internal publication within 15 days.
- Adopt a single standardized template: impact and duration, timeline (detection through resolution), root cause (not "human error" but system condition that enabled error), contributing factors, what went well, what went badly, action items.
- Train a pool of blameless facilitators; require a trained facilitator (not the IC) for every SEV1 and SEV2 review.
- Prohibit counterfactual and blame language in postmortems; require contributing factors addressing tooling, process, organization, and human factors.
- Limit action items to small number of concrete, verifiable, owned changes with dates.
- Publish all postmortems internally by default, with security review only for genuinely sensitive material (e.g., unpatched vulnerability details).
- For SEV1 incidents affecting regulated customers, produce a variant customer-facing root cause report.
16. Action item tracking, reliability backlog, and completion governance (depends on: 15)
A postmortem without durable action tracking is a complaint. Solve the 11-of-64 problem.
- Create a single reliability backlog in the engineering tracker (Jira, Linear, etc.) with mandatory label, owner, due date, and link to originating incident.
- Define closure criteria: evidence required (merged code change, tested alert, verified drill) not self-reported status.
- Protect capacity: reserve a fixed percentage of each team's sprint (e.g., 10%) for reliability work; track unspent capacity and report to VP Engineering.
- Run a weekly ageing review of open actions; escalate anything overdue by >30 days to team lead and VP Engineering.
- Require IC or postmortem facilitator to sign off on action completion.
- Report completion rate and median action age in monthly incident review (target: >90% closed within 60 days).
- If the same service repeats an incident in the same area, trigger a design review rather than another action item; break the cycle.
17. Metrics, dashboards, and review cadence (depends on: 2, 3)
Measure to prove the system works. Publish dashboards so everyone sees the scoreboard.
- Define outcome metrics: time to detect (by source, target <5 min internally detected); time to mitigate SEV1/SEV2 (target <60 min); customer-detected incidents per month (target <2); SLA credits paid (target <$100K/year by month 12).
- Define process metrics: declaration latency, page acknowledgement rate, IC roster coverage (no single point of failure), first-update timeliness (% within SLA), update-cadence adherence.
- Define health metrics: alert volume and noise ratio per team, off-hours pages per engineer per month, postmortem timeliness, action closure rate and median action age.
- Build live dashboards visible to every engineer (not just managers); auto-populate from incident tool, update daily.
- Institute review cadence: weekly operational review (30 min, prior week incidents), monthly reliability review (trends, top causes, action status), quarterly executive review (CEO's office, SLA cost, systemic changes).
- Baseline every metric against S2 evidence pack; set 90-day and 12-month targets.
- End every review with decisions and owners, not just numbers.
18. Training, certification, and exercise program (depends on: 4, 12, 13, 14, 15)
A process that exists only on a wiki fails on the first real page. Build skills before deployment.
- Build curriculum: how to be on-call, how to declare an incident, how to run incidents as IC, how to communicate, how to write blameless postmortems.
- Create role-specific tracks: (1) All engineers: 30-min video on severity levels, communication expectations, postmortem format, action items (mandatory); (2) On-call responders: 2-hour workshop on playbooks (S12), alert tool (S5), escalation paths (S9), when to call manager, case studies, mobile app walkthrough; (3) Incident commanders: 4-hour IC bootcamp on leadership under pressure, decision-making, communicating with executives, status-page discipline, postmortem facilitation, practiced drills; (4) Communications leads: 2-hour training on templates, update timings, how to talk to customers, regulatory rules.
- Require certification before joining IC on-call roster: written assessment plus live simulated incident (pass/fail).
- Record videos so async teams can learn on their schedule; create runbooks and quick-reference cards (print + digital); pair new on-call engineers with experienced responder for first week.
- Run monthly tabletop exercises on realistic scenarios from the prior 12 months: region loss, ledger corruption, cascading failures.
- Run quarterly game days with intentional failure injection (database failover, status-page outage, alert tool downtime); include all on-call roles.
19. Pilot with volunteer teams (depends on: 5, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 18)
Do not roll out untested to 28 teams. Run the entire process end to end with a small cohort first.
- Recruit 3–4 volunteer teams covering a mix of criticality: one payment-path team, one ledger-adjacent team, one platform team, one low-traffic team.
- Run complete process in pilot: new severity scale (S3), consolidated tooling (S5, S6), roles (S4), escalation (S9), communications (S13, S14), postmortems (S15), action tracking (S16), paid on-call (S11), training (S18), metrics (S17).
- Instrument pilot against S17 metrics; compare results with S2 baseline.
- Hold weekly retrospectives with pilot teams; iterate on written policies, tooling, training based on feedback.
- Fix top issues found before wider rollout; document what changed and why.
- Produce pilot report with before/after numbers (MTTD, MTTR, alert noise, action completion rate) to carry into rollout conversations.
- Set explicit pilot exit criteria: rotation coverage achieved, zero unacknowledged pages over 2 weeks, all postmortems delivered on time, >80% of action items tracked.
20. Phased rollout to all 28 teams (depends on: 16, 17, 19)
Rollout is a staged migration with readiness gates, not an email announcement.
- Sequence the 28 teams into four waves of ~7 teams each, ordered by customer-impact criticality and readiness; space waves three weeks apart.
- Define per-team readiness checklist: services mapped and owned, alerts cleaned to standard (S8), runbooks written, rotation staffed, training complete, manager briefed.
- Hold gate review with program owner before each team joins; move unready teams to next wave with dated remediation plan.
- Assign named champion per wave; run internal communications cadence explaining why using pilot numbers from S19.
- Handle resistance directly: publish the "own-your-code, own-your-pager" rule and paid on-call mechanics before each wave, not after.
- Retire legacy tools, informal escalation lists, and ad-hoc status-page process at end of each wave on published cutover date.
- Harvest feedback formally at each wave and push accepted process changes through change control.
21. SOC 2 control mapping and evidence framework (depends on: 1, 3, 13, 15)
The audit will test incident response. Mapping controls now means evidence accumulates from the first real incident under the new process.
- Map the process to Trust Services Criteria for incident identification, response, evaluation, and communication of security incidents.
- Write control statements in auditor language; name a single owner per control.
- Specify the evidence artifact for each control: incident record, severity classification, paging log, communications log, postmortem, action tracker entry, training record.
- Define evidence retention and storage location (not on laptops, not on chat history that expires); plan for audit access.
- Identify which controls are blocked until certain rollout waves complete; keep a gap register with owners and review fortnightly with steering group.
- Run early walkthrough with an experienced compliance partner or pre-audit readiness team to test control design before the audit window.
22. SOC 2 dry run, gap remediation, and audit support (depends on: 20, 21)
Convert a good process into a provable one, a few months before auditors arrive. Prove the system works at scale.
- Schedule a dry run six weeks before audit window, sampling real incidents from completed waves against each control's evidence requirements.
- Remediate every gap found; prioritize anything risking a qualified opinion.
- Brief every engineer who may be interviewed so they can describe the process as practiced.
- Prepare auditor package: process documentation, sample incident records, training records, on-call schedules, action tracking register, alert quality metrics.
- Designate a single audit liaison and small evidence-request team so requests do not scatter across teams.
- Rehearse IC and Communications Lead roles under interview conditions; auditors probe realism under pressure.
- Ensure all postmortems, incidents, and evidence are retained, searchable, and accessible to auditors for the required audit period.
23. Standing governance, process ownership, and continuous improvement (depends on: 20, 22)
The classic post-audit failure is the process freezing and decaying. Lock in continuous improvement as a permanent structure.
- Establish a standing Incident Management Council chaired by the program owner, meeting monthly with engineering, support, compliance, and product representation.
- Give program owner documented mandate to change standards; require formal change control for any change to severity, roles, communications timings, or compensation.
- Re-validate severity taxonomy quarterly against real declarations; re-baseline metrics annually.
- Feed incident themes into architecture review and hiring so the program improves the system, not just the response.
- Report quarterly to executive team on metric set (MTTD, MTTR, SLA credits, customer-detected %) and top five systemic causes of incidents.
- Fund a resilience roadmap separate from incident response: reducing shared-database blast radius, multi-region failover, deploy safety.
- Refresh training and tabletop program annually and after any SEV1.
- Keep public backlog of improvement ideas (Jira, shared doc) so any engineer can suggest; publish monthly "incident newsletter" to all engineers with wins and learnings.
--- PROPOSAL 2 (agent deepseek-flash_refine_2, deepseek/deepseek-flash) ---
Estimated complexity: high
Success metrics: - Median time to detect falls from 22 minutes to under 8 minutes by month 6, and under 5 minutes for money-moving journeys by month 9.
- Share of customer-impacting incidents first detected by customers falls from 40% to under 10% by month 9.
- Median time to mitigate for SEV-1 falls from 3 h 10 min to under 60 minutes by month 9.
- 100% of SEV-1 and SEV-2 incidents have a named Triage Owner within 5 minutes and a named Incident Commander within 15 minutes, from month 4.
- Zero incidents in which command authority is unclear for more than 15 minutes, measured monthly from month 4.
- Monthly page volume falls from 3,400 to under 600, with a false-positive rate under 15%, by month 5.
- No service exceeds two pages per on-call shift for three consecutive months, measured by month 6.
- All 28 teams have a documented rotation, a service ownership map and at least one trained on-call engineer by month 6.
- Incident Commander roster holds at least 12 certified ICs covering 24x7 with no uncovered week, from month 4.
- 100% of on-call shifts are paid under a published policy from month 2.
- 100% of SEV-0, SEV-1 and SEV-2 postmortems are published internally within 15 business days, from month 5.
- Postmortem action items closed within 60 days rise from 17% to over 90%, with median age under 30 days, by month 6.
- At least one incident review or near-miss report is filed per team per quarter, from month 6.
- Status page first update is published within 30 minutes on at least 95% of SEV-1 incidents, from month 4.
- Zero missed regulatory notification windows on any incident requiring notification.
- The customer-impact ledger is complete for 100% of customer-impacting incidents from month 5, with SLA credits reconciled monthly.
- Annual SLA credits paid fall at least 50% from the $1.3M baseline within 12 months.
- Customer-impacting incidents fall at least 40% year over year from the 31-incident baseline.
- On-call satisfaction is at or above 7 out of 10, with zero voluntary attrition attributed to on-call, measured quarterly from month 6.
- Review cadence is sustained: weekly operational review in at least 90% of weeks, 12 of 12 monthly reliability reviews, 4 of 4 quarterly executive reviews.
- SOC 2 Type II is passed with zero findings related to incident response.
- 100% of new engineers complete the incident-response onboarding module within 30 days of joining.
- 100% of the 180 services have a named owner, a detection contract and at least one symptom-based alert held to the paging contract by month 7.
Steps (21):
1. Charter, mandate and the evidence clock
This step turns the CEO's email into a funded programme with one accountable owner and explicit authority, and it starts the SOC 2 clock on day one.
- Appoint a single accountable process owner — a Director of Incident Management reporting to the CTO, with a dotted line to the COO for customer and SLA matters.
- Publish a one-page charter: scope (every customer-impacting, money-moving, security and data-integrity event across all 28 teams and both regions), decision rights during an active incident, and the power to freeze deploys and override team preferences.
- Fix the funding envelope up front: tooling licences, training and drill time, and on-call compensation, with an indicative annual figure and the expected return in avoided SLA credits.
- **Start the evidence clock immediately.** A SOC 2 Type II report covers a period of operating effectiveness, so evidence cannot be backfilled; an audit in eight months means the operating period begins in the first weeks of this programme.
- Stand up a steering group of CTO, VP Engineering, Head of Support, Head of Compliance and two engineering managers, meeting fortnightly.
- Make incident-process participation a documented performance expectation for every engineering manager, not an optional extra.
- Agree the timeline explicitly: a working process in four months, evidence accumulating from month two, audit-ready by month seven.
2. Baseline evidence and problem statement (depends on: 1)
You cannot demonstrate improvement without a defensible starting point, and you cannot win the argument about noise without numbers.
- Build a 12-month incident register and re-classify all 31 customer-impacting incidents: date, severity, class, detection source, time to detect, time to mitigate, customers affected, services involved and SLA credits paid.
- Run an alert census per tool, per team and per service: total volume, page-to-action ratio, off-hours interruptions per engineer, and the 50 noisiest rules with a named owner.
- Build a silent-failure register: incidents in which no internal alert fired at all. This is the number that explains the 40% customer-detected rate.
- Survey every on-call engineer and manager on burden, fairness, escalation clarity and pay expectations, targeting above 70% response.
- Interview Support, Account Management and Sales about how customers learn of incidents, what they complain about, and what they are promised.
- Reconstruct the two command-ambiguity incidents minute by minute, from alert to mitigation, to find exactly where ownership lapsed.
- Publish the pack internally as the problem statement and retain it as management-review evidence for the audit.
3. Control mapping and evidence architecture (depends on: 1)
Most programmes leave compliance to the end. This one maps controls in the first month, because the mapping determines what the process must capture from day one.
- Map the process to the relevant Trust Services Criteria for incident identification, evaluation, response, containment, recovery and communication of events (CC7.1 to CC7.5, CC2.2, CC2.3, CC4.1, CC3.x).
- Write each control as a plain-language statement with one named owner and its evidence artifact: incident record, severity classification, paging log, communications log, postmortem, action tracker entry, training record.
- Define the **golden incident file**: one auto-assembled, single-click export per incident containing timeline, roles, severity, communications, postmortem, actions and closure evidence.
- Set retention, storage location and immutability so no control depends on a laptop, a private Slack channel or chat history that expires.
- Keep a gap register with owners and dates, reviewed fortnightly by the steering group.
- Run an early design walkthrough with the auditor's readiness team inside the first 90 days, to test the design before building on it.
4. Severity times class taxonomy (depends on: 2)
Severity alone is not enough. Severity decides how loud the response is; class decides which playbook runs and what must never be traded away.
- Define severity by impact in customer and money terms: customers affected, dollars at risk, regions lost, ledger integrity, settlement delay, data loss and SLA exposure.
- Use SEV-0 for security, privacy and regulatory events; SEV-1 for total or material loss of a payment path; SEV-2 for degradation or single-region loss; SEV-3 for limited impact with a workaround; SEV-4 for internal-only issues and near-misses.
- Define response classes: Availability, Performance, Data integrity and Ledger, Security and Privacy, Third-party dependency, Settlement and Financial, and Process failure.
- **Class can raise a response, never lower it.** A SEV-2 data-integrity incident gets SEV-1 posture, because integrity failures are not recoverable by moving faster.
- Specify automatic triggers: loss of one AWS region, ledger write failures, replication lag above threshold, a missed settlement cut-off, payment success rate below threshold for five minutes.
- State who may declare: any engineer, Support agent or account manager. State who may downgrade: the Incident Commander alone.
- Map every level to its SLA credit exposure and to the customer-visible status page state.
- Include worked examples from the last 12 months so each of the 28 teams recognises its own incidents, and re-validate the taxonomy quarterly against real declarations.
5. Roles, command and the never-without-an-owner rule (depends on: 4)
The two incidents where nobody was in charge for over an hour did not fail at declaration. They failed in the gap before it, when an alert had fired and no one owned it.
- Create one-page role cards: Triage Owner, Incident Commander, Deputy IC, Operations Lead, Internal Communications Lead, Customer Communications Lead, Scribe, Subject-Matter Responders and Executive Sponsor for SEV-1 only.
- Introduce the **Triage Owner** rule: from the moment a page is acknowledged, that person owns the incident until an IC takes over or the incident is stood down. There is never an unowned minute between first page and close.
- State the core rule plainly: the IC owns the incident, not the fix, and does not debug. An IC who starts troubleshooting has abandoned command.
- Give the IC explicit decision rights: declare and escalate severity, freeze changes, halt deploys, approve customer messaging, and pull in any engineer in the company.
- Define minimum viable staffing per severity: SEV-1 staffs every role; SEV-2 staffs IC, scribe, comms and responders; SEV-3 staffs an IC and a scribe only.
- Set handover discipline: four-hour maximum IC shifts on SEV-1, a written handover template, and a deputy named within 15 minutes of declaration.
- Set behaviour standards for responders: one incident channel, one bridge, no side channels, and asks phrased with a named owner and a time.
- Publish the role cards on the internal wiki and link them from every paging notification.
6. Declaration, lifecycle and escalation policy (depends on: 5)
This step defines the mechanical path from an alert to a declared incident and back to normal service, and it removes judgment calls from the worst moments.
- Define lifecycle states with entry and exit criteria: Detected, Triaged, Declared, Mitigated, Resolved, Postmortem, Closed — plus a Watch state with a hard 30-minute timer, after which the incident is either declared or stood down.
- **Make declaring free.** A declaration that turns out to be a false alarm is closed as a false declaration, with no blame and no follow-up, and it is tracked as a metric so the cost of caution stays visible.
- Apply the ambiguity rule: if two responders disagree about whether this is an incident, it is an incident.
- Set acknowledgement and declaration targets: page acknowledged within five minutes, triage decision within fifteen, declaration within thirty for anything customer-facing.
- Build automatic escalation ladders with timers at every layer: responder, service owner, team manager, IC on-call, VP Engineering. Escalation never requires a human decision and is never criticised.
- Define the unresponsive-team path: fifteen minutes without acknowledgement escalates to the team manager, thirty minutes to the IC, who may then direct any available engineer.
- Define change freeze and rollback authority during SEV-1 and SEV-2, and the single named condition that lifts the freeze.
- Enforce one incident, one record, with the timeline auto-captured from the channel and bridge rather than written from memory afterwards.
- Add the two-simultaneous-incidents rule: a second IC is designated, and no IC ever runs two incidents at once.
7. Detection strategy: journeys, synthetic signals and customer-report intake (depends on: 4)
Customers detected 40% of incidents first. That number is the reason this step exists, and it is fixed by measuring customer journeys rather than infrastructure.
- Define SLIs and SLOs for the top 20 customer journeys — payment initiation, authorisation, settlement, ledger read and write, API availability, webhook delivery and payout — measured per region.
- Require symptom-based alerting on those SLOs instead of cause-based alerting on infrastructure metrics.
- Add synthetic transaction monitoring from outside AWS in both regions plus a third location, on a one-minute cadence for money-moving paths.
- Add ledger and shared-PostgreSQL specific signals: replication lag, connection saturation, write latency, lock waits, transaction ID exhaustion, checkpoint pressure and bloat.
- Open a customer-report intake: Support and account managers can raise an incident directly, every customer report creates an incident record, and the customer-report path is counted as a detection source in all reporting.
- Apply the detection-gap rule: whenever a customer reports an incident before internal monitoring, a detection-gap ticket is opened automatically and owned by the responsible team.
- Publish a detection contract per service: every one of the 180 services needs an owner, at least one symptom alert, and a documented expected detect time.
- Measure current detect time per journey, set targets, and run a detection drill per team: break something in staging and see whether it pages before a human notices.
8. Three on-call rotations across 28 teams (depends on: 5)
The objection is that engineers will not carry a pager for another team's code. The answer is not to argue with it, but to build three rotations so the objection becomes structurally impossible.
- Run a **Service On-Call** rotation per team, covering only that team's own services. No engineer is ever paged for code their team does not own.
- Run a **Platform Duty** rotation for genuinely shared infrastructure: the shared PostgreSQL cluster, Kubernetes, networking, CI/CD and observability. This is nobody's product code, so it gets its own paid rotation, staffed from platform teams plus volunteers from other teams.
- Run a central **Incident Commander** roster of 12 to 16 certified senior engineers drawn from across all 28 teams, covering 24x7 on one-week shifts with a primary and a secondary.
- State the consequence honestly: 16 of 28 teams have no on-call today. Each must either build a rotation or formally transfer ownership of its services to a team that will, with the transfer dated and recorded.
- Merge small or low-traffic teams into shared rotations so no rotation has fewer than six engineers, the level below which coverage stops being sustainable.
- Cap load in the scheduling tool: no engineer is on-call more than two weeks per quarter, enforced by configuration rather than negotiation.
- Publish a coverage matrix for all 28 teams showing services, owners, rotation size and gaps, reviewed monthly.
- Write on-call participation into engineering job levels and hiring criteria, so the commitment is set before someone joins rather than after.
9. Compensation, rest and the price of opting out (depends on: 8)
Unpaid on-call is the most cited reason for resistance. The policy must be settled before rollout, not negotiated during it.
- Move immediately to paid on-call: a per-shift stipend benchmarked to the New York market, agreed with HR and Finance, and published with an effective date before any team is asked to join a new rotation.
- Pay event-based compensation for out-of-hours callouts, with a minimum call-out block and a 1.5x rate for time actually spent mitigating.
- Provide documented compensatory rest: no engineer works a normal day after a night incident, and the rest day is policy rather than a favour granted by a manager.
- Cap intrusion with a maximum number of off-hours pages per shift, and trigger a mandatory review whenever the cap is exceeded.
- Allow opt-out but **put a price on it**: an engineer may step out of a rotation, and their team must buy coverage from the paid pool at a published internal rate. This turns a cultural argument into a visible budget decision.
- Publish an explicit amnesty: incident records, near-miss reports and false declarations are never used in performance reviews. Only failure to report is a performance issue.
- Protect against gaming: pay attaches to shifts and callouts, never to incident counts.
- Review the policy every six months against real page volumes, attrition and survey results.
10. Paging contract and alert quality (depends on: 2, 7)
3,400 alerts a month at 85% noise is why engineers resent the pager. Fixing that is the price of admission for everything else in this plan.
- Publish the paging contract: every page must be symptom-based, actionable, owned, mapped to a severity and a class, and linked to a runbook. **No runbook, no page**, enforced by a CI check on the alert definition.
- Separate paging from ticketing: only customer-impacting or imminent-impact signals may wake a human; everything else becomes a ticket or a dashboard entry.
- Set a page budget per team and per service: a maximum number of pages per on-call shift. Breaching it auto-opens a remediation ticket with the engineering manager as owner.
- Put new alerts on two-week probation: a new rule runs as a ticket only and becomes a pager only after it has proved actionable, so teams stop being woken by untested rules.
- Define the default action for a noisy alert: fix, tune or delete within ten working days, with deletion treated as a legitimate and celebrated outcome and the deleted-alert count published.
- Correlate and deduplicate at the ingest pipeline so one root cause produces one page instead of forty.
- Require an expiry date on every silencing rule and temporary threshold, so suppression cannot quietly become permanent blindness.
- Run a 90-day noise sprint with a public burn-down of the top 100 noisiest rules, each owned by a named manager, and report page-to-action ratio per team monthly.
11. Incident tooling consolidation (depends on: 5, 10)
Six alerting tools and no single incident record is a structural cause of the 22-minute detection and the three-hour mitigation.
- Select one incident platform for paging, schedules, escalation policies, incident records, postmortem workflow and status-page integration, and choose it within four weeks.
- Consolidate the six alert sources into a single event pipeline feeding that platform, with deduplication, correlation and severity and class mapping applied at ingest.
- Integrate observability, runbooks, chat and voice bridges so responders see dashboards and runbooks inside the incident record and the timeline is captured automatically.
- Implement the golden incident file export defined in S3, so audit evidence is a single click rather than a reconstruction.
- Run a dual-run period alongside the legacy tools with a defined rollback, then switch the legacy tools off on a published date.
- Host the status page outside the production failure domain so it survives a total platform outage, and test that during a game day.
- Make the platform usable from a phone: an Incident Commander must be able to run a SEV-1 from a mobile device at 3am.
12. Communications: internal, customer and regulator (depends on: 4, 5)
Today the status page is written by whoever is around. This step replaces improvisation with a clock, a named owner and a pre-cleared template.
- Set the internal cadence: first update within 15 minutes of declaration, then every 30 minutes for SEV-1 and hourly for SEV-2, whether or not there is progress.
- Never let the status page be how an employee learns of an incident: internal communication leads, external follows.
- Set the customer cadence: status page updated within 30 minutes of a SEV-1 and 60 minutes of a SEV-2, with a no-new-information update still mandatory.
- Define the channel hierarchy: status page for all 2,100 customers, direct email to affected customers on SEV-1, and named account-manager calls for the top 50 accounts.
- Prepare templates per severity and class in advance, pre-approved by Legal and Compliance, each with the next-update time built in.
- Forbid speculation: customer communications never guess at cause, never assign blame, and never commit to a root cause before the postmortem.
- Define the closing communication: resolution notice, how SLA credits will be handled, and the committed date for the written report.
- Build a **regulator clock matrix**: for each event type, which regulator, which window, who signs off, and the shortest clock that drives the first action. Cover NYDFS Part 500, money-transmitter and banking notification, security breach notification, card-network rules, and public-company disclosure where applicable.
- Route every regulatory notification through Compliance, never Engineering, and pre-clear the templates.
- Publish a customer-facing root-cause report for SEV-1 incidents, especially for regulated and top-tier accounts.
- Assign a named Customer Communications Lead plus a trained deputy on every SEV-1.
13. Customer-impact ledger and SLA credit automation (depends on: 12)
The $1.3M in credits is a symptom of having no single record of customer impact. This step creates one, and makes it do four jobs at once.
- Maintain one durable customer-impact ledger per incident: which customers, which journeys, from when to when, and the estimated credit.
- Use that single record for customer communications, SLA credit computation, regulatory reporting and the annual review, so nobody reconciles two versions of the same outage.
- Automate the credit path end to end — computation, approval, customer notification and finance treatment — so credits stop being a manual scramble three weeks later.
- Reconcile accrued against paid credits monthly and report the result in the executive review.
- Track credit avoidance against programme cost, so the funding case is a number rather than an argument.
14. Postmortem policy with three artifact levels (depends on: 5)
Postmortems currently happen for some incidents, in various formats. This step makes them mandatory where they matter and light where they do not.
- Make postmortems mandatory for every SEV-0, SEV-1 and SEV-2, every SEV-3 with customer impact or a repeat pattern, every near-miss the IC flags, and every incident in which the process itself failed.
- Use three levels so the ritual matches the weight: a lightweight async incident review for SEV-3 and SEV-4, a standard facilitated postmortem for SEV-2, and a full review with an executive sponsor for SEV-1.
- Set deadlines: draft within five business days, blameless review within ten, internal publication within fifteen.
- Adopt one template: impact, timeline, detection, response, contributing factors across tooling, process, organisation and human factors, what went well, what went badly, and actions.
- Train a pool of blameless facilitators and require a trained facilitator for every SEV-1 review.
- Prohibit counterfactual and blame language in the template, and specifically ban the phrase human error as a root cause — the question is always what made the error possible.
- Publish all postmortems internally by default, with security review only for genuinely sensitive material, and produce a customer-facing root-cause variant for SEV-1.
- Treat postmortems as the learning product of the process, not as a compliance artifact.
15. Action items: capped, verifiable, with a repeat-incident rule (depends on: 14)
Eleven of 64 action items closed is not a tracking problem. It is a generation problem: the process produces more actions than the organisation can absorb.
- **Cap each postmortem at three action items.** Anything beyond three goes into a ranked reliability backlog, not into the postmortem.
- Require each action to have a named human owner, a due date, and a definition of done that is an artifact: a merged change, a new alert that fired in a drill, or a test that fails without the fix.
- Prohibit self-reported closure. Closure requires the artifact, signed off by the process owner or the Incident Commander.
- Keep one reliability backlog in the engineering tracker with a mandatory label and a link back to the originating incident.
- Reserve a fixed percentage of every team's sprint capacity for reliability work, with unspent capacity visible to the vice presidents.
- Apply the repeat-incident rule: a second incident with the same contributing factor triggers a design review owned by an architect, not another action item.
- Run a weekly ageing review of open actions and escalate anything more than 30 days overdue to the VP of Engineering.
- Target more than 90% of actions closed within 60 days and a median age under 30 days, reported monthly by team.
16. Training, certification and drills (depends on: 5, 6, 12, 14)
A process that exists only on a wiki page fails on the first real page. Skills have to be built and tested before they are needed.
- Build a short practical curriculum: how to be on call, how to triage and declare, how to run an incident as IC, how to communicate, and how to write a postmortem.
- Require certification before joining the IC roster: a written assessment plus a live simulated incident.
- Certify at least two ICs per team group, so the central roster has depth across all 28 teams and no holiday week is uncovered.
- Run monthly tabletops drawn from the last 12 months of real incidents, including region loss, ledger write failure and a missed settlement cut-off.
- Run quarterly game days with deliberately injected failure, including shared-PostgreSQL failover, status-page outage and alerting-pipeline outage.
- **Drill the process's own failure modes**, not just technical ones: IC unreachable, comms lead on PTO, two simultaneous SEV-1s, a paging storm, and a false alarm that burns an hour.
- Audit the incident process for single points of failure: who is the only person who can do each critical task, and what happens in their holiday week.
- Add a mandatory onboarding module for every engineer joining or transferring in, with completion records kept in an audit-ready form.
- Track training completion by team and publish it in the monthly reliability review.
17. Metrics and review cadence (depends on: 2, 4, 15)
Establish what good looks like, and measure it in a way that makes people report more incidents rather than fewer.
- Define outcome metrics: time to detect by source, time to mitigate by severity and class, percentage of incidents detected by customers, incidents by severity and class, SLA credits paid, and customer-impact minutes.
- Define process metrics, because the process itself has SLOs: IC assigned within five minutes, page acknowledgement rate, first-update timeliness, update-cadence adherence, postmortem timeliness, action closure rate, action median age and IC roster coverage.
- Define health metrics: alert volume and noise ratio per team, off-hours pages per engineer, on-call satisfaction, training completion, and percentage of services with a detection contract.
- **Never publish incident count as a team metric.** It rewards hiding incidents. Publish reporting metrics instead — near-misses filed, detection gaps closed, false declarations made — alongside the outcome metrics.
- Publish live dashboards visible to every engineer, refreshed daily, with each metric baselined against the S2 evidence pack and given a 90-day and a 12-month target.
- Run a 30-minute weekly operational review incident by incident, a monthly reliability review of trends and actions, and a quarterly executive review with the CEO's office covering customer impact, credits and the top five systemic causes.
- Hold a quarterly review of the process itself: what in the process wasted time, what confused responders, and what should be deleted.
- Require every review to end with decisions and named owners, never with numbers alone.
18. Pilot with three to four teams, using real incidents (depends on: 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17)
Do not roll out to 28 teams untested, and do not rely on drills alone when real incidents are available and better.
- Recruit three to four volunteer teams spanning criticality: a payment-path team, a ledger-adjacent team, the platform team that owns the shared PostgreSQL cluster, and a low-traffic team.
- Run the entire process end to end in the pilot: the new severity and class scale, roles, escalation, communications, postmortems, action tracking and paid on-call.
- Treat real incidents during the pilot as the primary training material. Hold a retro within 48 hours of each pilot incident, run by the process owner, while the friction is fresh.
- Instrument the pilot against the S2 baseline and publish before-and-after numbers.
- Hold weekly retrospectives with the pilot teams and change the written policies, the tooling and the training in response, documenting what changed and why.
- Set explicit exit criteria: rotation coverage achieved, zero unacknowledged pages over a defined period, postmortems delivered on time, and actions tracked to closure.
- Produce a pilot report that every later rollout conversation starts from.
19. Phased rollout sequenced by cost of failure (depends on: 15, 18)
Rollout is a staged migration with readiness gates, not an email announcement. The sequencing matters more than the schedule.
- Sequence the 28 teams into four waves of roughly seven, ordered by **incident density and customer-journey ownership**, not by size or alphabet: the teams whose incidents cost the most SLA credits go first, because their improvement is the visible proof.
- Leave three weeks between waves and define a per-team readiness checklist: services mapped and owned, alerts cleaned to the paging contract, runbooks written, rotation staffed to at least six engineers, training complete, manager briefed, and compensation in effect.
- Hold a gate review with the process owner before each team joins, and move an unready team to the next wave with a dated remediation plan rather than granting an exception.
- Give each wave a named champion and run an internal communications cadence that explains the why using the pilot's numbers.
- Handle resistance directly and early: publish the own-your-code-own-your-pager rule and the paid on-call mechanics before each wave, never after.
- Hold a retro within 48 hours of each wave's first incident under the new process, and feed accepted changes back through change control.
- Retire legacy tools, escalation lists and the informal status-page process at the end of each wave on a published cutover date, so two processes never run in parallel for long.
- Sequence the waves so no team faces both a rollout and the audit dry run in the same week.
20. Audit dry run and evidence review (depends on: 3, 19)
This step converts a good process into a provable one, a few months before the auditors arrive.
- Run an internal dry run six weeks before the audit window, sampling real incidents from the pilot and the early waves against each control's evidence requirements.
- Remediate every gap found, prioritising anything that risks a qualified opinion, and re-test the remediated control against the same sample.
- Brief every engineer who may be interviewed so they can describe the process as they actually practise it, not as it is written.
- Prepare the auditor package: process documentation, sample incident records, paging logs, communication logs, training records, on-call schedules and the action tracking register.
- Designate one audit liaison and a small evidence-request team, so requests do not land on 28 teams simultaneously.
- Rehearse the walkthrough with an incident commander and a communications lead, because auditors probe realism under pressure, not documents.
- Keep the audit liaison and the process owner as the same person where possible, so accountability for the control is also accountability for the evidence.
21. Standing governance, resilience roadmap and continuous improvement (depends on: 19, 20)
The classic post-audit failure is that the process freezes and then decays. This step keeps it alive and separates incident response from the engineering that prevents incidents.
- Establish a standing Incident Management Council chaired by the process owner, meeting monthly with engineering, support, compliance and product representation.
- Give the process owner a documented mandate to change standards, and require formal change control for any change to severity, classes, roles, communication timings or compensation.
- Re-validate the severity and class taxonomy quarterly against real declarations, and re-baseline the whole metric set annually.
- Fund a **resilience roadmap separate from incident response**: reduce the shared PostgreSQL blast radius, improve multi-region failover readiness, and tighten deploy safety. Better incident management cannot save a single shared ledger during a corruption event.
- Report quarterly to the executive team on the metric set and the top five systemic causes of incidents, and feed those themes into architecture review, release engineering and hiring.
- Keep a public backlog of process improvement ideas so any engineer can propose one, and change the process from it every quarter.
- Share wins publicly each quarter — faster detection, fewer customer-detected incidents, credits avoided — to sustain momentum after the audit.
- Refresh the training and tabletop programme annually and immediately after any SEV-1.
--- PROPOSAL 3 (agent qwen3.8-flash_refine_3, alibaba/qwen3.8-flash) ---
Estimated complexity: high
Success metrics: - Median Time to Detect (MTTD) < 10 minutes.
- Median Time to Mitigate (MTTR) < 45 minutes for SEV-1.
- Customer-detected incidents < 5% of total incidents.
- Monthly alert volume < 400 actionable alerts (90% reduction in noise).
- SLA credit payouts < $100k annually.
- Postmortem action item completion rate > 90% within 30 days.
- 100% of SEV-1 incidents have a designated IC and Scribe.
- On-call engineer satisfaction score > 4.0/5.0.
- Zero critical findings in SOC 2 Type II audit regarding incident response.
Steps (20):
1. Program Charter & Executive Mandate
Secure explicit authority to mandate the incident management program, transforming it from an engineering initiative into a business-critical operational standard.
* Appoint a Director of Incident Management (DIM) reporting to the CTO, with a dotted line to the Head of Compliance for audit readiness.
* Publish a charter defining scope: all customer-impacting incidents, payment path failures, and data integrity issues.
* Establish a steering committee (CTO, VP Eng, Head of Support, Legal) to resolve cross-team conflicts and approve tooling budgets.
* Define the 'Own Your Code, Own Your Pager' principle as a non-negotiable engineering value, directly addressing the 28-team resistance.
2. Baseline Audit & Evidence Framework (depends on: 1)
Create a defensible starting point for metrics and a structured evidence repository for the SOC 2 Type II audit.
* Consolidate the last 12 months of incident data into a single register, flagging the 31 customer-impacting incidents.
* Map current alert sources to specific services to quantify the 3,400 monthly alerts and identify the top 50 noisiest rules.
* Conduct a survey of on-call engineers to measure current sentiment and identify specific pain points in the unpaid rotations.
* Designate a secure, immutable storage location for all future incident artifacts (timelines, logs, postmortems) to satisfy audit requirements.
3. Severity Taxonomy & Trigger Matrix (depends on: 1)
Define unambiguous severity levels tied to specific triggers, communication SLAs, and executive notifications to eliminate classification debates.
* **SEV-1 (Critical):** System-wide outage, data loss risk, or >5% revenue-impacting transactions. Triggers immediate IC, VP Eng, and legal notification.
* **SEV-2 (Major):** Significant degradation or region-wide latency. Triggers IC and Team Lead.
* **SEV-3 (Minor):** Localized issue with workaround. Triggers on-call engineer.
* **SEV-4 (Internal):** Low-priority noise. Triggers ticket only.
* Map each severity to specific SLA credit exposures to align technical response with financial risk.
4. Incident Roles & Command Structure (depends on: 3)
Standardize the Incident Command System (ICS) to prevent the 'ambiguity of ownership' that caused previous hour-long delays.
* **Incident Commander (IC):** Single point of decision authority; does not debug. Required for SEV-1/2.
* **Scribe:** Maintains the real-time timeline log; mandatory for SEV-1.
* **Comms Lead:** Owns status page and internal broadcasts; shields IC from external noise.
* **SMEs:** Technical responders focused solely on diagnosis/mitigation under IC direction.
* Create role cards with clear 'Do' and 'Don't' responsibilities for quick reference during chaos.
5. On-Call Architecture & Compensation Policy (depends on: 4)
Establish a sustainable, paid 24x7 on-call model that respects ownership boundaries and compensates for disruption.
* Adopt a federated model: each of the 28 teams maintains its own SME rotation for owned services.
* Create a central IC rotation staffed by senior engineers across teams to ensure command coverage.
* Implement a paid policy: base stipend per shift, enhanced pay for call-outs, and mandatory recovery time after SEV-1 night shifts.
* Define the 'Unowned Service' rule: services without a dedicated team rotation must be transferred to an owning team or decommissioned.
6. Tooling Consolidation & Integration (depends on: 2, 5)
Replace the six fragmented alerting tools with a single Incident Management Platform to ensure a unified source of truth.
* Select a platform (e.g., PagerDuty + Statuspage + Slack/Teams) capable of handling 180 services.
* Integrate Kubernetes and PostgreSQL monitoring to feed alerts into the central platform.
* Automate the creation of incident channels and timelines upon alert acknowledgment.
* Ensure the status page is decoupled from the production environment to remain available during outages.
7. Alert Quality & Noise Reduction Program (depends on: 6)
Implement strict alert hygiene rules to reduce the 3,400 monthly alerts and restore trust in the pager.
* **Rule:** All paging alerts must be symptom-based (customer impact) and have a linked runbook.
* **Rule:** Implement deduplication and rate-limiting at the ingestion layer.
* **Sprint:** Conduct a 90-day noise sprint to fix, tune, or delete the top 100 noisiest rules.
* **Metric:** Enforce a 'Page-to-Action' ratio target; if an alert does not result in human action, it must be demoted to ticket-only.
8. Detection Strategy & SLOs (depends on: 3, 6)
Shift from infrastructure metrics to Service Level Objectives (SLOs) to detect customer-impacting issues faster than clients can.
* Define SLIs for critical payment paths: Transaction Success Rate, API Latency, Settlement Lag.
* Implement synthetic transaction monitoring from external vantage points in both AWS regions.
* Establish error budgets; breach of error budget triggers automatic review and potential feature freezes.
* Integrate customer-reported signals (support tickets, social media mentions) into the detection workflow as low-priority alerts.
9. Communication Protocols & Templates (depends on: 4, 6)
Define who talks, when, and what they say to stop ad-hoc status updates and protect brand reputation.
* **Status Page:** SEV-1 initial update within 15 mins; subsequent updates every 30 mins. Use legal-approved templates.
* **Internal:** IC broadcasts to #exec-leadership for SEV-1 every hour.
* **Regulatory:** Define triggers for NY financial regulator notifications; assign Legal/Compliance responsibility for filing.
* **Client Success:** Provide Account Managers with a 'Known Issues' playbook and real-time access to status page APIs for high-value accounts.
10. Postmortem Framework (Blameless) (depends on: 9)
Standardize the learning process to ensure consistent quality and adherence to the 'Blameless' culture.
* **Mandatory:** All SEV-1 and SEV-2 incidents require a postmortem within 5 business days.
* **Format:** Fixed template including Impact, Timeline, Root Cause, Contributing Factors, and Action Items.
* **Review:** Facilitated by a neutral party (not the IC) to enforce blameless language and prevent finger-pointing.
* **Publication:** All postmortems published internally on the Wiki with full searchability.
11. Action Item Tracking & Governance (depends on: 10)
Close the loop on incident resolution by enforcing rigorous tracking of corrective actions to fix the 17% completion rate.
* Automatically create Jira tickets for every action item identified in the postmortem.
* **Enforcement:** SEV-1 incident cannot be marked 'Closed' until high-priority actions are completed or deferred with VP approval.
* **Review:** Weekly review of overdue actions in the Engineering Leadership standup.
* **Metric:** Track 'Mean Time to Remediation' for action items as a key health indicator.
12. SOC 2 Control Mapping (depends on: 2, 10)
Proactively map the new incident processes to SOC 2 Trust Services Criteria to ensure audit readiness.
* Map S4 (Roles), S9 (Comms), and S10 (Postmortems) to Security and Availability criteria.
* Define 'Evidence of Operation' for each control (e.g., automated timeline logs, signed-off postmortems).
* Identify gaps between current state and audit requirements; assign remediation tasks to the DIM.
* Establish a quarterly internal compliance review to test control effectiveness before the Type II audit.
13. Training & Certification Curriculum (depends on: 4, 6)
Equip all engineers with the skills to operate within the new framework, reducing anxiety and improving response quality.
* **Universal Training:** 1-hour module on severity levels and tools for all engineers.
* **IC Certification:** Mandatory workshop and simulation for engineers joining the central IC rotation.
* **Runbook Review:** Each team must update and validate runbooks for their top 3 critical alerts.
* **Onboarding:** Include incident response basics in the engineering onboarding checklist.
14. Pilot Implementation (Wave 1) (depends on: 7, 8, 9, 13)
Deploy the new process to a controlled subset of high-traffic teams to validate assumptions before broad rollout.
* Select 3 teams: Payments Core, Ledger/API, and one Infrastructure team.
* Run the full cycle for 6 weeks: Alerts, IC handover, Comms, Postmortems.
* Collect feedback daily and adjust playbooks and tooling configurations based on real-world friction.
* Measure success using the baseline metrics from S2 (MTTD, Noise, Satisfaction).
15. Full Rollout Strategy (Waves 2-4) (depends on: 14)
Gradually migrate the remaining 25 teams using lessons learned from the pilot to manage change resistance.
* **Wave 2 (Month 3):** Deploy to 8 remaining critical customer-facing teams.
* **Wave 3 (Month 4):** Deploy to internal platform and data teams.
* **Wave 4 (Month 5):** Deploy to remaining low-traffic teams and legacy services.
* Maintain a 'Help Desk' for new on-call engineers to troubleshoot process questions during their first shift.
16. Simulations & Game Days (depends on: 15)
Test the resilience of the process and the tools under controlled failure conditions.
* **Tabletop Exercises:** Quarterly scenario walks-throughs for leadership and ICs (e.g., 'What if the primary region fails?').
* **Chaos Engineering:** Inject failures in non-production or canary environments to test alert accuracy and runbook validity.
* **Communication Drills:** Simulate SEV-1 to test the speed of status page updates and internal notification paths.
* Document findings in postmortems and create action items for identified weaknesses.
17. Metrics Dashboard & Executive Review (depends on: 2, 15)
Establish a continuous feedback loop to monitor the health of the incident management system.
* Build a real-time dashboard tracking MTTD, MTTR, Alert Volume, and SLA Credits.
* **Weekly:** Operational review of new incidents and action items with the DIM and Team Leads.
* **Monthly:** Executive review with CTO/VP Eng focusing on trends, noise ratios, and on-call fatigue.
* Publish 'Reliability Reports' to all engineers to celebrate wins and share learnings.
18. Culture & Change Management (depends on: 5, 15)
Address the cultural shift required to move from 'hero mode' to 'process mode' and mitigate pushback.
* Highlight success stories where the new process reduced toil or prevented customer churn.
* Enforce the 'Blameless' mantra by publicly correcting any leadership communication that implies blame.
* Recognize and reward effective ICs and engineers who improve runbooks or alert quality.
* Address 'Pager Fatigue' proactively by monitoring call-out volume and rotating engineers off if thresholds are breached.
19. SOC 2 Dry Run & Evidence Prep (depends on: 12, 17)
Conduct a mock audit six months out to identify gaps in evidence retention or process execution.
* Sample 10 recent incidents and verify all required artifacts (timeline, logs, postmortem, actions) are present and accessible.
* Interview on-call engineers to ensure they can describe the process and their roles without hesitation.
* Remediate any 'Control Failures' identified during the dry run.
* Prepare the 'Audit Readiness' package for the external auditors.
20. Continuous Improvement Loop (depends on: 17, 19)
Institutionalize the evolution of the incident process to prevent stagnation.
* Review the Severity Taxonomy quarterly to ensure it reflects current system behavior and customer expectations.
* Update Tooling integrations annually based on vendor roadmaps and new AWS/K8s features.
* Revise Compensation Policy annually based on market data and internal fairness reviews.
* Set annual 'Reliability Goals' for the engineering organization to drive long-term focus.
Step-level differences computed by the tool:
Proposal 1 vs the previous-round proposal it resembles most (claudeHaiku4.5_refine_1): 15 steps kept, added ['Severity and response class taxonomy', 'Incident lifecycle, Triage Owner rule, and escalation policy', 'Three on-call rotations: Service, Platform, and Incident Commander', 'Status page infrastructure and customer-impact ledger', 'Postmortem policy: mandatory, blameless, three-level framework', 'Phased rollout sequenced by cost of failure', 'SOC 2 dry run and evidence review'], removed ['Severity taxonomy and trigger matrix', 'Alert consolidation and event pipeline', 'On-call architecture and 24x7 coverage model', 'Playbooks and communication templates by severity', 'Status page, customer notifications, and account-manager playbook', 'Postmortem policy, blameless process, and facilitation', 'Phased rollout to all 28 teams', 'SOC 2 dry run, gap remediation, and audit support']
Proposal 2 vs the previous-round proposal it resembles most (deepseek-flash_refine_2): 18 steps kept, added ['Day-one operating rules and the minimum viable process', 'Paging contract and the noise-reduction programme', 'Customer trust workstream and the SLA credit ledger', 'Postmortems: mandatory set, three levels, blameless by design', 'Standing governance and process ownership', 'Resilience track for the shared ledger and multi-region blast radius'], removed ['Paging contract and alert quality', 'Customer-impact ledger and SLA credit automation', 'Postmortem policy with three artifact levels']
Proposal 3 vs the previous-round proposal it resembles most (claudeHaiku4.5_refine_1): 18 steps kept, added ['Severity & Class Taxonomy', 'Incident Command Roles & Structure', 'Detection Strategy & SLO Implementation', 'Communication Workflows (Internal & External)'], removed ['Severity taxonomy and trigger matrix', 'Incident roles, command structure, and decision rights', 'Alert consolidation and event pipeline', 'Detection strategy: SLOs, signals, and customer-journey monitoring', 'Internal and customer communications workflows']
Origin of the steps of the new proposals, matched by title by the tool (evidence for "taken"; ideas can also travel without a matching title):
Proposal 1: 13 of its 22 steps match its own previous version, 3 are new; steps 3, 4, 5, 9, 20, 21 resemble steps 3, 4, 6, 7, 19, 20 of proposal 2
Proposal 2: 16 of its 24 steps match its own previous version, 4 are new; steps 17, 22 resemble steps 17, 23 of proposal 1; steps 4, 11 resemble steps 3, 7 of proposal 3
Proposal 3: 9 of its 22 steps match its own previous version, 0 are new; steps 8, 10, 11, 13, 14, 16, 18, 19, 20 resemble steps 10, 12, 9, 14, 15, 17, 19, 20, 21 of proposal 1; steps 1, 2, 3, 17 resemble steps 1, 2, 4, 16 of proposal 2
PROPOSALS OF ROUND 2 (to assess):
--- PROPOSAL 1 (agent claudeHaiku4.5_refine_1, anthropic/claude-haiku-4-5) ---
Estimated complexity: high
Success metrics: - Median Time to Detect from 22 minutes to <5 minutes for internal detection and <10 minutes for customer journeys by month 6; <2% of customer-impacting incidents first detected by customers (vs. 40% currently) by month 9.
- Median Time to Mitigate for SEV-1 from 3h 10min to <45 minutes by month 9; SEV-2 to <90 minutes by month 9.
- Monthly alert volume from 3,400 to <400 with >95% signal-to-noise ratio by month 5; all alerts satisfy paging contract (symptom-based, owned, runbook-linked, severity-mapped).
- Zero incidents with command ambiguity lasting >15 minutes from month 2 onwards; 100% of SEV-1 and SEV-2 incidents have named IC logged in timeline within 5 minutes; 100% of Triage Owners assigned within acknowledgement.
- 100% of the 28 teams have a documented Service On-Call rotation or a signed service-ownership transfer by month 4; Platform Duty rotation and IC roster both staffed and live by month 3.
- On-call compensation policy published and effective by month 1; on-call satisfaction score ≥7/10 by month 6; zero on-call-attributed voluntary attrition by month 6.
- SLA credits paid from $1.3M annually to <$100K by month 12; credit avoidance (prevented credits) tracked and reported monthly.
- 100% of mandatory postmortems (SEV-0, SEV-1, SEV-2, and repeat incidents) published internally within 15 business days by month 4.
- Postmortem action item completion rate from 17% (11 of 64) to >90% within 60 days by month 6; median action age <30 days; zero repeat incidents caused by the same contributing factor without a design review.
- 100% of the 180 services have a named owner, a detection contract, and at least one symptom-based alert by month 6.
- Status-page first update published within 30 minutes for ≥95% of SEV-1 incidents by month 4.
- 100% of new engineers complete incident-response onboarding within 30 days of joining; IC certification includes written exam and live simulation; ≥2 certified ICs per team group; zero uncovered hours in 24x7 IC roster.
- Weekly operational review held in ≥90% of weeks; 12 of 12 monthly reliability reviews; 4 of 4 quarterly executive reviews; all reviews end with documented decisions and owners.
- SOC 2 Type II audit passes all incident-response controls (CC7.1–7.5, CC2.2–2.3, CC4.1, CC3.x) with zero findings by month 8.
- All pilot and rollout incidents captured with complete golden incident files (timeline, roles, communications, postmortem, actions, closure evidence) by month 3 onwards; audit dry run identifies zero critical gaps by month 7.
- All 28 teams transitioned to new process by month 5; all legacy alert tools decommissioned; single source of truth for incidents established and sustained.
Steps (22):
1. Executive charter, governance structure, and evidence clock
Turn the CEO email into a funded, authorized program with clear ownership and documented evidence collection for SOC 2, starting today.
- Appoint a Director of Incident Management reporting to CTO, with dotted line to COO (customer impact) and Head of Compliance (audit readiness).
- Publish a one-page charter: scope (all customer-impacting, payment-path, data-integrity incidents across 28 teams and 2 regions), decision rights (IC may freeze changes, override team preferences during incidents), and authority to mandate process participation.
- Secure annual budget for tooling, training, on-call compensation ($500–800K estimated), and resilience work. Connect funding to avoided SLA credits ($1.3M baseline).
- Establish standing Incident Management Steering Group: CTO, VP Engineering, Head of Support, Head of Compliance, one engineering manager per region. Meet monthly.
- **Start the SOC 2 evidence clock on day 1.** An audit in eight months means operating-period evidence begins now; design the process to capture evidence continuously, not retroactively.
- Make incident-response participation a documented performance expectation for all engineering managers and team leads.
- Publish timeline: working process in month 2, all 28 teams in month 5, audit-ready in month 7.
2. Baseline measurement, incident register, and evidence pack (depends on: 1)
Establish defensible baseline metrics and identify structural gaps that explain the 40% customer-detected rate and 22-minute detection time.
- Build a 12-month incident register with all 31 customer-impacting incidents: date, detection source, detection time, mitigation time, customer count, services involved, SLA credits paid, root cause class.
- Audit the current alert estate: total volume per tool, volume per team, volume per service, page-to-action ratio, top 50 noisiest rules, off-hours interruptions per engineer.
- Construct a **silent-failure register**: incidents with no internal alert fired at all. This explains the 40% customer-detected rate.
- Reconstruct the two command-ambiguity incidents minute by minute: exactly when did ownership become unclear, how long, what was the decision bottleneck.
- Survey on-call engineers (target >70% response): burden, fairness, pay expectations, escalation clarity, willingness to stay.
- Interview Support and Account Management: how do customers discover incidents, what do they complain about, how do they contact you.
- Publish the problem statement internally; retain all artifacts for SOC 2 audit evidence. This is the baseline against which all improvements are measured.
3. Control mapping and evidence architecture (depends on: 1, 2)
Design the process to generate SOC 2-compliant evidence automatically, from the first real incident, so the audit clock ticks in your favour.
- Map the new process to Trust Services Criteria CC7.1–7.5 (incident identification, response, evaluation, containment, communication), CC2.2–2.3 (authorization), CC4.1 (change management), CC3.x (information availability).
- For each control, write a one-paragraph plain-language statement, name a single owner, and specify the evidence artifact (incident record, timeline, communications log, postmortem, action tracker, training record).
- Define the **golden incident file**: one single-click export per incident containing severity, timeline, roles assigned, decisions made, communications sent, postmortem, and action items. This is the audit unit.
- Specify data retention, immutability, access control, and storage location (not laptops, not chat history that expires). Ensure evidence is searchable and organized by incident date.
- Keep a gap register with owners and dates; review fortnightly in the steering group. Identify which controls are blocked by incomplete rollout and when they unblock.
- Run an early design walkthrough with an experienced SOC 2 readiness partner inside month 1 to stress-test control design before building on it.
4. Severity and response class taxonomy (depends on: 2)
Define four severity levels and four response classes so every decision—paging, communications, postmortem, compensation—keys off a defensible rule, not a judgment call.
- **Severity by impact scope**: SEV-1 (total payment-path loss, data corruption, or >5% transaction failure for >5 min); SEV-2 (significant degradation or single region loss); SEV-3 (limited impact with workaround available); SEV-4 (internal issue or cosmetic); SEV-0 (reserved for security/regulatory/privacy events).
- **Response class** (orthogonal to severity): Availability, Performance, Data Integrity & Ledger, Security & Privacy. **Key rule: class can raise severity, never lower it.** A SEV-3 data-integrity incident gets SEV-1 response posture because integrity is not recoverable by moving faster.
- Automatic triggers: loss of one AWS region → SEV-1 or SEV-2 (class-dependent); ledger write failures → SEV-1; replication lag >10s → escalation review; payment success rate <99% for >5 min → SEV-1/2; missed settlement window → SEV-1; total external API unavailability → SEV-1.
- Who may declare: any engineer, Support, account manager (based on observed customer impact). Who may downgrade: IC only, after investigation.
- Map each level to SLA credit exposure and to customer-facing status-page state.
- Include worked examples from the last 12 months so all 28 teams recognize their own incidents in the taxonomy. Re-validate quarterly against real declarations.
5. Incident lifecycle, Triage Owner rule, and escalation policy (depends on: 4)
Eliminate the "nobody was in charge for over an hour" problem by assigning ownership the moment a page is acknowledged.
- Define lifecycle states with clear entry/exit criteria: Detected (alert fired) → Triaged (is this real and customer-impacting?) → Declared (severity assigned) → Mitigated (core issue resolved) → Resolved (all verifications done) → Postmortem (review scheduled) → Closed (action items tracked or dismissed).
- **Introduce the Triage Owner rule**: the person who acknowledges the page owns the incident until an IC is assigned or the incident is stood down. There is never an unowned gap between first page and declaration. Triage Owner's sole job: decide within 15 minutes whether this requires an IC or a direct stand-down.
- Set aggressive acknowledgement and declaration targets: page acknowledged in 5 min; triage decision (is this real?) in 15 min; severity declaration in 30 min for any customer-facing incident.
- Implement automatic escalation ladders with no human judgment required: if responder does not acknowledge in 5 min, escalate to service owner; if no ack in 10 min, escalate to team manager; if no ack in 15 min, escalate to IC on-call. Escalation is never criticized.
- Define unresponsive-team path: if a service's on-call is unreachable for 30 min, IC may direct any available engineer from any team to engage.
- For SEV-1 and SEV-2: change freeze until IC declares mitigation confirmed; IC unfreezes changes explicitly.
- Enforce one incident, one record. Timeline auto-captured from Slack channel and bridge; never written from memory later.
6. Incident roles, command structure, and decision rights (depends on: 5)
Define clear roles with one-page responsibility cards published and linked from every paging notification.
- **Incident Commander**: owns incident outcome, not the fix. Declares severity, decides escalation, approves all customer communications, freezes changes, calls responders, hands off in shifts. Non-technical ICs are acceptable; technical depth is not required.
- **Deputy IC**: assigned within 15 min of declaration; shadows IC; takes over if IC unavailable or after 4-hour shift on SEV-1. Maximum IC shift: 4 hours on SEV-1, 6 hours on SEV-2.
- **Triage Owner** (new role): owns incident from first page acknowledgement until IC takes over or stand-down decision is made. Required for all incidents.
- **Communications Lead**: owns internal Slack updates and status-page messaging; shields IC from customer contact and interruptions.
- **Scribe**: records real-time timeline with decisions, actions, and key timestamps; not responsible for fixing.
- **Subject-Matter Responders**: engineers with service context; take IC direction; report only to IC; no side channels or parallel debugging.
- **Operations Lead** (SEV-1 only): coordinates multiple responders, manages incident bridge, maintains escalation list.
- Minimum viable staffing: SEV-1 requires all roles; SEV-2 requires IC, Deputy, Comms, Scribe, SMEs; SEV-3 requires Triage Owner and IC.
- Create laminated role cards for every on-call shift location (office, home, printed in pockets).
7. Three on-call rotations: Service, Platform, and Incident Commander (depends on: 5, 6)
Directly address the "carrying a pager for another team's code" objection by making it structurally impossible.
- **Service On-Call rotation** (federated): each of the 28 teams maintains a rotation for their own services only. No engineer is paged for code their team does not own. The answer to "why am I carrying a pager?" is now simply: "for your team's code."
- **Platform Duty rotation** (centrally staffed): shared PostgreSQL cluster, Kubernetes, networking, CI/CD, observability, and incident management tooling. Nobody's product code, so it gets its own dedicated rotation. Staffed from platform teams plus volunteers from other teams; paid at premium rate.
- **Incident Commander roster** (24x7): 12–16 certified senior engineers from across all 28 teams, on one-week primary shifts with secondary backup. Covers every hour with no single point of failure and no uncovered holiday week.
- **Consequences and gates**: 16 of 28 teams have no on-call today. Each must either (a) build a Service On-Call rotation of at least 6 engineers, or (b) formally transfer service ownership to a team that will, with transfer documented and dated. No exceptions, no waivers. Unowned services are decommissioned or transferred by end of month 2.
- Merge small or low-traffic teams into shared rotations where service ownership is unclear (e.g., shared analytics, testing infrastructure).
- Enforce scheduling limits in the tooling: no engineer on-call more than 2 weeks per quarter, automatically enforced by configuration, not negotiation.
- Publish a coverage matrix for all 28 teams showing services, owners, rotation size, gaps, and monthly status.
8. On-call compensation, rest policy, and sustainability (depends on: 7)
Settle compensation before rollout, not during negotiations. Make on-call sustainable and valued.
- **Paid on-call**: effective immediately upon joining a rotation. Weekly stipend while on shift (benchmark to New York market: $600–1,000 per week per engineer), regardless of incident volume.
- **Event-based compensation**: 1.5× hourly rate for time spent mitigating out-of-hours incidents, minimum one-hour block per callout. Tracked by incident record (auto-capture from timeline).
- **Compensatory rest**: no engineer works a normal 8-hour business day after a night incident requiring >2 hours mitigation. Rest day is documented policy, not a favour granted by manager.
- **Intrusion cap**: maximum 3 unscheduled pages per week per engineer. Exceed the cap in a week and trigger an immediate review; exceed in a month and escalate to VP Engineering. Breaches are structural signal that alert quality or service stability has a problem.
- **Voluntary opt-out**: an engineer may exit a rotation; their team must hire or buy replacement coverage from paid pool at published internal rate ($X per shift). This converts culture debate into visible budget decision.
- **Amnesty policy**: incident records, near-miss reports, and false declarations are never used in performance reviews or compensation discussion. Only failure to report is a performance issue.
- **Policy publication**: publish compensation structure and effective date before any team is asked to join a rotation, and include on-call expectations in job descriptions and hiring conversations.
- **Semi-annual review**: reassess compensation and caps every six months against actual page volumes, attrition rates, and survey feedback.
9. Detection strategy: SLOs, synthetic monitoring, and customer-report intake (depends on: 4, 7)
Close the 40% customer-detected gap by monitoring customer journeys instead of infrastructure metrics.
- **SLO-based alerting**: Define SLIs and SLOs for the top 20 customer journeys (payment initiation, authorization, settlement, ledger read/write, API availability, webhook delivery, payout). Measure per region. Alert on SLO breach, not on infrastructure metric (e.g., alert on "payment success rate <99%" not "database CPU >80%").
- **Synthetic transaction monitoring**: deploy synthetic transactions from outside AWS in both regions plus a third geographic location, one-minute cadence, for all money-moving paths. These are your first alarm bell.
- **Ledger-critical signals**: PostgreSQL replication lag (target: <1s, alert >5s), connection saturation, write latency (p95), lock-wait time, transaction ID exhaustion proximity, checkpoint pressure, table bloat. These are separate alerts on shared-database health.
- **Customer-report intake** (new detection channel): Support and Account Managers can raise an incident directly in the platform. Every customer report creates an incident record automatically, and the "customer report" detection source is counted in all metrics. This is a legitimate detection method, not a failure.
- **Detection-gap rule**: whenever a customer reports an incident before internal monitoring fires, auto-create a ticket in the owning service's backlog with root cause: "Monitoring gap on [journey]."
- **Detection contract per service**: every one of the 180 services needs a named owner, at least one symptom-based alert mapped to a SLO, and a documented expected detect time (target: <5 min for payment path, <10 min for others). Published on wiki and reviewed monthly.
- **Detection drills**: run a quarterly drill per team: simulate a broken service in staging and verify it triggers a page before a human notices.
- **Resilience roadmap separation**: detection improvements do not protect against ledger corruption or multi-region failure. Fund a separate resilience roadmap to reduce shared-database blast radius and improve failover safety.
10. Alert quality standards and noise-reduction program (depends on: 9)
Cut the 3,400 monthly alerts (85% noise) to <600 with 95% signal. This is the price of admission for on-call buy-in.
- **Paging contract**: every page must satisfy all of (1) symptom-based (customer impact, not infrastructure cause), (2) actionable (linked runbook with immediate next step), (3) owned (named team responsible), (4) severity-mapped (SEV-1/2/3/4), (5) SLO-linked where applicable. **No runbook, no page.** Enforce with CI check on alert definition.
- **Separation rule**: only customer-impacting or imminent-impact signals trigger pages. Everything else becomes a dashboard entry, a ticket, or a log line. Noisy infrastructure metrics go to dashboards, not pagers.
- **Page budget per service**: no service may exceed 2 pages per on-call shift per month. Exceeding budget auto-opens a remediation ticket in the owning team's backlog (with alert-quality review assigned to tech lead).
- **Automatic suppression rules**: (1) silence alerts if service auto-recovered within 30s, (2) suppress known maintenance windows, (3) group flapping alerts (>5 in 2 min) into one page, (4) rate-limit noisy services (max 1 page per 5 min until condition clears). All suppression rules must have an expiry date; no permanent silence without a ticket.
- **Probation for new alerts**: new alert rules run as tickets only and alert to a Slack channel; after two weeks of proving actionability (every alert resulted in human action), they graduate to pager.
- **Noise sprint**: run a focused 90-day program with a public burn-down of the top 100 noisiest rules. Assign each to a named manager. Default action: fix root cause, tune threshold, or delete within 10 working days. Deletion is a legitimate successful outcome (celebrate it).
- **Correlation and deduplication**: consolidate alert sources at ingest pipeline so one outage triggering 40 alerts produces one page, not 40.
- **Alert ownership**: every alert must have an owning team and a maintenance contact. Update monthly.
11. Incident tooling consolidation and integration (depends on: 5, 10)
Replace six alert tools and ad-hoc incident records with a single source of truth that unifies paging, escalation, timeline, and audit evidence.
- **Tool selection**: choose an incident-management platform (e.g., PagerDuty, Incident.io, Opsgenie) that integrates paging schedules, escalation policies, incident records, postmortem workflow, and status-page APIs. Decision gate: month 1.
- **Event pipeline consolidation**: route all alerts from the six legacy tools into a single event pipeline that feeds the incident platform. Apply deduplication, correlation, severity/class mapping, and rate-limiting at ingest.
- **Observability integration**: connect the incident platform to your Kubernetes dashboards, PostgreSQL monitoring, distributed tracing, and logs so responders see context in one pane. Link runbooks directly into incident records.
- **Slack and bridge integration**: auto-create incident Slack channels, auto-invite roles, auto-capture timeline from channel transcript and voice-bridge recording. Timeline is not written from memory; it is auto-captured.
- **Golden incident file**: implement the export defined in S3. One click produces a complete, immutable, audit-ready PDF: severity, timeline, roles, decisions, communications, postmortem, action items, and closure evidence.
- **Dual-run period**: run both legacy and new platform in parallel for two weeks. Define rollback criteria (e.g.,
12. Escalation automation and incident lifecycle enforcement (depends on: 5, 11)
Eliminate judgment calls from the worst moments. Escalation is automatic, mechanical, and blameless.
- **Automatic escalation ladders**: page responder → if no ack in 5 min, page service owner → if no ack in 10 min, page team manager → if no ack in 15 min, page IC on-call + call them immediately (phone + SMS + Slack). No human decides to escalate; timers fire escalations.
- **Severity-based escalation tempo**: SEV-1 uses faster timers (2 min for IC on-call), SEV-2 uses moderate timers (5–10 min), SEV-3 uses slower timers (15–30 min). Configured in tooling, reviewed quarterly.
- **Dual IC rule**: if a second SEV-1 incident is detected while the first is active, immediately page and assign a separate IC. ICs never run two incidents in parallel.
- **Change freeze and rollback authority**: SEV-1 and SEV-2 trigger automatic deploy freeze. Only the IC (with CTO/VP Eng notification) may unfreeze. Freeze lifts only when IC explicitly declares mitigation confirmed and verifies no new incident symptoms for 5 min.
- **Unresponsive team escalation**: if service's on-call does not acknowledge in 30 min, IC may direct any engineer from any team (volunteers first, then rotated) to engage. This is documented and reported in monthly review (escalation = signal of rotation problem).
- **One incident, one record**: all decisions logged in the incident platform. Auto-capture from Slack, bridge, status-page updates. Timeline is the source of truth; postmortem is written from timeline, never constructed after the fact.
- **Ambiguity rule**: if two responders disagree about whether an incident should be declared, it is declared. False declarations (stand-downs within 30 min of declaration) are tracked as metrics and closed without blame.
- **Watch state**: an unconfirmed incident can live in "Watch" state for max 30 min; after that, either declare it or stand it down explicitly.
13. Internal, customer, and regulatory communications workflows (depends on: 6, 12)
Define who informs whom, in what order, via what channel, with explicit timings and pre-approved templates.
- **Internal cadence**: first update to #incidents Slack channel within 3 min of declaration (even if "Investigating"). Then updates every 5 min (SEV-1), 15 min (SEV-2), or 30 min (SEV-3), or immediately on material change (e.g., mitigation achieved, scope widened). **Comms Lead owns the update; IC must not be interrupted.**
- **Executive notification**: IC calls CTO and VP Eng within 1 min of SEV-1 declaration (not email, not Slack, call). Incident declared in Slack with severity label, IC name, and affected service. Escalation channel lead auto-pinged.
- **Customer communication channels**: status page (all 2,100 customers), direct email to affected customers (top-tier accounts and customers affected by SEV-1), account-manager calls (top 50 accounts on SEV-1).
- **Status page timings**: update within 3 min of SEV-1 declaration, 10 min of SEV-2, 30 min of SEV-3 (even if root cause unknown; use "Investigating" with next-update ETA). Updates every 5–30 min depending on severity. Always include next-update time.
- **Pre-approved templates**: draft customer-facing language for each severity and class in advance with Legal and Compliance. Templates specify impact language ("some of your transactions are delayed" not "our database failed"), workarounds if available, and next-update commitment. Never speculate on cause in customer communication.
- **Regulatory notification path**: identify incidents requiring regulator notification (NYDFS Part 500, money-transmitter rules, payment-card-network rules, securities disclosure). Build a clock matrix: event type → regulator → notification window → signer. Compliance owns all regulatory notifications (never Engineering). Pre-clear templates. Flag incidents to Compliance immediately upon declaration.
- **Account manager playbook**: contact tree for top 50 accounts, templated talking points (facts only, never speculation), escalation path if customer escalates, what to offer (service credit, technical deep-dive call).
- **Closing communication**: resolution notice, SLA credit impact, commitment date for written root-cause report, customer action required (none, or security update, etc.).
14. Status page infrastructure and customer-impact ledger (depends on: 13)
Make the status page reliable, customer-centric, and audit-ready. Track customer impact in a single durable record.
- **Status page decoupling**: host status page outside production failure domain (separate cloud, separate infrastructure, separate database). Integrate incident platform with status page so incident record drives all public updates. Status page survives total platform outage.
- **Component-to-journey mapping**: status page components map to customer journeys ("Payments", "Settlements", "Payouts", "Ledger API") not to internal services. Allow customers to subscribe to components; notify by email or webhook.
- **One-click update templates**: pre-fill status-page template with severity, impact language, next-update time, and estimated resolution. Comms Lead types minimal new info ("Root cause identified" or "Workaround available"), and updates auto-post.
- **Customer-impact ledger** (one record per incident): which customer accounts affected, which journey(s) impacted, exact start and end time of impact, estimated SLA-credit exposure. Use this single record for customer communications, credit computation, regulatory reporting, and annual review. No reconciliation of two versions of the same outage.
- **SLA credit automation**: compute credit based on duration × severity × customer tier → auto-generate customer notification → auto-post to finance system. Reconcile accrued vs. paid credits monthly and report in executive review.
- **Testing during game days**: simulate status-page outage and verify alerts continue to fire; test total region loss and confirm status page remains updated; drill runbook for manually updating status page if platform is down.
15. Postmortem policy: mandatory, blameless, three-level framework (depends on: 6, 13)
Turn postmortems from a writing exercise (11 of 64 action items closed) into the learning engine of the system.
- **Mandatory postmortems**: all SEV-0, SEV-1, and SEV-2 incidents; all SEV-3 with customer impact or repeat pattern; any near-miss IC flags; any incident where the process itself failed (IC unreachable, Comms Lead unavailable, false declaration, missed update SLA).
- **Three-level framework** (proportionate to weight): (1) lightweight async review for SEV-4 and low-impact SEV-3 (10 min template in shared doc, owner + IC review), (2) standard facilitated postmortem for SEV-2 and impactful SEV-3 (full template, facilitated by trained neutral party, published within 10 days), (3) full executive postmortem for every SEV-1 and every security incident (executive sponsor assigned, full investigation, published within 15 days, customer-facing variant prepared).
- **Fixed timeline**: draft postmortem within 5 business days, blameless review within 10 days, internal publication within 15 days.
- **Single template**: impact (who, how many, how long, financial exposure), timeline (detection through resolution), root cause (not "human error" but system condition that enabled error; what was the gap?), contributing factors (tooling, process, organization, knowledge, monitoring), what went well, what went badly, action items (≤3, rest go to reliability backlog).
- **Blameless facilitation**: train a pool of blameless postmortem facilitators (target: 10+ engineers). Require a trained, neutral facilitator for every SEV-1 and SEV-2 review. Prohibit counterfactual language ("if the engineer had"), blame language, and the phrase "human error" as a root cause.
- **Publication rule**: publish all postmortems internally by default; security review only for genuinely sensitive material (e.g., unpatched vulnerability details or customer PII in logs). Create a customer-facing root-cause report for every SEV-1, especially for regulated customers, with legal and compliance sign-off.
- **Searchability**: store postmortems in a searchable wiki or issue tracker with tags (service, class, root cause category) so teams can learn from similar incidents without repeating them.
16. Action item tracking, reliability backlog, and repeat-incident design rule (depends on: 15)
Close the loop on incident learning by enforcing verifiable, tracked action items and breaking cycles of repeat incidents.
- **Action item capping**: each postmortem generates a maximum of 3 action items. Anything beyond 3 goes into a ranked reliability backlog, not into the postmortem, to prevent overwhelming teams.
- **Action item requirements**: each item must have (1) a named human owner (not a team), (2) a due date (≤60 days, target ≤30 days), (3) a definition of done (merged code change, tested alert, audit evidence, architectural decision, new runbook, training completed) not self-reported status.
- **Single reliability backlog**: create one backlog in your engineering tracker (Jira, Linear, etc.) with mandatory label (e.g., `incident-action`), link to originating incident, and link to postmortem. Track progress weekly.
- **Closure sign-off**: Incident Commander or postmortem facilitator must sign off on closure, verifying artifact exists (code merged, alert tested in drill, runbook verified).
- **Repeat-incident rule**: if the same service or component has a second incident with the same contributing factor, **do not create another action item**. Instead, escalate immediately to an architect or tech lead and trigger a design review (not a task, a review). This breaks the cycle of repeated patches; the system needs a structure change.
- **Capacity protection**: reserve a fixed percentage of each team's sprint capacity (10–15%) for reliability work. Track unspent capacity and report to VP Engineering monthly; if a team is not spending it, work with them to identify and fix blockers.
- **Ageing and escalation**: run a weekly review of open actions; escalate anything >30 days overdue to team lead and VP Engineering. Monthly report: completion rate (target >90% within 60 days) and median action age (target <30 days).
17. Training, certification, and exercise program (depends on: 6, 13, 15, 16)
Build skills before deploying the process. Run ongoing drills so the system is tested, not guessed at.
- **Curriculum**: (1) All engineers (30-min async video): severity taxonomy, communication expectations, postmortem format, when to declare an incident, where to find runbooks. (2) On-call responders (2-hr workshop): alert tool walkthrough, playbooks by severity, escalation paths and timers, when to call manager, mobile app walkthrough, case studies from the last 12 months. (3) Incident Commanders (4-hr bootcamp + test): leadership under pressure, decision-making (severity, escalation, rollback), communicating with executives, status-page discipline, postmortem facilitation, handling ambiguity, live simulated incident (pass/fail certification). (4) Communications Leads (2-hr training): templates per severity and class, customer-communication rules (no speculation, no blame), update timings, how to shield IC, regulatory triggers.
- **IC certification**: written assessment (75% pass required) plus live simulated incident (role-play with facilitator, graded on severity declaration, escalation decisions, communication, handover). Certification valid for 12 months; recertify via annual refresher or another live sim.
- **Depth across teams**: certify at least 2 ICs per team or team group so central roster is not siloed in one group; no holiday week is uncovered.
- **Async content**: record all training videos so async teams can learn on their schedule. Create quick-reference cards (laminated, pocket-sized) for roles and playbooks; distribute to on-call locations (office, home).
- **Monthly tabletop exercises**: drawn from real incidents from the last 12 months (region loss, ledger write failure, missed settlement window, cascading failures). Facilitator describes scenario; 3–4 responders play out response (Triage Owner, IC, Comms) as if real. Run 30 min; retro for 15 min afterward.
- **Quarterly game days**: deliberately inject failures into production (database failover, status-page outage, alerting-pipeline outage, dual SEV-1 incidents). All on-call roles engage. Run 2–3 hours; measure response times, decision quality, and communication. Document findings and create action items for identified gaps.
- **Drill the process's own failure modes**: IC unreachable (on-call unavailable, phone broken), Comms Lead on PTO, two simultaneous SEV-1s, paging storm (100+ alerts), false alarm that consumes an hour. Test escalation paths, deputy takeover, and recovery.
- **New-engineer onboarding**: add incident-response module to all engineering onboarding (completion tracked, audit-ready). All engineers must complete within 30 days of joining or transferring in.
18. Metrics, dashboards, and review cadence (depends on: 2, 12, 16, 17)
Measure to prove the system works. Publish live dashboards so every engineer sees the scoreboard and the system is transparent.
- **Outcome metrics**: Median Time to Detect by source (target: <5 min internally detected, <10 min customer journeys); Median Time to Mitigate for SEV-1/2 (target: <60 min SEV-1); customer-detected incidents as % of total (target: <5%); incidents by severity (should be mostly SEV-3/4, few SEV-1); SLA credits paid (target: <$100K/year by month 12); annual credit avoidance vs. program cost.
- **Process metrics**: IC assigned within 5 min (target: >95% of incidents); page acknowledgement rate (target: >98% within 5 min); first-update timeliness (target: >95% within SLA); postmortem timeliness (target: 100% of mandatory postmortems published on time); IC roster coverage (zero uncovered hours, monitored weekly).
- **Health metrics**: alert volume and signal-to-noise ratio per team (trending toward target); off-hours pages per engineer per month (trend, cap enforcement); on-call satisfaction survey (target: >7/10); training completion by team (target: 100% within 30 days); % of services with active detection contract (target: 100%).
- **Never publish incident count as a team metric.** Reward hiding. Instead publish detection metrics (near-misses reported per team, detection gaps closed, false declarations made).
- **Live dashboards**: build dashboards visible to all engineers (not just managers) showing outcome, process, and health metrics. Auto-populate from incident platform and alert tool. Update daily. Link from Slack and internal wiki.
- **Baseline all metrics against S2 evidence pack.** Set 90-day and 12-month targets for each metric. Publish targets and progress monthly.
- **Review cadence**: (1) weekly 30-min operational review (incident by incident from prior week: what went well, what hurt, actions); (2) monthly 60-min reliability review (trends, top causes, action-item aging, alert quality per team); (3) quarterly 60-min executive review (CEO's office: customer impact, SLA credits, top five systemic causes, program ROI).
- **Quarterly process review**: what in the process wasted responder time, what confused people, what should be deleted. Solicit feedback from ICs, Comms Leads, and responders. Document changes and reasoning.
19. Pilot program with 3–4 volunteer teams (depends on: 6, 8, 11, 12, 13, 14, 15, 16, 17, 18)
Do not roll out untested to 28 teams. Run the entire process end-to-end with a small cohort using real incidents as the primary training material.
- **Team selection**: recruit 3–4 volunteers spanning criticality: one payment-path team, one ledger-adjacent team, one shared infrastructure team (platform or Kubernetes), one low-traffic team. Volunteers see early adoption and influence.
- **Full process in pilot**: new severity and class taxonomy (S4), consolidated tooling (S11), roles and Triage Owner (S5–6), three rotations (S7), escalation automation (S12), communications (S13–14), postmortems (S15), action tracking (S16), paid on-call (S8), training (S17), metrics (S18). This is not a partial test; it is the complete system.
- **Real incidents are the training**: hold a retro within 48 hours of each pilot incident (while memory is fresh). Process Owner facilitates. Discuss: what worked, what hurt, how is the runbook, is the alert tuned, did Comms template work, did roles work, was timeline auto-captured correctly. Document feedback and iterate.
- **Instrument against baseline**: measure all S18 metrics in pilot; compare results with S2 baseline. If MTTD dropped from 22 min to 8 min, that is the number you carry into the next conversation.
- **Weekly retros with pilot teams**: structured 30-min sessions reviewing policies, tooling, training, templates, and process changes. Change the written policy or tooling in response; document what changed and why so later teams see the evolution, not just the finished state.
- **Explicit exit criteria**: (1) rotation coverage sustained (≥6 engineers per rotation, zero unacknowledged pages over 3 weeks), (2) postmortems delivered on time (100% of mandatory postmortems published within 15 days), (3) action tracking established (100% of action items in backlog with owner and date), (4) metrics live (dashboards updated daily, first weekly review completed).
- **Pilot report**: document before/after numbers (MTTD, MTTR, alert noise, action completion rate, on-call satisfaction) and key process learnings. This report is the foundation for every conversation in the rollout.
20. Phased rollout sequenced by cost of failure (depends on: 16, 18, 19)
Rollout is a staged migration with readiness gates and sequencing that prioritizes visible impact, not ease.
- **Wave sequencing**: divide 28 teams into 4 waves of ~7 teams each, **ordered by incident density and customer-journey ownership** (highest-cost-of-failure teams first). Teams with the most SLA credits at stake go first; their improvement is the proof.
- **Wave spacing**: three weeks between waves. This gives each wave time to stabilize and find problems before the next cohort joins.
- **Readiness checklist per team**: (1) all services mapped and owned (no unowned services), (2) alerts cleaned to paging contract (runbook linked, severity mapped), (3) playbooks updated and tested in staging, (4) rotation staffed to ≥6 engineers, (5) team completes training module, (6) manager briefed on policy, (7) on-call compensation in effect.
- **Gate review before each wave**: Process Owner holds gate review with target teams. Move unready teams to next wave with a dated remediation plan. No exceptions, no waivers; readiness is non-negotiable.
- **Wave champion**: assign a named engineer per wave to champion the rollout, answer questions, escalate issues to Process Owner. Champions are not representatives; they are advocates and feedback collectors.
- **Communication cadence**: weekly all-hands or newsletter for 4 weeks before each wave. Explain why (owned-code-owned-pager rule, paid on-call, SLA credit savings). Use pilot numbers. Answer FAQs. Announce champion and escalation path.
- **First incident under new process**: hold a retro within 48 hours. Feed accepted process changes back through change control.
- **Retire legacy tools and processes**: at end of each wave, retire legacy alert tools, informal escalation lists, ad-hoc status-page process. No parallel processes running for >3 weeks; this prevents confusion and half-learning.
- **Sequence to avoid audit collision**: ensure no team is rolling out in the same week as the audit dry run (S21).
21. SOC 2 dry run and evidence review (depends on: 3, 20)
Convert a good working process into a provable one. Test control evidence a few months before auditors arrive, when you can still fix gaps.
- **Dry run timing**: run 6 weeks before audit window (around month 7 of this program).
- **Scope**: sample 10–15 real incidents from pilot and early rollout waves. For each incident, verify evidence artifact exists and is complete: incident record, timeline (auto-captured), severity and class declaration, roles assigned and logged, communications log (Slack + status page), postmortem (if mandatory), action items in tracker with due dates, action completion evidence (code, alert test, training record).
- **Control walkthrough**: walk through each control statement from S3 with a checklist. Is the evidence artifact present? Is it immutable? Is it searchable? Is retention adequate? Is access logged?
- **Gap remediation**: for every gap found, estimate time to fix and prioritize by audit risk. Anything risking a qualified opinion (e.g., missing postmortem, no timeline evidence) must be fixed before the audit. Test the remediation against a new incident or a resample.
- **Interview readiness**: brief 10–15 engineers who may be interviewed by auditors (ICs, Comms Leads, Process Owner, team managers). Ask them to describe the process as they actually practice it, not as written. Listen for confusion or gaps in understanding. Correct them.
- **Auditor package preparation**: assemble process documentation, sample incident records (5–10 complete golden files), training records, on-call schedules, alert quality metrics, action tracker register, and status-page archive. Organize by control. Create a table of contents and index.
- **Single audit liaison**: designate Process Owner or a small dedicated compliance person as sole point of contact for audit requests. Prevents requests scattering across 28 teams.
- **Rehearsal**: conduct mock interview with an IC and a Comms Lead. Auditors ask tough questions under pressure ("How do you know the timeline is accurate?", "What happens when both ICs are unavailable?", "Show me how you proved the alert was actionable."). Practice answering.
22. Standing governance and continuous improvement (depends on: 20, 21)
Lock in durable improvement. The classic post-audit failure is the process freezing and then decaying. This step prevents that.
- **Standing Incident Management Council**: chaired by Process Owner, monthly meetings, attendees: CTO, VP Engineering, Head of Support, Head of Compliance, one engineering manager per region, one IC, one team manager from recent wave. Agenda: metrics review, policy changes, gaps from recent incidents, escalation for contentious issues.
- **Change control mandate**: give Process Owner documented authority to change severity taxonomy, response classes, roles, communications timings, and compensation policy. Any change requires: written justification, steering group approval (monthly cadence), and documented effective date before implementation. This prevents silent drift and ensures changes are deliberate.
- **Quarterly validation**: re-validate severity and response class taxonomy against all incidents from the prior quarter. Ask: did our taxonomy correctly predict response posture? Did we misclassify? Update taxonomy if patterns emerge.
- **Annual metric re-baselining**: every 12 months, re-run measurements from S2 (alert census, incident register) to reset targets. System should improve; targets should tighten.
- **Resilience roadmap separation**: fund a distinct architectural or platform roadmap for incident prevention (reduce shared-database blast radius, multi-region failover, deploy safety, observability investments). Better incident response does not protect a single-ledger corruption or unplanned failover. These are separate problems.
- **Quarterly executive report**: CTO and VP Eng report to CEO/CFO on metric set (detection time, mitigation time, SLA credits, customer-detected %), top five systemic causes of incidents, program cost vs. credit avoidance, and strategic architecture changes in flight.
- **Public backlog of improvement ideas**: teams and engineers propose process improvements via Slack or wiki. Process Owner reviews quarterly and implements accepted ideas (e.g., "add a dashboard for detection gaps", "update postmortem template"). Publish what changed and why.
- **Celebration and learning**: share wins publicly each quarter ("We reduced MTTD from 22 min to 5 min", "Customer-detected incidents down 80%", "$800K SLA credits avoided"). Refresh training and tabletop program annually and immediately after any SEV-1 to keep the system sharp and responsive to new scenarios.
--- PROPOSAL 2 (agent deepseek-flash_refine_2, deepseek/deepseek-flash) ---
Estimated complexity: high
Success metrics: - Median time to detect falls from 22 minutes to under 8 minutes by month 6, and under 5 minutes for money-moving journeys by month 9.
- Share of customer-impacting incidents first detected by customers falls from 40% to below 10% by month 9.
- Median time to mitigate for SEV1 falls from 3 h 10 min to under 60 minutes by month 9.
- 100% of SEV1 and SEV2 incidents have a named Triage Owner within 5 minutes and a named Incident Commander within 15 minutes, from month 4.
- Zero incidents with unclear command authority lasting more than 15 minutes, measured monthly from month 4.
- Monthly page volume falls from 3,400 to under 600, with a false-positive rate below 15%, by month 5.
- No service exceeds two pages per on-call shift for three consecutive months by month 6.
- 100% of the 180 services have a named owner, a detection contract and at least one symptom-based alert by month 7.
- All 28 teams have a documented rotation, a service ownership map and at least one trained on-call engineer by month 6.
- 100% of on-call shifts are paid under a published policy from month 2, with zero on-call-attributed voluntary attrition by month 6.
- On-call satisfaction scores 7 out of 10 or better in quarterly surveys from month 6.
- The IC roster holds at least 12 certified ICs covering 24x7 with no uncovered week from month 4.
- 100% of SEV0, SEV1 and SEV2 postmortems are published internally within 15 business days from month 5.
- Action items closed within 60 days rise from 17% to above 90%, with a median action age under 30 days, by month 6.
- Status page first update is posted within 30 minutes on at least 95% of SEV1 incidents from month 4.
- Zero missed regulatory notification windows on any incident requiring notification.
- The customer-impact record is complete for 100% of customer-impacting incidents from month 5, with SLA credits reconciled monthly.
- Annual SLA credits paid fall at least 50% from the $1.3M baseline within 12 months.
- Customer-impacting incidents fall at least 40% year over year from the 31-incident baseline.
- Review cadence is sustained: weekly operational review in at least 90% of weeks, 12 of 12 monthly reliability reviews, 4 of 4 quarterly executive reviews.
- Every page has a recorded disposition — fixed, tuned or deleted — within 10 working days, from month 4.
- SOC 2 Type II is passed with zero findings related to incident response.
- 100% of new engineers complete the incident-response onboarding module within 30 days of joining.
Steps (24):
1. Mandate, one owner, and the evidence clock
This step turns the CEO's email into a funded programme with a single accountable owner, and it starts the SOC 2 clock on day one.
- Appoint a Director of Incident Management reporting to the CTO, with a dotted line to the COO for customer and SLA matters and to the Head of Compliance for audit readiness.
- Publish a one-page charter: scope covers every customer-impacting, money-moving, security and data-integrity event across all 28 teams and both regions.
- Grant explicit authority to declare an incident, set severity, freeze deploys, page any engineer in the company, and approve customer messaging.
- Fund the envelope up front: tooling, training and drill time, on-call compensation, and a small programme team, roughly $500–700K a year against $1.3M in credits paid.
- State the return plainly to the steering group: credits avoided, churn avoided, and audit findings avoided.
**Start the evidence clock now.** A SOC 2 Type II report covers a period of operating effectiveness, so evidence cannot be backfilled; an audit in eight months means the operating period begins in the first weeks of this programme.
- Stand up a fortnightly steering group of the CTO, VP Engineering, Head of Support, Head of Compliance and two engineering managers.
- Open a programme risk register with the top risks, owners and review dates, and revisit it at every steering group.
- Make participation in the incident process a documented performance expectation for every engineering manager, not an optional extra.
2. Baseline evidence pack and cost-of-downtime model (depends on: 1)
You cannot demonstrate improvement without a defensible starting point, and you cannot win the argument about noise without numbers.
- Build a 12-month incident register and re-classify all 31 customer-impacting incidents by detection source, duration, customers affected and credits paid.
- Build the silent-failure register: incidents where no internal alert fired at all, which is the number that explains the 40% customer-detected rate.
- Run an alert census per tool, per team and per service: volume, page-to-action ratio, off-hours interruptions, and the 50 noisiest rules with a named owner each.
- Survey every on-call engineer and manager on burden, fairness, escalation clarity and pay expectations, targeting above 70% response.
- Interview Support, Account Management and Sales about how customers learn of incidents and what they are promised.
- Reconstruct the two command-ambiguity incidents minute by minute to find exactly where ownership lapsed.
- Build a cost-of-downtime model: dollars per minute of impact per customer journey, used later to sequence teams and justify funding.
- Publish the pack internally as the problem statement, and retain every artifact as management-review evidence for the audit.
3. SOC 2 control mapping and evidence architecture (depends on: 1)
Most programmes leave compliance to the end; this one maps controls in month one, because the mapping decides what the process must capture from day one.
- Map the process to the Trust Services Criteria for incident identification, evaluation, response, containment, recovery and communication (CC7.1–CC7.5, CC2.2, CC2.3, CC4.1, CC3.x).
- Write each control in plain language, with one named owner and its evidence artifact.
- Define the golden incident file: one auto-assembled, single-click export per incident containing timeline, roles, severity, communications, postmortem, actions and closure proof.
- Set retention, storage location and immutability so no control depends on a laptop, a private channel, or chat history that expires.
- Keep a gap register with owners and dates, reviewed fortnightly by the steering group.
- Meet the auditor's readiness team inside the first 90 days to test the control design before anything is built on top of it.
4. Severity and class taxonomy with the trigger matrix (depends on: 2)
Severity decides how loud the response is; class decides which playbook runs and what must never be traded away.
- Define severity in customer and money terms: customers affected, dollars at risk, regions lost, ledger integrity, settlement delay, data loss and SLA exposure.
- Use SEV0 for security, privacy or regulatory events; SEV1 for total or material loss of a payment path; SEV2 for degradation or single-region loss; SEV3 for limited impact with a workaround; SEV4 for internal-only issues and near-misses.
- Define response classes: Availability, Performance, Data integrity and Ledger, Security and Privacy, Third-party dependency, Settlement and Financial, Process failure.
**Class can raise the response but never lower it.** A SEV2 data-integrity incident gets SEV1 posture, because integrity failures are not recoverable by moving faster.
- Specify automatic triggers: loss of one AWS region, ledger write failures, replication lag above threshold, a missed settlement cut-off, payment success rate below threshold for five minutes.
- State who may declare — any engineer, Support agent or account manager — and who may downgrade: the Incident Commander alone.
- Map every level to its SLA credit exposure and to the customer-visible status page state.
- Include worked examples from the last 12 months so each of the 28 teams recognises its own incidents, and revalidate the taxonomy quarterly against real declarations.
5. Day-one operating rules and the minimum viable process (depends on: 1, 4)
The full process will take months; the first useful version must be live in two weeks using the tools that already exist.
- Publish ten day-one rules that need no procurement: a named owner within five minutes, one channel per incident, one register entry per incident, one person speaking to customers.
- Apply the ambiguity rule: if two responders disagree about whether this is an incident, it is an incident.
- Make declaring free: a false alarm is closed as a false declaration, tracked as a metric, and never criticised.
- Require a register entry within 24 hours for every customer-impacting incident, even a minimal one.
- Ban silent incidents: if we know, the customer hears it from us rather than from their own reconciliation.
- Run the first 30 days on manual command, with a rotating duty Incident Commander drawn from the 12 teams that already have on-call.
- Hold a 15-minute daily incident stand-up during month one to catch friction while it is still fresh.
6. Roles, command structure and the no-unowned-minute rule (depends on: 4)
The two hour-long command failures did not happen at declaration; they happened in the gap before it, when an alert had fired and nobody owned it.
- Publish one-page role cards: Triage Owner, Incident Commander, Deputy IC, Operations Lead, Internal Comms Lead, Customer Comms Lead, Scribe, Subject-Matter Responders, and Executive Sponsor for SEV1 only.
- Introduce the Triage Owner rule: whoever acknowledges the page owns the incident until an IC takes over or the incident is stood down.
**The IC owns the incident, not the fix, and does not debug.** An IC who starts troubleshooting has abandoned command.
- Give the IC explicit decision rights: declare and escalate severity, freeze changes, halt deploys, approve customer messaging, and pull in any engineer.
- Define minimum viable staffing per severity: SEV1 fills every role; SEV2 staffs IC, scribe, comms and responders; SEV3 staffs an IC and a scribe.
- Set handover discipline: four-hour maximum IC shifts on SEV1 with a written handover, and a deputy named within 15 minutes of declaration.
- Set responder behaviour: one channel, one bridge, no side channels, and every request phrased with a named owner and a time.
- Link the role cards from every paging notification so they are one tap away at 3 AM.
7. Lifecycle, declaration and escalation policy (depends on: 4, 6)
This step defines the mechanical path from an alert to a declared incident and back to normal service, removing judgment calls from the worst moments.
- Define states with entry and exit criteria: Detected, Triaged, Declared, Mitigated, Resolved, Postmortem, Closed, plus a Watch state with a hard 30-minute timer.
- Set targets: page acknowledged within five minutes, triage decision within fifteen, declaration within thirty for anything customer-facing.
- Build automatic escalation ladders with timers at every layer: responder, service owner, team manager, IC on-call, VP Engineering.
- Define the unresponsive-team path: fifteen minutes escalates to the team manager, thirty minutes to the IC, who may then direct any available engineer.
- Define change freeze and rollback authority during SEV1 and SEV2, with the single named condition that lifts the freeze.
- Enforce one incident, one record, with the timeline captured automatically from the channel and bridge rather than written from memory afterwards.
- Add the two-simultaneous-incidents rule: a second IC is designated, and no IC ever runs two incidents at once.
- Test every escalation path weekly with synthetic pages, and adjust the timings after the first month of real operation.
8. On-call architecture across 28 teams (depends on: 4, 6)
The objection is that engineers will not carry a pager for another team's code; the answer is to build rotations that make the objection structurally impossible.
- Run a Service On-Call rotation per team, covering only that team's own services.
- Run a Platform Duty rotation for genuinely shared infrastructure: the PostgreSQL ledger cluster, Kubernetes, networking, CI/CD and observability.
- Run a central Incident Commander roster of 12–16 certified senior engineers on one-week shifts with a primary and a secondary.
**State the consequence honestly.** Sixteen of 28 teams have no rotation today; each must build one or formally transfer service ownership to a team that has one, with the transfer dated and recorded.
- Merge small or low-traffic teams into shared rotations so no rotation has fewer than six engineers.
- Cap load in the scheduling tool: no engineer on call more than two weeks per quarter, enforced by configuration rather than negotiation.
- Publish a coverage matrix of all 28 teams showing services, owners, rotation size and gaps, reviewed monthly.
- Write on-call participation into engineering job levels and hiring criteria, so the commitment is set before someone joins.
9. Compensation, rest and the economics of opting out (depends on: 8)
Unpaid on-call is the most cited reason for resistance, so settle compensation before rollout, not during it.
- Move to paid on-call: a per-shift stipend benchmarked to the New York market, agreed with HR and Finance, published with an effective date before any team is asked to join a rotation.
- Pay for callouts at 1.5× the hourly rate for time actually spent mitigating, with a minimum block per interruption.
- Provide documented compensatory rest: no normal working day after a night incident, and the rest day is policy rather than a favour granted by a manager.
- Cap intrusion with a maximum number of off-hours pages per shift, and trigger a mandatory review whenever the cap is exceeded.
**Allow opt-out, but put a price on it.** An engineer may step out of a rotation, and their team buys coverage from the paid pool at a published internal rate, which turns a cultural argument into a visible budget decision.
- Publish amnesty: incident records, near-misses and false declarations are never used in performance reviews; only failure to report is a performance issue.
- Protect against gaming: pay attaches to shifts and callouts, never to incident counts.
- Check the New York labour, overtime and tax treatment with Legal and Finance before announcing, and review the policy every six months against real page volumes, attrition and survey results.
10. Detection strategy: journeys, synthetic signals and customer-report intake (depends on: 4)
Customers detected 40% of incidents first, which makes detection the highest-leverage business problem in this programme.
- Define SLIs and SLOs for the top 20 customer journeys, measured per region: payment initiation, authorisation, settlement, ledger read and write, API availability, webhook delivery and payout.
- Alert on symptoms against those SLOs, not on cause-based infrastructure metrics.
- Add synthetic transaction monitoring from outside AWS in both regions plus a third location, on a one-minute cadence for money-moving paths.
- Add ledger and shared-PostgreSQL signals: replication lag, connection saturation, write latency, lock waits, transaction ID exhaustion, checkpoint pressure and bloat.
- Open a customer-report intake so Support and account managers can raise an incident directly, and count that path as a detection source in all reporting.
- Apply the detection-gap rule: whenever a customer reports first, a detection-gap ticket is opened automatically and owned by the responsible team.
- Publish a detection contract per service — owner, at least one symptom alert, documented expected detect time — for all 180 services.
- Run a detection drill per team: break something in staging and see whether it pages before a human notices.
11. Paging contract and the noise-reduction programme (depends on: 2, 10)
3,400 alerts a month at 85% noise is the reason engineers resent the pager, and fixing it is the price of admission for everything else in this plan.
- Publish the paging contract: every page must be symptom-based, actionable, owned, mapped to a severity and class, and linked to a runbook.
**No runbook, no page**, enforced by a CI check on the alert definition itself.
- Separate paging from ticketing: only customer-impacting or imminent-impact signals may wake a human.
- Set a page budget per team and per service, with a remediation ticket opened automatically, owned by the engineering manager, when the budget is breached.
- Put new alerts on two-week probation as ticket-only until they have proved actionable.
- Define the default action for a noisy alert: fix, tune or delete within ten working days, with deletion treated as a legitimate and celebrated outcome and the deleted count published.
- Deduplicate and correlate at the ingest pipeline so one root cause produces one page instead of forty.
- Require an expiry date on every silencing rule and temporary threshold, so suppression cannot quietly become permanent blindness.
- Run a 90-day noise sprint with a public burn-down of the top 100 noisiest rules, each owned by a named manager.
12. Incident tooling consolidation and the golden incident file (depends on: 3, 6, 10, 11)
Six alerting tools and no single incident record are structural causes of the 22-minute detection and the three-hour mitigation.
- Choose one incident platform for paging, schedules, escalation policies, incident records, postmortem workflow and status-page integration, and choose it within four weeks.
- Consolidate the six alert sources into a single event pipeline feeding that platform, with deduplication, correlation and severity and class mapping applied at ingest.
- Integrate observability, runbooks, chat and voice bridges so responders see dashboards and runbooks inside the incident record, and the timeline is captured automatically.
- Implement the golden incident file export defined in S3, so audit evidence is one click rather than a reconstruction.
- Run a dual-run period alongside the legacy tools with a defined rollback, then switch the legacy tools off on a published cutover date.
- Host the status page outside the production failure domain so it survives a total platform outage, and prove that in a game day.
- Make the platform usable from a phone: an Incident Commander must be able to run a SEV1 from a mobile device at 3 AM.
13. Internal, customer and regulator communications (depends on: 6, 7, 12)
Today the status page is written by whoever is around; this step replaces improvisation with a clock, a named owner and pre-cleared templates.
- Set the internal cadence: first update within 15 minutes of declaration, then every 30 minutes for SEV1 and hourly for SEV2, whether or not there is progress.
- Never let an employee learn of an incident from the status page: internal communication leads, external follows.
- Set the customer cadence: status page updated within 30 minutes of a SEV1 and 60 minutes of a SEV2, with a no-new-information update still mandatory.
- Define the channel hierarchy: status page for all 2,100 customers, direct email to affected customers on SEV1, and named account-manager calls for the top 50 accounts.
- Pre-approve templates per severity and class with Legal and Compliance, each carrying its own next-update time.
- Forbid speculation: customer messages never guess at cause, never assign blame, and never commit to a root cause before the postmortem.
- Define the closing communication: resolution notice, how SLA credits will be handled, and the committed date for the written report.
- Build a regulator clock matrix covering event type, regulator, notification window, signer and the shortest applicable clock, including NYDFS Part 500, money-transmitter and banking notifications, breach notification, card-network rules and public-company disclosure.
**The regulatory clock starts at awareness, not at root cause.** Route every notification through Compliance, never Engineering, and pre-clear the templates.
14. Customer trust workstream and the SLA credit ledger (depends on: 13)
The $1.3M in credits is a symptom of having no single record of customer impact, and the CEO's inbox is a symptom of customers learning things late.
- Maintain one durable customer-impact record per incident: which customers, which journeys, from when to when, and the estimated credit.
- Use that one record for communications, credit computation, regulatory reporting and the annual review, so nobody reconciles two versions of the same outage.
- Automate the credit path end to end — computation, approval, customer notification and finance treatment — so credits stop being a manual scramble three weeks later.
- Reconcile accrued against paid credits monthly and report the result in the executive review.
- Give the status page a named product owner and map its components to customer journeys, not to internal services.
- Send a CTO-signed reliability note to the top 50 accounts and publish a quarterly reliability report to all customers.
- Give account managers a script of the facts they may state, the speculation they may not, and a path for customer escalations.
- Track credit avoidance against programme cost, so the funding case stays a number rather than an argument.
15. Postmortems: mandatory set, three levels, blameless by design (depends on: 6)
Postmortems currently happen for some incidents, in various formats; this step makes them mandatory where they matter and light where they do not.
- Make postmortems mandatory for every SEV0, SEV1 and SEV2, every SEV3 with customer impact or a repeat pattern, every near-miss the IC flags, and every incident where the process itself failed.
- Use three levels so the ritual matches the weight: a lightweight async review for SEV3 and SEV4, a facilitated postmortem for SEV2, and a full review with an executive sponsor for SEV1.
- Set deadlines: draft within five business days, blameless review within ten, internal publication within fifteen.
- Adopt one template: impact, timeline, detection, response, contributing factors across tooling, process, organisation and human factors, what went well, what went badly, and actions.
- Train a pool of blameless facilitators and require a trained one for every SEV1 review, never the IC.
**Ban blame language in the template and ban "human error" as a root cause.** The question is always what system condition made the error possible.
- Publish all postmortems internally by default, with security review only for genuinely sensitive material.
- Produce a customer-facing root-cause report for SEV1 incidents, especially for regulated and top-tier accounts.
16. Action items: capped, verifiable, with the repeat-incident rule (depends on: 15)
Eleven of 64 action items closed is not a tracking problem; it is a generation problem, because the process produces more actions than the organisation can absorb.
- Cap each postmortem at three action items, with anything beyond that going into a ranked reliability backlog.
- Require every action to have a named human owner, a due date, and a definition of done that is an artifact: a merged change, an alert that fired in a drill, or a test that fails without the fix.
- Prohibit self-reported closure; closure requires the artifact and sign-off by the process owner or the Incident Commander.
- Keep one reliability backlog in the engineering tracker with a mandatory label and a link back to the originating incident.
- Reserve a fixed share of each team's sprint capacity for reliability work, with unspent capacity visible to the vice presidents.
- Apply the repeat-incident rule: a second incident with the same contributing factor triggers a design review owned by an architect, not another action item.
- Run a weekly ageing review of open actions and escalate anything more than 30 days overdue to the VP of Engineering.
- Report action completion rate and median action age monthly, by team.
17. Metrics, dashboards and the review cadence (depends on: 2, 4, 16)
Define what good looks like, then measure it in a way that rewards reporting incidents rather than hiding them.
- Define outcome metrics: time to detect by source, time to mitigate by severity and class, share of incidents detected by customers, incidents by severity and class, SLA credits paid, and customer-impact minutes.
- Define process metrics, because the process itself has SLOs: IC assigned within five minutes, page acknowledgement rate, first-update timeliness, update-cadence adherence, postmortem timeliness, action closure rate, action median age, and IC roster coverage.
- Define health metrics: alert volume and noise ratio per team, off-hours pages per engineer, on-call satisfaction, training completion, and share of services with a detection contract.
**Never publish incident count as a team metric.** It rewards hiding incidents; publish reporting metrics instead — near-misses filed, detection gaps closed, false declarations made.
- Publish live dashboards visible to every engineer, refreshed daily, with every metric baselined against the S2 evidence pack and given a 90-day and a 12-month target.
- Run a 30-minute weekly operational review incident by incident, a monthly reliability review of trends and actions, and a quarterly executive review with the CEO's office.
- Hold a quarterly review of the process itself: what wasted time, what confused responders, and what should be deleted.
- End every review with decisions and named owners, never with numbers alone.
18. Training, certification and the drill programme (depends on: 6, 7, 13, 15)
A process that lives only on a wiki page fails on the first real page, so skills are built and tested before they are needed.
- Build a practical curriculum: how to be on call, how to triage and declare, how to run an incident as IC, how to communicate, and how to write a postmortem.
- Require certification before joining the IC roster: a written assessment plus a live simulated incident.
- Certify at least two ICs per team group so the central roster has depth across all 28 teams and no holiday week is left uncovered.
- Train communications leads separately on templates, cadences, customer language and the regulatory rules.
- Run monthly tabletops drawn from the last 12 months of real incidents, including region loss, ledger write failure and a missed settlement cut-off.
- Run quarterly game days with deliberately injected failure, including PostgreSQL failover, status-page outage and alerting-pipeline outage.
- Drill the process's own failure modes, not just technical ones: IC unreachable, comms lead on leave, two simultaneous SEV1s, a paging storm, and a false alarm that burns an hour.
- Audit the process for single points of failure: who alone can perform each critical task, and what happens in their holiday week.
- Keep a mandatory onboarding module for every engineer joining or transferring in, with audit-ready completion records.
19. Pilot with three to four teams, using real incidents (depends on: 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 18)
Do not roll out to 28 teams untested, and do not rely on drills alone when real incidents are better training material.
- Recruit three to four volunteer teams spanning criticality: a payment-path team, a ledger-adjacent team, the platform team that owns the shared PostgreSQL cluster, and a low-traffic team.
- Run the entire process end to end in the pilot: severity and class, roles, escalation, communications, postmortems, action tracking, and paid on-call.
- Treat real incidents during the pilot as the primary training material, and hold a retro within 48 hours of each one, run by the process owner while the friction is fresh.
- Instrument the pilot against the baseline and publish before-and-after numbers.
- Hold weekly retrospectives with the pilot teams, and change the written policies, the tooling and the training in response, documenting what changed and why.
- Set explicit exit criteria: rotation coverage achieved, zero unacknowledged pages over two weeks, postmortems delivered on time, and actions tracked to closure.
- Produce a pilot report that every later rollout conversation starts from.
20. Phased rollout to 28 teams, sequenced by cost of failure (depends on: 16, 18, 19)
Rollout is a staged migration with readiness gates, not an email announcement, and the sequencing matters more than the schedule.
- Split the 28 teams into four waves of roughly seven, ordered by incident density and customer-journey ownership: the teams whose incidents cost the most SLA credits go first, because their improvement is the visible proof.
- Leave three weeks between waves and define a per-team readiness checklist: services mapped and owned, alerts cleaned to the paging contract, runbooks written, rotation staffed to at least six engineers, training complete, manager briefed, compensation in effect.
- Hold a gate review with the process owner before each team joins, and move an unready team to the next wave with a dated remediation plan rather than granting an exception.
- Give each wave a named champion and explain the why using the pilot's numbers.
- Handle resistance directly and early: publish the own-your-code-own-your-pager rule and the paid on-call mechanics before each wave, never after.
- Hold a retro within 48 hours of each wave's first incident under the new process, and push accepted changes through change control.
- Retire legacy tools, informal escalation lists and the ad-hoc status page process at the end of each wave on a published cutover date, so two processes never run in parallel for long.
- Sequence the waves so no team faces both a rollout and the audit dry run in the same week.
21. Audit dry run and evidence review (depends on: 3, 20)
This step converts a good process into a provable one, about six weeks before the auditors arrive.
- Sample real incidents from the pilot and the early waves against each control's evidence requirements.
- Remediate every gap found and re-test the remediated control against the same sample, prioritising anything that risks a qualified opinion.
- Brief every engineer who may be interviewed so they can describe the process as they actually practise it, not as it is written.
- Prepare the auditor package: process documentation, sample incident records, paging logs, communication logs, training records, on-call schedules, the action register and alert quality metrics.
- Designate one audit liaison and a small evidence-request team, so requests do not land on all 28 teams at once.
- Rehearse the walkthrough with an Incident Commander and a communications lead, because auditors probe realism under pressure.
- Keep the audit liaison and the process owner as close to the same person as possible, so accountability for the control is also accountability for the evidence.
22. Standing governance and process ownership (depends on: 20, 21)
The classic post-audit failure is that the process freezes and then decays, so ownership has to outlive the programme.
- Establish a standing Incident Management Council chaired by the process owner, meeting monthly with engineering, support, compliance and product representation.
- Give the process owner a documented mandate to change standards, with formal change control for any change to severity, classes, roles, communication timings or compensation.
- Revalidate the severity and class taxonomy quarterly against real declarations, and re-baseline the entire metric set annually.
- Report quarterly to the executive team on the metric set and the top five systemic causes of incidents.
- Feed incident themes into architecture review, release engineering and hiring, so the programme improves the system and not only the response.
- Keep a public backlog of process improvement ideas so any engineer can propose a change, and act on it every quarter.
23. Resilience track for the shared ledger and multi-region blast radius (depends on: 1, 17)
Better incident management cannot save a single shared ledger during a corruption event, so this runs as a funded engineering track beside the process, not inside it.
- Reduce ledger blast radius: read replicas for read paths, connection-pool isolation per service domain, statement timeouts, write throttles and per-tenant limits.
- Build a ledger protection mode: an automatic tripwire that degrades to read-only or queues writes before corruption can spread.
- Rehearse multi-region failover with the ledger as the binding constraint, and set a measured recovery time objective per region.
- Test restore and point-in-time recovery on a schedule, and publish the timings.
- Tighten deploy safety: progressive rollout, automatic rollback on SLO burn, and change-freeze windows around settlement.
- Track a small number of resilience bets with named owners and quarterly milestones, reviewed in the same executive review as the incident metrics.
24. Culture, incentives and continuous improvement (depends on: 20, 22)
Incident management is a cultural change before it is a process change, and the incentives decide whether reporting goes up or down.
- Restate the amnesty at every rollout wave: reporting, near-misses and false declarations are safe; hiding is not.
- Recognise detection improvements, deleted alerts, closed detection gaps and well-written postmortems publicly, and reward the engineers who do them.
- Correct blame language in leadership communication on the spot, including inside the executive review.
- Watch for pager fatigue with real data, and take an engineer off rotation as soon as the intrusion cap breaks.
- Run a quarterly process retro with the whole engineering organisation and publish what changed as a result.
- Refresh the training and tabletop programme annually and immediately after any SEV1.
- Share wins publicly each quarter — faster detection, fewer customer-detected incidents, credits avoided — to hold momentum after the audit.
--- PROPOSAL 3 (agent qwen3.8-flash_refine_3, alibaba/qwen3.8-flash) ---
Estimated complexity: high
Success metrics: - Median Time to Detect (MTTD) < 5 minutes.
- Median Time to Mitigate (MTTR) < 60 minutes for SEV1.
- >90% of incidents detected internally (vs. 40% previously).
- SLA credits reduced to <$100K annually.
- Monthly alert volume < 600 with >95% signal-to-noise ratio.
- Postmortem action item completion rate > 80%.
- Zero incidents with command ambiguity > 15 minutes.
- 100% of SEV1/2 postmortems published within 15 days.
- Status page updates within 30 minutes for 95% of SEV1.
- 100% of teams integrated into on-call rotations.
- On-call satisfaction score > 7/10.
- 40+ certified Incident Commanders available 24x7.
- SOC 2 Type II audit passed with zero findings.
Steps (22):
1. Executive Mandate & SOC 2 Evidence Clock
Establish clear ownership and funding to transform the CEO's concern into a governed program with immediate compliance implications.
2. Baseline Measurement & Problem Statement (depends on: 1)
Quantify the current state to build a defensible problem statement and establish the data foundation for improvement tracking.
3. Severity & Class Taxonomy (depends on: 2)
Define unambiguous incident levels based on business impact to trigger specific response protocols.
4. Incident Command Roles & Structure (depends on: 3)
Standardize roles and decision rights to eliminate command ambiguity during active incidents.
5. Tooling Consolidation & Integration (depends on: 1)
Deploy a unified incident management platform to replace fragmented alerting tools and create a single source of truth.
6. Detection Strategy & SLO Implementation (depends on: 3)
Shift monitoring focus to customer-impacting symptoms to detect issues before customers do.
7. Alert Quality Standards & Noise Reduction (depends on: 5, 6)
Implement strict alerting rules to reduce noise and restore on-call trust.
8. On-Call Architecture & Coverage Model (depends on: 4)
Design a federated, team-owned on-call structure supported by central command coverage.
9. On-Call Compensation & Wellbeing Policy (depends on: 8)
Introduce paid, sustainable on-call rotations to remove resistance and mitigate burnout.
10. Playbooks & Communication Templates (depends on: 3, 4)
Create standardized playbooks and templates for each severity level to reduce cognitive load.
11. Escalation & Incident Lifecycle Policy (depends on: 4, 5)
Define the mechanical progression from alert to resolution, including automatic escalation paths.
12. Communication Workflows (Internal & External) (depends on: 4, 10)
Establish timed protocols for internal, customer, and regulatory communications.
13. Status Page & Customer Notification Automation (depends on: 5, 12)
Integrate incident management with public status pages to ensure timely and accurate updates.
14. Postmortem Policy & Blameless Framework (depends on: 3)
Mandate and standardize postmortems to focus on system improvement rather than individual error.
15. Action Item Tracking & Governance (depends on: 14)
Ensure postmortem action items are tracked, owned, and completed to prevent recurrence.
16. Metrics, Dashboards & Review Cadence (depends on: 2, 3)
Define and monitor key performance indicators to validate the effectiveness of the incident process.
17. Training, Certification & Drills (depends on: 4, 10, 12)
Build practical skills and readiness for incident response across all roles.
18. Pilot Program with Volunteer Teams (depends on: 5, 7, 9, 11, 13, 15, 17)
Test the end-to-end process on a small scale to refine procedures before full rollout.
19. Phased Rollout to All Teams (depends on: 18)
Deploy the process to all 28 teams in manageable waves to minimize disruption and manage resistance.
20. SOC 2 Control Mapping & Evidence Framework (depends on: 1, 3)
Align the incident management process with SOC 2 requirements to ensure audit readiness.
21. SOC 2 Dry Run & Audit Support (depends on: 19, 20)
Verify evidence sufficiency and prepare the team for the external audit.
22. Governance & Continuous Improvement (depends on: 19, 21)
Establish standing governance to sustain the process and drive ongoing improvements.
Your answer has these parts:
- "round_summary": one or two sentences on how the round went as a whole.
- "converging": true if the proposals of this round are more similar to each other than those of the previous round, false otherwise.
- "differences": a short list (four items at most) of what still separates the proposals of this round, naming the proposals and the steps concerned.
- "influences": a short list (five items at most) of the flows of ideas of the round: who took what from whom, and what nobody took, naming proposals and steps ("P1 and P3 took P2's CDC sync (steps 2 and 13)").
- "proposals": one entry per proposal of round 2, each with:
- "proposal": its number,
- "assessment": "improved", "worsened", "mixed" or "unchanged" with respect to its previous version ("no_previous_version" if that agent produced nothing in the previous round),
- "what_changed": how it improved or worsened and why: two or three short sentences, or a list,
- "improvements": a list of concrete gains (specific steps, metrics, structure),
- "regressions": a list of concrete losses (dropped steps, vaguer metrics, broken dependencies...),
- "taken": the ideas this proposal visibly adopted from the OTHER proposals of round 1 (not from its own previous version): one entry per idea with "from_proposal" (the number of the proposal it came from), "steps" (the numbers of the steps of that proposal where the idea lives, as listed above; empty if it is not tied to specific steps), "what" (the idea, one sentence) and "why" (how it was used or adapted, one sentence),
- "rejected": the ideas of the OTHER proposals of round 1 that this proposal visibly declined: an explicit contradiction, or a prominent idea it saw and left out while taking the opposite approach. Same fields; "why" gives the evidence (what the proposal does instead). Do not list mere omissions without evidence; an empty list is a valid answer.
[FINAL]
[SYSTEM]
You are an expert reviewer of multi-agent planning processes.
Several LLM agents drafted plans for a task, refined them over a number of rounds while seeing each other's proposals, and finally voted for the best one.
Be exhaustive but precise: name concrete steps, ideas and metrics, never generalities. Judge plans by their fitness for the task as stated, their realism, their completeness, the soundness of their order and dependencies, how measurable their success is and how they handle things going wrong.
You are an impartial evaluator, not a chronicler: assess the proposals and the process on their merits, never rationalise what happened or assume that the outcome was right.
After your analysis, answer in the requested structure.
Every text field you write will be read by a busy person who skims. Make it easy to skim: short sentences and short paragraphs; when you name several things, prefer a list to a paragraph, with sub-items when an item has parts, but keep a single fact as a sentence; lead with the point and then the evidence; name proposals and steps by number (P2, step 4); no preamble, no repetition of the question, no closing summary; bold at most one key phrase per item or paragraph. Text fields accept Markdown: a blank line between paragraphs, "- " for lists, **bold**.
[HUMAN]
Task given to the agents: "A B2B payments platform in New York (260 engineers in 28 teams, 2,100 customers, $4B processed a month, a 99.95% SLA with service credits) runs about 180 services on Kubernetes across two AWS regions, with a shared PostgreSQL cluster for the ledger. In the last 12 months: 31 customer-impacting incidents, median time to detect 22 minutes (customers detected 40% of them first), median time to mitigate 3 h 10 min, $1.3M paid in SLA credits, two incidents in which nobody was sure who was in charge for over an hour, and a CEO email about "outages we hear about from clients". On-call exists in 12 of the 28 teams, unpaid, fed by alerts from six different tools — 3,400 alerts a month, 85% of them noise; there is no severity scale, status-page updates are written by whoever is around, and postmortems happen for some incidents, in various formats, with action items that are rarely tracked (11 of 64 closed). A SOC 2 Type II audit in eight months will test incident response. Engineers push back against "carrying a pager for other teams' code".
Define the incident management process: severity levels and what each one triggers; roles (incident commander, communications, scribe, subject-matter responders) and how they are staffed 24x7 across 28 teams; the on-call structure, rotations, compensation and the rules for alert quality; detection and escalation paths; internal and customer communications (status page, account managers, regulators when required) with their timings; postmortems (when mandatory, format, blameless review, ownership and tracking of actions); the metrics and reviews that show whether it works; and how it is introduced across the teams without waiting for the audit."
THE INITIAL PROPOSALS (round 0):
--- PROPOSAL 1 (agent claudeHaiku4.5_initial_1, anthropic/claude-haiku-4-5) ---
Estimated complexity: high
Success metrics: - Median time to detect reduced from 22 minutes to <8 minutes within 6 months
- Customer-detected incidents drop from 40% to <5% within 6 months
- Median time to mitigate (MTTR) reduced from 3h 10min to <45 minutes for SEV-1 incidents within 6 months
- Annual SLA credits decrease from $1.3M to <$100k within 12 months
- Alert noise reduced from 3,400 per month (85% false positive) to <400 per month (>95% signal) within 3 months
- Zero incidents with command-and-control ambiguity (>1 hour without clear IC) within 2 months
- Postmortem action item completion rate reaches >80% (from 17%) within 4 months
- On-call satisfaction score (survey) reaches >7/10 for on-call engineers within 3 months
- All 28 teams integrated into incident management system with active on-call rotations by week 20
- SOC 2 Type II audit passes incident response controls with no findings 8 months from start
- Incident commander certification: 100% of active ICs trained and drilled within 2 months
- Monthly incident review meeting established and attended by leadership; trends documented
- New incident system integration complete: single alert tool, single dashboard, all 180 services feeding in, <5 min deployment
Steps (21):
1. Define severity levels and decision criteria
Create a four-tier severity framework (SEV-1 through SEV-4) that guides all downstream decisions about response, escalation, and communications.
Each level must specify: customer impact (revenue at risk, customers affected, data loss risk); financial threshold triggering service credits; whether an incident commander is required; response time SLA (e.g., SEV-1 < 5 min notification, SEV-4 < 2 hours); and the go-live decision tree (when to declare and when to resolve).
- SEV-1: Complete service down or critical path broken for >5% of customers; every minute costs money; IC required; 99.99% uptime threatened
- SEV-2: Significant degradation, features unavailable, affecting 1–5% of customers; IC typically required
- SEV-3: Minor impact, limited customer footprint or workaround exists; escalation path but not automatic IC
- SEV-4: Observations or minor issues; alert-driven, no escalation unless pattern emerges
2. Define incident roles and responsibilities (depends on: 1)
Create the organizational roles that operate during an incident: who is in charge, who talks to customers, who writes down what happened, who fixes the system, and how decisions are made under pressure.
Each role must have a single-sentence mission, decision authority, and escalation upward.
- Incident Commander: owns decision-making and timeline; declares severity; resolves conflicts; may or may not be technical
- Deputy IC: shadow to IC and takes over if IC becomes unavailable
- Communications Lead: writes status page, notifies account teams, manages customer perception
- Scribe: records decisions, who did what, key timestamps; not responsible for fixing
- SME Responders: engineers with context on the failing service(s); take IC's direction without debate
3. Design 24x7 on-call rotation structure (depends on: 2)
Build a rotation model that covers all 28 teams with primary and backup on-call engineers every hour across weekdays, evenings, weekends, and holidays; addresses the pager-carrying resistance.
Key design decisions: Is coverage per-team (each team owns its services) or pooled (shared responder pool handles anything)? How many people per rotation? How long are shifts (one week, two weeks)? When can engineers opt out without leaving the team exposed? Which roles are on-call (IC, communications, SME)?
- Recommend: dedicated IC pool (4–6 people in fast rotation) + per-team SME on-call for each team's own services
- Recommend: two-week rotation blocks to reduce handoff friction
- Recommend: one primary, one secondary per slot; secondary handles during primary's escalation
- Provide swaps, blackout dates, and a rule that no engineer is on-call more than 2 weeks per quarter
4. Define on-call compensation and incentives (depends on: 3)
Create a pay model that makes on-call acceptable and rewards engineers who carry the pager; ties compensation to real business risk.
- Base on-call stipend: e.g., $500–1,000 per week while on-call (regardless of incidents)
- Callback pay: 1.5× hourly rate for time spent mitigating incidents during off-hours
- Incident bonus: $50–100 extra per SEV-1 or SEV-2 incident mitigated (recognition)
- Comp time: full business day off after an incident that required >2 hours mitigation during night/weekend
- Annual bonus tie-in: 10–20% bonus multiplier for flawless on-call reviews
- Communicate: position as investment in reliability, not punishment for being online
5. Consolidate alert routing infrastructure
Replace six alert tools with a single ingestion and routing system; stop engineers from being woken by duplicate alerts, and make escalation automated instead of manual.
Evaluate existing tools (likely candidates: PagerDuty, Opsgenie, or Incident.io) or build a lightweight wrapper. The system must: accept alerts from all 180 services; deduplicate and correlate (same outage, different monitoring source); route to correct on-call engineer; expose an API for playbook automation; log every alert for postmortem analysis.
- Choose tool by week 2 of S5
- Migrate alerting endpoints from 6 sources to 1 by week 4
- Set up audit trail and retention
- Ensure mobile app works (on-call engineers need to engage from phone)
6. Build alert quality rules to cut noise (depends on: 1, 5)
Implement rules that automatically suppress the 85% of alerts that are noise (flapping, transient errors, auto-recovered conditions). Target: <400 actionable alerts per month.
Rules to implement: suppress alerts if service auto-recovered within 30 seconds; deduplicate same alert from multiple monitoring sources; suppress alerts for known maintenance windows; group flapping alerts (same service, >5 occurrences in 2 minutes) into a single page to on-call; rate-limit alerts from noisy services (e.g., max 1 alert per 5 minutes per service until silence clears).
- Audit existing 3,400 alerts per month: which are true signals, which are noise
- Tag each alert source with severity level (S1, S2, S3, S4 from S1)
- Create exceptions list: services known to be noisy, require different rules
- Weekly review: alert teams that trigger >50 alerts per week for reduction strategies
7. Implement automated detection and escalation paths (depends on: 1, 5, 6)
Wire the alert system to automatically escalate based on time or severity; removes the need for manual judgment calls during chaos.
Logic: SEV-1 alert arrives → IC notified instantly via phone call + SMS + Slack + mobile; if IC does not acknowledge within 2 minutes, page deputy IC; Communications Lead pinged simultaneously. SEV-2: on-call SME for that service + IC notify via Slack and mobile, escalate to IC's manager if not acknowledged in 10 min. SEV-3/4: on-call SME only, escalate after 30 min.
- Implement in alert routing system (S5)
- Test all paths weekly via synthetic page to on-call
- Track escalation metrics: how many pages reach secondary, how many hit manager
- Adjust timing based on first month of operations
8. Build incident dashboard and status tracking (depends on: 5)
Create a single source of truth during an incident that every responder sees in real time: who is on-call, incident timeline, who said what, current status, next steps.
Dashboard displays: active incidents and their severity; who is the IC and communications lead; timeline of all events (alert fired, IC assigned, customer notified, mitigation started, resolved); Slack channel and mobile notification status; on-call rosters (who is on-call right now for each team); postmortem link as soon as incident closes.
- Integrate with alert tool (S5) to auto-populate incident creation and initial severity
- Push updates to status page and customer account managers automatically
- Log all timeline entries for audit and postmortem completeness
- Mobile-optimized so IC can work from any device
9. Write incident playbooks for each severity (depends on: 1, 2, 3)
Create a one-page (or one-screen) reference for the IC and SMEs during an incident; sequences the steps and removes ambiguity.
Each severity level gets its own playbook: who gets paged (roles, order); first questions to ask (is it real, how big, who knows); what the IC should declare in first message (status page text, account manager notification, regulatory trigger); how long before escalating to executive team; decision rules for going dark vs. continuing to update customers.
- SEV-1 playbook: immediate IC + comms + CTO notification; customer status every 5 minutes
- SEV-2 playbook: IC + comms + tech lead notification; status every 15 minutes
- SEV-3 playbook: on-call SME + comms if customer-visible; status every 30 min or as resolved
- SEV-4 playbook: on-call SME only; update customers only if promised SLA is at risk
- Include decision trees: is this SEV-1 or SEV-2? Is it our code or dependency? Escalate or containment?
10. Define internal communication workflows (depends on: 2, 3)
Specify who informs whom, in what order, via what channel (call, Slack, email) during an incident; prevents gaps like "nobody knew who was in charge for an hour."
Workflow for SEV-1: IC assigned → IC calls CTO/VP Eng and incident channel lead within 1 minute; incident declared in #incidents Slack channel with severity, IC name, service affected; SME on-call for that service joins call automatically; IC pushes updates to #incidents every 5 minutes or when material change occurs. For SEV-2: IC notifies team leads via Slack, updates #incidents every 15 min. Define escalation: if IC is unreachable, deputy IC takes over and announces it.
- Create a phone tree or on-call list accessible to responders
- Set expectations: "If you don't hear from IC in 2 minutes, call them"
- Use a single incident Slack channel per incident (auto-created by incident tool)
- Log all comms in the incident dashboard for postmortem review
11. Design customer communication and status page process (depends on: 1, 2)
Plan when and how to inform customers, account managers, and regulators; ensure 2,100 customers are not learning about outages from Twitter before you tell them.
Rules by severity: SEV-1 detected → status page updated within 3 minutes (even if root cause unknown; post "Investigating"); account managers of affected customers called within 5 minutes; regulatory notification (if payment processing down) queued for approval; customer email within 10 minutes with ETA for next update. SEV-2: status page within 10 min, account managers called within 15 min, email if affecting >10 customers. SEV-3/4: no customer communication unless SLA at risk.
- Empower Communications Lead to update status page without IC approval if delay >3 min
- Prepare templated messages for common scenarios (database failover, data pipeline stuck, service crashed)
- Route regulatory notifications through legal/compliance; don't wait for perfect root cause
- Track customer impact in real time: how many customers affected by severity
12. Establish blameless postmortem process and format (depends on: 1, 2)
Build a systematic way to learn from incidents so the same failure does not happen twice; counter the fear that admitting a mistake leads to being blamed.
Mandatory postmortems: all SEV-1 and SEV-2 incidents, within 48 hours of resolution. Optional but encouraged: SEV-3 if interesting or if >3 of same type in 30 days. Format: what was the user-visible impact and for how long; what was the root cause (not "human error" but the system condition that made error possible); timeline of discovery and response; action items with owner and deadline; blameless tone (focus on process and system design, not individual mistakes).
- Assign a facilitator (not the on-call IC) to run postmortem
- Attendees: IC, comms lead, SMEs involved, team lead, customer success if customer-facing
- Write postmortem in shared doc; make it findable (searchable, linked from incident)
- No discussion of "who screwed up"; only "why did the system allow this to happen"
13. Build action item tracking and accountability (depends on: 12)
Create a system that tracks postmortem action items so they are not forgotten; currently 11 of 64 (17%) are being tracked, leaving 53 unfinished improvements.
System: each postmortem generates action items (e.g., "add monitoring for X," "update runbook for Y," "write test for Z"). Each item gets: clear description, owner (engineer's name), due date (1–4 weeks based on priority), severity (critical = must do before similar incident happens again; important = improve next month; nice-to-have = backlog). Action items live in a dedicated Jira project visible to all teams; owners are accountable (their manager reviews quarterly). Weekly: incident commander reviews open items due that week. Monthly: each team's postmortem items reviewed in their standup.
- Export action items from postmortem document to tracking system automatically
- Require IC to sign off that an action is complete before closing
- Report on completion rate as a metric (target: >80% by month 3)
14. Define incident metrics and KPIs
Establish what "good" looks like; measure so you can improve. Target metrics for 12 months out: mean time to detect 8 minutes (vs. 22 now), customers detect first <5% of incidents (vs. 40%), MTTR 45 minutes for SEV-1 (vs. 190), SLA credits <$100k/year.
Metrics to track: (1) MTTD = time from incident start to first alert/report; disaggregate: external report vs. internal detection. (2) MTTR = time from first report to full mitigation; track by severity and by service. (3) Customer-reported incidents per month (should drop to <2 per month). (4) Alert signal-to-noise ratio (goal: <5% false positive after S6 rules). (5) On-call satisfaction (survey: would you do this again?). (6) Postmortem action completion rate. (7) Incident commander and responder utilization (hours per week per person).
- Dashboard: auto-populated from incident tool, updated daily
- Disaggregate by team and service: which teams have bad MTTR? Which service is most incident-prone?
15. Create review cadence and governance process (depends on: 14)
Establish regular rhythm to inspect the metrics, spot trends, and adjust the process itself; prevent the system from calcifying.
Weekly: incident commander and on-call lead review prior week—number of incidents, any escalations, any communication gaps. Monthly: director-level incident review—trends by service, top causes of incidents, action item status, whether severity classification is working. Quarterly: full leadership review—MTTD, MTTR, customer impact, SLA credit spend, on-call satisfaction score, any systemic changes needed. Annually: audit the entire process for SOC 2 compliance.
- Assign meeting owners: weekly = on-call lead; monthly = director of reliability; quarterly = VP Eng + CFO (SLA cost) + customer success
- Use same data dashboard (S14) for all reviews
- Publish a monthly "incident newsletter" to all engineers: what happened, what we learned, what's improving
16. Prepare SOC 2 Type II audit checklist (depends on: 1, 2, 9, 12, 13, 14, 15)
Document that the incident management system meets the control requirements for a SOC 2 audit; audit happens in 8 months, so this work builds confidence in coverage.
Audit will test: (1) Is there a defined incident response process? (2) Are roles and responsibilities clear? (3) Are incidents logged and tracked? (4) Is root cause analysis performed? (5) Are action items tracked and completed? (6) Is on-call staffing adequate? (7) Are communications timely? (8) Are postmortems documented and blameless? Create a control mapping document that links each SOC 2 requirement to your process (S1–S15). Collect evidence: incident logs, postmortem documents, action item tickets, metrics reports, training records.
- Designate a compliance owner (often a reliability lead or security engineer)
- Run a mock audit at month 6 to identify gaps
- Ensure all postmortems and incidents are retained and searchable for auditor review
17. Develop implementation and rollout plan (depends on: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16)
Create a phased timeline to roll out the incident management system across all 28 teams; avoids big-bang failure and builds credibility.
Recommended structure: Phase 1 (weeks 1–4): build and test infrastructure (S5, S8, alerting); deploy severity levels and roles (S1, S2); pick pilot teams (2–3 high-traffic teams). Phase 2 (weeks 5–12): train pilot teams, run incident drills, refine playbooks based on learning; expand to half of remaining teams. Phase 3 (weeks 13–20): full rollout to all 28 teams; continue drills; track metrics. Phase 4 (weeks 21–28): stabilize, iterate on metrics, prepare for audit.
- Assign a release manager to coordinate across teams
- Create a detailed Gantt chart with swim lanes (infra, process, training, rollout)
- Identify risks: competing priorities, engineers worried about pager burden, tool adoption friction
- Plan stakeholder engagement: weekly updates to eng leadership, monthly town halls for all engineers
18. Build training and documentation (depends on: 17)
Create role-specific education so engineers understand the new system and are confident executing during an incident.
Training tracks: (1) All engineers: 30-min video on severity levels, communication expectations, postmortem format, action items (mandatory); (2) On-call responders: 2-hour workshop on playbooks, alert tool, escalation paths, when to call manager, case studies of real incidents, mobile app walkthrough; (3) Incident commanders: 4-hour IC bootcamp covering leadership under pressure, decision-making, communicating with executives, status page discipline, postmortem facilitation, practiced drills; (4) Communications leads: templates, when to update, how to talk to customers, regulatory notification rules.
- Record videos so async teams can learn on their schedule
- Create runbooks and quick-reference cards for each role (print + digital)
- Pair new on-call engineers with experienced responder for first week
- Require IC certification before anyone joins IC rotation (pass a practical drill)
19. Execute staged rollout across teams (depends on: 18)
Progressively activate the incident management system with feedback loops at each stage; reduces risk of system-wide failure.
Wave 1 (week 6–8): 3–4 pilot teams begin on-call rotations and incident response using new system; capture feedback daily. Wave 2 (week 10–14): 8–10 additional teams, incorporating lessons from Wave 1; ensure diversity of team types (payment processing, monitoring, data pipeline, auth, etc.). Wave 3 (week 15–20): remaining teams; by now, the system is proven and less hand-holding needed.
- Daily retros with Wave 1 teams: what worked, what was confusing, what broke
- Each wave produces a "lessons learned" document that informs the next
- Track adoption metrics: how many incidents reported per team, alert quality, MTTD/MTTR
- Address resistance: engineers who are skeptical of the system, on-call burden, tool friction; assign a "change champion" in each team
20. Run incident response drills and simulations (depends on: 19)
Practice incidents in a controlled setting so responders gain confidence and gaps are found before a real crisis.
Drill 1 (week 8, pilot teams only): simulated database failover; test alert routing, IC assignment, on-call notification, status page update, postmortem. Drill 2 (week 12, all teams): full incident spanning multiple services; test escalation, communications, decision-making. Drill 3 (week 18): chaos drill with intentional confusion—what if IC is unreachable, what if comms lead is on PTO, what if an alert tool goes down? Quarterly thereafter: repeat drills with new on-call engineers.
- Schedule drills at a time when key people are available (not Friday afternoon)
- Run drill with live observers (reliability lead, manager) who provide feedback
- Measure drill success: all steps completed on time, communications clear, postmortem written
- Debrief immediately: what did we learn, what do we change?
21. Establish continuous improvement process (depends on: 15, 20)
Build a feedback loop so the incident management system evolves based on what you learn.
Monthly improvement cycle: in the monthly incident review (S15), leadership identifies one or two friction points from recent incidents (e.g., "alert tool response is slow," "postmortem meetings are boring," "action items keep slipping"). Assign an owner to investigate and propose a fix (change to playbook, tweak to tool config, additional training, process change). Test the fix with pilot team, gather feedback, roll out if successful. Annual review: can we reduce on-call burden, simplify communication, improve MTTD further? Do any teams need special handling (new services, difficult customers)?
- Keep a public backlog of improvement ideas (Jira, shared doc) so any engineer can suggest
- Quarterly all-hands: share wins ("MTTD improved 50%," "on-call satisfaction up," "zero customer-detected incidents") to build momentum
- Link improvements back to metrics (S14): show that changes actually worked
--- PROPOSAL 2 (agent deepseek-flash_initial_2, deepseek/deepseek-flash) ---
Estimated complexity: high
Success metrics: - Median time to detect reduced from 22 minutes to 5 minutes or less by month 6.
- Customer-detected incidents reduced from 40 percent to 10 percent or less of all customer-impacting incidents by month 9.
- Median time to mitigate reduced from 3 h 10 min to 60 minutes or less for SEV1 and SEV2 by month 9.
- Customer-impacting incidents reduced by at least 40 percent year over year from the 31 incident baseline.
- SLA credits paid reduced by at least 50 percent year over year from the $1.3M baseline.
- 100 percent of SEV1 incidents have a named Incident Commander within 5 minutes and a scribe within 15 minutes.
- Zero incidents in which command authority is unclear for more than 15 minutes.
- 28 of 28 teams have a documented on-call rotation, an owning service map and at least one trained on-call engineer by month 6.
- Central Incident Commander roster holds at least 40 certified ICs covering 24x7 with no single point of failure.
- 100 percent of on-call rotations are paid under a published policy by month 5.
- Monthly alert volume reduced from 3,400 to below 700, with a false-positive rate below 20 percent.
- No service exceeds 2 pages per on-call shift, measured monthly for three consecutive months.
- 100 percent of SEV1 and SEV2 postmortems published internally within 15 business days.
- At least 90 percent of postmortem action items closed within 60 days, up from 17 percent (11 of 64).
- Status-page first update published within 30 minutes on at least 95 percent of SEV1 incidents.
- Zero missed regulatory notification windows on any incident requiring notification.
- SOC 2 Type II audit passed with no findings related to incident response.
- Review cadence sustained: weekly operational review in at least 90 percent of weeks, 12 of 12 monthly reliability reviews, 4 of 4 quarterly executive reviews.
- On-call satisfaction at 70 percent or higher on the quarterly survey, with zero on-call-attributed voluntary attrition.
- 100 percent of new engineers complete the incident-response onboarding module within 30 days of joining.
Steps (21):
1. Programme charter, ownership and executive mandate
This step turns the CEO's email into a funded programme with a named owner and explicit authority. Without it, every downstream decision stalls in cross-team negotiation.
- Appoint a single accountable process owner (for example a Director of Incident Management) reporting to the CTO, with a dotted line to the COO for customer and SLA matters.
- Publish a one-page charter covering scope (all customer-impacting and money-moving incidents), decision rights, and the power to override team preferences during an active incident.
- Define the funding envelope: tooling licences, training time, exercise time and on-call compensation, with an indicative annual figure.
- Set the timeline against the SOC 2 date: a working process in four months, evidence accumulating from month two, audit-ready by month seven.
- Stand up a steering group with CTO, VP Engineering, Head of Support, Head of Compliance and one engineering manager per region.
- Agree that incident-process participation is a documented performance expectation for engineering managers, not an optional extra.
2. Baseline measurement and evidence pack (depends on: 1)
You cannot demonstrate improvement without a defensible starting point, and you cannot win the argument about noise without numbers.
- Build a 12-month incident register: date, retro-assigned severity, detection source, time to detect, time to mitigate, customer impact, services involved and SLA credits paid.
- Quantify the alert estate per tool, per team and per service; compute page-to-action ratio, list the 50 noisiest rules and count off-hours interruptions per engineer.
- Survey on-call engineers and managers on burden, fairness and clarity of escalation, targeting a response rate above 70 percent.
- Interview Support, Account Management and Sales about how customers learn of incidents and what they complain about.
- Document exactly where the current process breaks: unclear command in the two known incidents, postmortem action closure at 11 of 64, and ad-hoc status-page authorship.
- Publish the pack internally as the problem statement and retain it as management-review evidence for the SOC 2 audit.
3. Severity taxonomy and trigger matrix (depends on: 1, 2)
Severity is the keystone of the whole process. Every other rule, from paging to communications timing to postmortems, is keyed off it.
- Define four levels plus a special SEV0 for security or regulatory events: SEV1 for total or material loss of a payment path, SEV2 for degradation or single-region loss, SEV3 for limited impact with a workaround, SEV4 for internal-only issues and near-misses.
- Anchor each level in customer and money terms: customers affected, dollars at risk, regions lost, ledger integrity, settlement delay.
- Specify automatic triggers, for example loss of one AWS region, ledger write failures, a missed settlement cut-off, or payment success rate below threshold for five minutes.
- State who may declare each level (any engineer, Support or account manager may declare) and who may only recommend a downgrade (the Incident Commander alone).
- Map each level to SLA credit exposure and to the customer-visible status-page state.
- Include worked examples from the last 12 months so teams recognise their own incidents in the definitions.
- Add a review clause: the taxonomy is re-validated quarterly against real declarations.
4. Incident roles, command structure and decision rights (depends on: 2, 3)
The two incidents where nobody knew who was in charge for over an hour are the direct brief for this step.
- Define roles with one-page role cards: Incident Commander, Deputy IC, Operations Lead, Internal Communications Lead, Customer Communications Lead, Scribe, Subject-Matter Responders and an Executive Sponsor for SEV1 only.
- State the core rule plainly: the IC owns the incident, not the fix, and does not debug.
- Give the IC explicit decision rights: declaring and escalating severity, freezing changes, halting deploys, approving customer messaging and calling additional responders.
- Define minimum viable role coverage per severity: SEV1 staffs every role, SEV3 staffs an IC and a scribe only.
- Define handover discipline: maximum four-hour IC shifts on SEV1, a written handover template, and a Deputy IC nominated within 15 minutes.
- Set behaviour standards for responders: one incident channel, one bridge, no side channels, and explicit asks with a named owner and a time.
- Publish role cards on the internal wiki and link them from every paging notification.
5. Escalation, paging and incident lifecycle policy (depends on: 3, 4)
This step defines the mechanical path from an alert to a declared incident and back to normal service.
- Define lifecycle states with entry and exit criteria: Detected, Triaged, Declared, Mitigated, Resolved, Postmortem, Closed.
- Set acknowledgement targets: page acknowledged within five minutes, triage decision within fifteen, declaration within thirty for anything customer-facing.
- Define escalation ladders per layer (responder, service owner, team manager, IC on-call, VP Engineering) each with an automatic timer.
- Make escalation blameless and automatic: no responder is ever criticised for escalating, and timers fire whether or not a human asks.
- Define change freeze and rollback authority during SEV1 and SEV2, and the single condition that lifts the freeze.
- Enforce one incident, one record: the incident record is the sole source of truth for timeline, roles and communications.
- Require every SEV1 and SEV2 to produce an automatically captured timeline from channel and bridge, never one written from memory afterwards.
6. Detection strategy: SLOs, signals and customer-journey monitoring (depends on: 2, 3)
Customers detected 40 percent of incidents first. That number is the reason this step exists.
- Define SLIs and SLOs for the top 20 customer journeys, including payment initiation, settlement, ledger read and write, API availability and webhook delivery, measured per region.
- Require symptom-based alerting on those SLOs rather than cause-based alerting on infrastructure metrics.
- Add synthetic transaction monitoring from outside AWS in both regions and a third location, on a one-minute cadence for money-moving paths.
- Add ledger and shared-PostgreSQL specific signals: replication lag, connection saturation, write latency, transaction ID exhaustion and checkpoint pressure.
- Open a customer-reported path so Support and account managers can raise an incident directly, and count that path as a detection source in reporting.
- Publish a detection contract per service: every one of the 180 services needs an owner, at least one symptom alert and a documented expected detect time.
- Fund a separate resilience track to reduce shared-cluster blast radius, because better detection will not save a single shared ledger during a corruption event.
7. Alert quality standard and noise-reduction programme (depends on: 2, 3, 6)
3,400 alerts a month with 85 percent noise is the reason engineers resent the pager. Fixing it is the price of admission for everything else.
- Publish alert standards: every page must be symptom-based, actionable, owned, linked to a runbook and mapped to a severity. No runbook, no page.
- Separate paging from ticketing: only customer-impacting or imminent-impact signals may page; everything else becomes a ticket or a dashboard entry.
- Set a noise budget per team and per service, for example no service may exceed two pages per on-call shift, measured monthly.
- Define the default action for a noisy alert: fix, tune or delete within ten working days, with deletion treated as a legitimate and successful outcome.
- Introduce correlation and deduplication at the event pipeline so a single root cause produces one page instead of forty.
- Require expiry dates on every silencing rule and temporary threshold so suppression cannot become permanent blindness.
- Run a 90-day noise sprint with a visible burn-down of the top 100 noisiest rules, each owned by a named manager.
- Report page-to-action ratio per team in the monthly reliability review.
8. On-call architecture and 24x7 coverage model across 28 teams (depends on: 3, 4)
This is the hardest political step. The answer to carrying a pager for another team's code is that every team carries its own, and the platform carries the shared risk.
- Adopt a federated model: every service has exactly one owning team, and that team's primary on-call carries its own pager. No team is paged for code it does not own.
- State the consequence honestly: 16 of 28 teams currently have no on-call. They must build one or formally transfer ownership of their services to a team that will.
- Merge small or low-traffic teams into shared rotations so no rotation has fewer than six engineers, below which coverage is not sustainable.
- Create a separate, centrally staffed Incident Commander on-call roster drawn from trained senior engineers across all teams, covering 24x7.
- Define primary and secondary per rotation, with the secondary engaged only on a no-acknowledge or an explicit request.
- Define coverage across the two AWS regions and New York business hours: one global IC rotation, service on-call aligned to their service's users.
- Define the unresponsive-team path: fifteen minutes without acknowledgement escalates to the team manager, thirty minutes to the IC, who may then direct any available engineer.
- Publish a coverage matrix of all 28 teams showing services, rotation size and gaps, reviewed monthly.
- Make on-call participation an explicit expectation in engineering job levels and hiring criteria.
9. On-call compensation, wellbeing and sustainability policy (depends on: 8)
Unpaid on-call is the most cited reason for resistance. The policy must be settled before rollout, not negotiated during it.
- Move to paid on-call: a per-shift stipend or salary uplift agreed with HR and Finance and benchmarked to the New York market.
- Pay event-based compensation for incident callouts outside business hours, with a minimum call-out block.
- Provide compensatory rest: no engineer works a normal day after a night incident, and the rest day is documented, not granted as a favour.
- Cap intrusion by defining a maximum number of off-hours pages per shift, with a mandatory review triggered whenever it is exceeded.
- Define a voluntary opt-out path for engineers with genuine constraints, balanced by an explicit obligation that someone else is paid to take the shift.
- Include on-call expectation and compensation in offers and job descriptions so the commitment is set before hiring.
- Publish the policy with an effective date before any team is asked to join a new rotation.
- Review the policy every six months against actual page volumes, attrition and survey results.
10. Internal and customer communications policy with timing SLAs (depends on: 3, 4)
Today the status page is written by whoever is around. This step replaces improvisation with a clock and a named owner.
- Set internal cadence: first update within 15 minutes of declaration, then every 30 minutes for SEV1 and 60 minutes for SEV2, regardless of whether there is progress.
- Set customer cadence: status page updated within 30 minutes of a SEV1 and 60 minutes of a SEV2, and a no-new-information update is still mandatory.
- Define the channel hierarchy: status page for everyone, direct email to affected customers on SEV1, named account-manager calls for the top 50 accounts.
- Prepare templates per severity in advance with legal and compliance pre-approval, covering detection, impact, workaround, mitigation and next-update time.
- Define regulatory obligations explicitly: money transmitter and banking regulator notification windows, security breach notification, and who signs off (Compliance, not Engineering).
- Prohibit speculation: customer communications never guess at cause or blame and never commit to a root cause before the postmortem.
- Define the closing communication: resolution notice, expected SLA credit handling and the committed date for a written report.
- Assign a named Customer Communications Lead per incident with a trained deputy on every SEV1.
11. Status page, notification tooling and account-manager playbook (depends on: 10)
Policy without tooling collapses at three in the morning. This step makes publishing a five-minute action.
- Upgrade or replace the status page so components map to customer journeys rather than internal services, with subscriber control per component.
- Integrate the incident tool with the status page so the incident record drives the update and the public timeline.
- Provide one-click templates pre-filled with severity, impact language and next-update time.
- Give account managers a playbook: contact tree, what they may say, what they must not say, and how to escalate a customer question into the incident channel within minutes.
- Define the SLA credit process end to end, covering computation, approval, customer notification and finance treatment, so credits stop being a manual scramble.
- Host the status page outside the production failure domain so it survives a total platform outage.
- Test publishing during game days, including a simulated status-page outage and a simulated loss of the primary region.
12. Postmortem policy, template and blameless review process (depends on: 3, 4)
Only 11 of 64 action items closed means the postmortem ritual is currently a writing exercise. This step rebuilds it around learning and tracking.
- Make postmortems mandatory for every SEV1 and SEV2, every SEV3 with customer impact or a repeat pattern, and any near-miss the IC flags.
- Set deadlines: draft within five business days, blameless review within ten, internal publication within fifteen.
- Adopt a single template: impact, timeline, detection, response, contributing factors, what went well, what went badly, action items.
- Train a pool of blameless facilitators and require a trained facilitator for every SEV1 review.
- Prohibit counterfactual and blame language in the template, and require contributing factors across tooling, process, organisation and human factors.
- Limit action items to a small number of concrete, verifiable, owned changes with dates.
- Publish all postmortems internally by default, with security review only for genuinely sensitive material.
- Produce a customer-facing root cause report variant for SEV1 incidents, especially those affecting regulated or top-tier accounts.
13. Corrective action tracking and reliability backlog governance (depends on: 12)
A postmortem without durable action tracking is a complaint, not a control.
- Create a single reliability backlog in the engineering tracker with a mandatory label, owner, due date and link to the originating incident.
- Define closure criteria that require evidence: a merged change, a tested alert or a verified drill, never a self-reported status change.
- Protect capacity by reserving a fixed percentage of each team's sprint for reliability work, with unspent capacity visible to vice presidents.
- Run a weekly ageing review of open actions and escalate anything overdue by more than 30 days to the VP Engineering.
- Report closure rate and median age monthly, targeting more than 90 percent closed within 60 days.
- Require a repeat incident in the same area to trigger a design review rather than another action item.
14. Incident tooling consolidation and integration (depends on: 3, 5, 7, 11)
Six alerting tools and no single incident record is a structural cause of the 22-minute detection and the three-hour mitigation.
- Select one incident management platform for paging, on-call schedules, escalation policies, incident records and postmortem workflow.
- Consolidate the six alerting sources into a single event pipeline feeding that platform, with deduplication and severity mapping applied at ingest.
- Integrate with platform and ledger observability so responders see dashboards and runbooks inside the incident record.
- Integrate chat and bridge: incident channel auto-created, timeline auto-captured, decisions logged as they happen.
- Define the data model and retention required for SOC 2 evidence: who did what, when, and under whose authority.
- Run a dual-run period alongside the old tools with a defined rollback, then switch off the legacy tools on a published date.
- Budget for licences, migration effort and a two-week hardening period after cutover.
15. Training, certification and exercise programme (depends on: 4, 5, 10, 12)
A process that exists only on a wiki page fails on the first real page. Skills have to be built and tested before they are needed.
- Build a short practical curriculum: how to be on call, how to declare an incident, how to run an incident as IC, how to communicate and how to write a postmortem.
- Require certification before joining the IC on-call roster: a written assessment plus a live simulated incident.
- Train at least two certified ICs per team group so the central roster has depth across all 28 teams.
- Run monthly tabletops on realistic scenarios drawn from the last 12 months, including region loss and ledger corruption.
- Run quarterly game days with deliberately injected failure, including shared-PostgreSQL failover and status-page outage.
- Add a mandatory onboarding module for every engineer joining or transferring in, with completion records suitable for audit.
- Track training completion by team and publish it in the monthly reliability review.
16. Metrics, dashboards and review cadence (depends on: 2, 3)
The programme needs a public scoreboard, or it will quietly rot after the audit.
- Define the outcome metrics: time to detect by source, time to mitigate, percentage of incidents detected by customers (target below ten), incidents by severity and SLA credits paid.
- Define the process metrics: declaration latency, page acknowledgement rate, IC roster coverage, first-update timeliness and update-cadence adherence.
- Define the health metrics: alert volume and noise ratio per team, off-hours pages per engineer, postmortem timeliness, action closure rate and action age.
- Publish live dashboards visible to every engineer, not only to managers, refreshed daily.
- Institute a weekly operational review of 30 minutes going incident by incident, a monthly reliability review of trends and actions, and a quarterly executive review with the CEO's office.
- Baseline every metric against the S2 evidence pack and set 90-day and 12-month targets.
- Require every review to end with decisions and owners, not just numbers.
17. Pilot with volunteer teams (depends on: 5, 7, 9, 11, 12, 13, 14, 15, 16)
Do not roll out to 28 teams untested. Run the entire process end to end with a small cohort first.
- Recruit three to four volunteer teams covering a mix of criticality: one payment-path team, one ledger-adjacent team, one platform team and one low-traffic team.
- Run the complete process in the pilot: new severity scale, roles, escalation, communications, postmortems, action tracking and paid on-call.
- Instrument the pilot against the S16 metrics and compare results with the S2 baseline.
- Hold weekly retrospectives with pilot teams and iterate on the written policies, the tooling and the training.
- Fix the top issues found before any wider rollout and document what changed and why.
- Produce a pilot report with before-and-after numbers to carry into every rollout conversation.
- Set explicit pilot exit criteria: rotation coverage achieved, no unacknowledged pages over a defined period, postmortems delivered on time and actions tracked.
18. Phased rollout to all 28 teams (depends on: 13, 16, 17)
Rollout is a staged migration with readiness gates, not an email announcement.
- Sequence the 28 teams into four waves of roughly seven teams, ordered by customer impact, with three weeks between waves.
- Define a per-team readiness checklist: services mapped and owned, alerts cleaned to standard, runbooks written, rotation staffed, training complete and manager briefed.
- Hold a gate review with the process owner before each team joins, and move unready teams to the next wave with a dated remediation plan.
- Give each wave a named champion and run an internal communications cadence that explains the why using pilot numbers.
- Handle resistance directly by publishing the own-your-code-own-your-pager rule and the paid on-call mechanics before each wave, not after.
- Retire legacy tools, escalation lists and the informal status-page process at the end of each wave on a published cutover date.
- Harvest feedback formally at each wave and push accepted changes back into the policy documents through change control.
19. SOC 2 incident-response control mapping and evidence framework (depends on: 1, 3, 10, 12)
The audit will test incident response. Mapping controls now means evidence accumulates from the first real incident under the new process.
- Map the process to the relevant Trust Services Criteria for incident identification, response, evaluation of incidents and communication of security events.
- Write control statements in auditor language and name a single owner for each control.
- Specify the evidence artifact for each control: incident record, severity classification, paging log, communications log, postmortem, action tracker entry and training record.
- Set evidence retention and storage location so nothing depends on a laptop or on chat history that expires.
- Run an early walkthrough with an experienced compliance partner or the auditor's readiness team to test the design before the audit window.
- Flag controls that are blocked until rollout reaches the relevant teams, and keep a gap register with owners and dates reviewed fortnightly by the steering group.
20. SOC 2 dry run, gap remediation and audit support (depends on: 18, 19)
This step converts a good process into a provable one, a few months before the auditors arrive.
- Run an internal dry run six weeks before the audit window, sampling real incidents from the pilot and early waves against each control's evidence requirements.
- Remediate every gap found, prioritising anything that risks a qualified opinion.
- Brief every engineer who may be interviewed so they can describe the process as they actually practise it.
- Prepare the auditor package: process documentation, sample incident records, training records, on-call schedules and the action tracking register.
- Designate a single audit liaison and a small evidence-request team so requests do not land on 28 teams at once.
- Rehearse the walkthrough with the incident commander and communications roles, since auditors probe realism under pressure.
- Schedule the dry run and the audit alongside rollout waves so no team is hit with both in the same week.
21. Standing governance, process ownership and continuous improvement (depends on: 18, 20)
The classic post-audit failure is that the process freezes and then decays. This step keeps it alive.
- Establish a standing Incident Management Council chaired by the process owner, meeting monthly with engineering, support, compliance and product representation.
- Give the process owner a documented mandate to change standards, and require formal change control for any change to severity, roles, communications timings or compensation.
- Re-validate the severity taxonomy quarterly against real declarations and re-baseline the metric set annually.
- Fund a resilience roadmap separate from incident response: reducing shared-database blast radius, multi-region failover readiness and deploy safety.
- Report quarterly to the executive team on the metric set and on the top five systemic causes of incidents.
- Feed incident themes into architecture review, release engineering and hiring so the programme improves the system, not just the response.
- Refresh the training and tabletop programme annually and after any SEV1.
--- PROPOSAL 3 (agent qwen3.8-flash_initial_3, alibaba/qwen3.8-flash) ---
Estimated complexity: high
Success metrics: - MTTD reduced to under 10 minutes with >90% detection occurring before customer reporting.
- MTTR for Sev1/Sev2 incidents reduced to under 60 minutes.
- SLA credits paid reduced by 50% within 12 months.
- On-call alert volume reduced by 80% (from 3,400 to <600 actionable alerts/month).
- 100% of Sev1/Sev2 incidents have a blameless postmortem completed within 5 days.
- >90% completion rate of postmortem action items within 30 days of assignment.
- Zero ambiguity in incident command: 100% of major incidents have a designated IC logged in the timeline.
- SOC 2 Type II audit passes incident response controls with zero critical findings.
Steps (12):
1. Establish Governance and Executive Sponsorship
Secure explicit CEO and CTO backing to mandate the program, framing it as a revenue protection and compliance necessity rather than an engineering burden. Form a cross-functional Incident Management Task Force.
- Identify a single executive owner (e.g., VP of Engineering or Head of Platform) with authority over budget and staffing.
- Create a charter defining the scope: severity definitions, on-call compensation, and tool consolidation.
- Communicate the strategic priority to all 28 team leads to preempt pushback regarding pager responsibility.
2. Define Severity Matrix and Response SLAs (depends on: 1)
Create a rigorous, binary severity scale to eliminate ambiguity in escalation and communication speed. Align internal response times with the 99.95% SLA.
- **Sev1 (Critical)**: System-wide outage, data corruption, or >5% transaction failure rate. Immediate executive page.
- **Sev2 (Major)**: Single customer segment impacted or >1% transaction failure rate. High-priority team response.
- **Sev3 (Minor)**: Degraded performance or isolated bug. Standard business-hours follow-up.
- **Sev4 (Cosmetic)**: No functional impact. Backlog item.
- Define maximum detection and mitigation times for each level.
3. Restructure On-Call Organization and Compensation (depends on: 1)
Transform the current unpaid, fragmented on-call system into a professional, compensated, and centralized rotation to ensure 24x7 coverage across all 28 teams. Address the 'pushback' by making it mandatory and paid.
- Consolidate the 12 current on-call teams into a unified, rotating pool covering all 28 teams.
- Implement a fair compensation model: base stipend + hourly rate for active hours + incident bonus for Sev1/2.
- Define a strict 'Buddy System' or 'Follow-the-Sun' rotation if geographically feasible, otherwise ensure adequate team size to rotate fairly.
- Mandate that every team must participate in the rotation; no exemptions.
4. Define Incident Roles and RACI Model (depends on: 2)
Eliminate the 'who is in charge' ambiguity by pre-defining roles and establishing a clear chain of command for every incident. Roles are assigned per incident, not permanently held.
- **Incident Commander (IC)**: Single point of authority. Owns the timeline, delegates tasks, and approves communications. Must be senior.
- **Scribe**: Maintains the real-time timeline log of actions, timestamps, and decisions.
- **Communications Lead**: Owns external/status page updates and internal Slack broadcasts. Shields the IC from interruptions.
- **Technical Leads (SMEs)**: Deployed by IC to troubleshoot specific services. They do *not* coordinate with each other; they report to IC.
- Create a 'Rapid Response Team' (RRT) for high-severity incidents involving the shared PostgreSQL cluster.
5. Consolidate Monitoring and Enforce Alert Hygiene (depends on: 3)
Reduce the 3,400 monthly alerts (85% noise) to a signal-to-noise ratio that allows engineers to trust the pager. Standardize tooling.
- Consolidate the six disparate alerting tools into a single platform (e.g., PagerDuty, Opsgenie, or Datadog).
- Implement 'Actionable Alerting': If an alert pages, there must be a documented, immediate runbook action. No 'informational' pages.
- Set a hard cap on alert volume per engineer/night. Breaching the cap triggers a mandatory alert quality review.
- Deprecate legacy tools that cannot map to the severity matrix defined in S2.
6. Build Detection and Escalation Automation (depends on: 5)
Reduce the 22-minute mean time to detect (MTTD) by automating the path from metric breach to page. Ensure customers never find an issue before the internal pager does.
- Implement synthetic transactions for core payment flows to detect failures before real users do.
- Configure escalation policies: If not acknowledged in 5 minutes, auto-escalate to the team lead, then to the Incident Commander pool.
- Map specific Kubernetes and PostgreSQL metrics to the Severity Matrix (S2) to trigger automated pages.
- Create a 'Major Incident Channel' in Slack that is automatically opened and pinned when a Sev1/2 is triggered.
7. Standardize Communication Templates and Timing (depends on: 4, 6)
Codify the rules for external and internal communications to stop the 'whoever is around' approach. Link these timings to the Severity Matrix.
- **Status Page Updates**:
- Sev1: First update within 5 minutes of detection. Updates every 15 minutes.
- Sev2: First update within 30 minutes. Updates every 60 minutes.
- **Customer Direct Mail**: For Sev1 impacting specific accounts, Account Managers send pre-approved templates within 1 hour of mitigation.
- **Regulator Notification**: Define legal triggers for NY financial regulations.
- Create a 'Comms War Room' Slack channel restricted to the IC and Comms Lead.
8. Formalize the Postmortem Process (depends on: 4, 7)
Transform postmortems from optional/various formats into a mandatory, blameless, and tracked process. This directly addresses the SOC 2 audit requirement for continuous improvement.
- Mandate a postmortem for every Sev1 and Sev2 incident within 5 business days.
- Enforce a 'Blameless' charter: Focus on systemic/process failures, not individual errors.
- Adopt a standardized template: Timeline, Root Cause Analysis (5 Whys or Fishbone), Impact Analysis, and Action Items.
- Action Items must have: Owner, Due Date, and Priority. They are tracked in Jira/Linear, not the doc.
9. Implement Action Item Tracking and Governance (depends on: 8)
Solve the '11 of 64 closed' problem by integrating postmortem actions into the core engineering workflow and executive reviews. Ensure fixes are shipped.
- Integrate Jira with the Postmortem tool (e.g., OpsLevel, Rootly, or Notion API). Auto-create tickets for every AI.
- Block 'Sev1 Closure' until all high-priority Action Items are closed or explicitly deferred with executive approval.
- Weekly Engineering Leadership review of open postmortem actions older than 30 days.
- Tie action item completion rates to team performance reviews.
10. Deploy Training and Runbook Library (depends on: 6, 7, 8)
Address the skill gap and the 'pushback' by equipping engineers with the tools and knowledge to succeed on-call. Reduce anxiety through clarity.
- Audit all 180 services. Require every service to have a 'Page-Worthy' runbook.
- Run 'Game Day' (Chaos Engineering) exercises quarterly to practice role assignments (IC, Scribe, Comms) in a simulated Sev1.
- Conduct mandatory training for the new on-call rotation and compensation model.
- Provide a 'Playbook' for the Incident Commander role, including scripts for communication.
11. Phased Rollout and Change Management (depends on: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10)
Introduce the new process in phases to manage cultural resistance and technical debt, ensuring adoption before the SOC 2 audit. Do not 'big bang' the change.
- **Month 1-2**: Launch Severity Matrix and Consolidated Tooling (S2, S5). Pilot on 5 high-traffic teams.
- **Month 3-4**: Roll out new On-Call Compensation and Roles (S3, S4) to all 28 teams. Enforce new Communication Templates.
- **Month 5-6**: Enforce strict Postmortem and Action Item Tracking (S8, S9).
- **Month 7**: Final audit readiness check. Simulate SOC 2 auditor questions on incident response.
12. Define Metrics for Success and Continuous Review (depends on: 2, 6, 8, 9)
Establish a dashboard to prove the new system works, linking directly to the $1.3M savings in SLA credits. Monitor the health of the incident process itself.
- Track MTTD (Target: <5 min), MTTR (Target: <30 min for Sev2), and Detection Ratio (Target: >95% internal detection).
- Monitor 'Alert Noise' (target: <5% of alerts result in a non-actionable page).
- Track 'Action Item Completion Rate' (Target: 100% on-time closure for Sev1/2).
- Calculate 'SLA Credit Avoidance' vs. On-Call Program Cost to prove ROI to leadership.
HOW THE ROUNDS WENT (from the round analyses):
Round 0: All three agents converge on the same skeleton — severity matrix → roles → paid on-call → tool consolidation → comms SLAs → blameless postmortems with tracked actions → metrics → phased rollout — but differ sharply in depth and in the on-call model. P2 is the most operationally concrete (baseline evidence pack, SLO-based detection, federated ownership, SOC 2 control mapping plus dry run); P1 is exhaustive but serializes training and drills after full rollout; P3 is the compact version and leaves audit evidence and detection engineering thin.
shared: **Four-tier severity scale as the keystone**: every plan keys paging, comms cadence and postmortem obligation off SEV1–SEV4 (P1 S1, P2 S3, P3 S2).
shared: Same role set — IC who commands but does not debug, comms lead, scribe, SME responders — with explicit decision rights (P1 S2, P2 S4, P3 S4).
shared: Collapse the six alerting tools into one platform and attack the 85% noise with actionable-alert standards, dedup and per-team caps (P1 S5–S6, P2 S7/S14, P3 S5).
shared: Paid on-call plus mandatory blameless postmortems for SEV1/SEV2 with action items tracked in Jira and reviewed by leadership; pilot first, then waves, not a big bang (P1 S4/S12/S13/S19, P2 S9/S12/S13/S18, P3 S3/S8/S9/S11).
differences: **On-call architecture**: P2 S8 is federated — one owning team per service, "own your code, own your pager", minimum rotation of six, 16 teams must build a rotation or transfer ownership, plus a central 24x7 IC roster. P3 S3 does the opposite, merging 12 rotations into one unified pool covering all 28 teams, which directly recreates the "pager for other teams' code" grievance it claims to solve. P1 S3 sits in between: dedicated IC pool of 4–6 plus per-team SME on-call.
differences: **Detection engineering**: only P2 S6 defines SLIs/SLOs for the top 20 customer journeys, symptom-based alerting, external synthetics in three locations, ledger-specific PostgreSQL signals (replication lag, TXID exhaustion) and a separate blast-radius resilience track. P3 S6 has synthetics only; P1 treats detection mostly as alert routing and never defines SLOs.
differences: **SOC 2 rigour**: P2 devotes S19–S20 to Trust Services Criteria mapping, a named owner and evidence artifact per control, retention rules and a dry run six weeks out. P1 S16 has a control mapping plus a month-6 mock audit. P3 has one bullet in S11 ("simulate auditor questions") — far too light for an eight-month Type II window.
differences: **Ordering flaws differ**: P1 S17 depends on all sixteen prior steps and pushes training (S18), rollout (S19) and the first drill (S20) to the end, so nobody practises before going live. P2 front-loads a 12-month baseline register (S2) that every metric later hangs on. P3 S11 defers comp and role rollout to months 3–4 while alert cleanup starts month 1, and P3 S9 ties action-item completion to performance reviews, which cuts against its own blameless charter (S8).
Proposal 1: 21 steps covering severity tiers, roles, a hybrid IC-pool/per-team SME rotation, concrete comp numbers ($500–1,000/week stipend, 1.5x callback, comp day), single alert tool with suppression rules, per-severity playbooks, status-page timings (SEV1 in 3 minutes), postmortem and action-item tracking, metrics and governance cadence. Rollout is a single mega-step (S17) depending on all sixteen predecessors, followed by training, waves and drills. Targets are the most aggressive of the three: MTTD <8 min, credits <$100k, 0 findings.
Proposal 2: Starts with a funded charter and a named process owner (S1), then a 12-month incident and alert baseline used as both problem statement and audit evidence (S2). Builds severity with a SEV0 for security/regulatory events, federated per-team on-call with a central IC roster, SLO- and synthetic-based detection, a 90-day noise sprint with a two-pages-per-shift budget, comms timing SLAs with pre-approved legal templates and a status page outside the failure domain, evidence-based action closure, then pilot, four gated waves and a SOC 2 dry run. Ends with a standing council and a separate resilience roadmap.
Proposal 3: Twelve steps: executive charter, severity matrix tied to transaction failure rates, a single mandatory paid on-call pool for all 28 teams, RACI roles plus a Rapid Response Team for the shared PostgreSQL cluster, tool consolidation with runbook-or-no-page hygiene, synthetic transactions and auto-escalation, tight status-page timings (SEV1 first update in 5 minutes), mandatory 5-day postmortems, Jira-integrated action tracking, game days, and a month-by-month rollout to month 7. Metrics target MTTD <5 min and 95% internal detection.
Round 1: P2's round-0 architecture became the de facto template: P1 and P3 both rebuilt their plans on it step-for-step, while P2 itself deepened its version with genuinely new mechanics (Triage Owner, severity x class, priced opt-out, three-action cap, regulator clock matrix, SOC 2 evidence clock). The round converged strongly on structure; what still separates the plans is depth of incident mechanics, realism of comms timings and where audit work sits in the sequence.
differences: **Command mechanics between alert and declaration.** P2 alone closes the gap that caused the two hour-long ownership failures: Triage Owner rule (step 5), Watch state with a 30-minute timer, "declaring is free", the ambiguity rule and the two-simultaneous-SEV-1 rule (step 6). P1 has an escalation ladder (step 9) but no owner-from-first-ack rule; P3 has no escalation or lifecycle step at all and dropped its round-0 5-minute auto-escalation.
differences: **On-call shape.** P2 builds three rotations including a paid Platform Duty for the shared PostgreSQL/Kubernetes estate (step 8) and prices opting out against a paid pool (step 9). P1 (steps 10–11) and P3 (step 5) stop at federated team rotations plus a central IC roster, leaving shared infrastructure ownership implicit.
differences: **Customer communication timings and money.** P1 demands a status-page update within 3 minutes and SEV-1 updates every 5 minutes (step 13) — contradicting its own success metric of 30 minutes. P2 uses 30/60 minutes (step 12) plus a regulator clock matrix naming NYDFS Part 500 and a customer-impact ledger driving SLA credit automation (step 13). P3 uses 15/30 minutes (step 9) with no credit process at all.
differences: **Where audit work sits.** P2 maps controls and defines the "golden incident file" in month one (step 3) on the argument that Type II evidence cannot be backfilled. P1 places control mapping at step 21, P3 at step 12 after postmortems. P3 also schedules game days (step 16) and the metrics dashboard (step 17) only after full rollout, so nothing is drilled or measured during the pilot.
influences: P1 and P3 rebuilt on P2's round-0 skeleton almost wholesale: charter (P2 s1), baseline evidence pack (s2), severity trigger matrix (s3), role cards (s4), SLO/synthetic detection (s6), "no runbook, no page" and the 90-day noise sprint (s7), federated on-call with a six-engineer floor (s8), paid on-call with compensatory rest (s9), pilot then four waves with readiness gates (s17, s18), control mapping and dry run (s19, s20).
influences: P1 kept only one structural idea of its own: severity playbooks with decision trees (its round-0 s9, now step 12); everything else in its 23 steps mirrors P2's ordering.
influences: P2 took P3's ROI framing (P3 s12, credit avoidance vs programme cost) into its step 13, and P1's mobile-pager requirement (P1 s5/s8) into step 11 ("run a SEV-1 from a phone at 3am").
influences: Nobody adopted P1's round-0 per-incident bonus (s4) — P2 explicitly bans pay attached to incident counts (step 9) — and nobody adopted P3's tying of action-item completion to performance reviews (s9), which P2 contradicts with a published amnesty.
influences: P1 carried over P2's round-0 metric of 40+ certified ICs, which P2 itself cut to 12–16 this round as more realistic for 260 engineers.
Proposal 1 (improved): P1 abandoned its generic round-0 framework and adopted P2's structure nearly step-for-step, gaining a charter, baseline evidence pack, SLO-based detection, a paging contract, a federated on-call model and a pilot-then-waves rollout. It kept its own useful severity playbooks step. Residual weaknesses are internal inconsistencies in timings and severity definitions.
Proposal 2 (improved): P2 kept its 21-step shape but added several mechanisms that close real gaps rather than restating policy: the SOC 2 evidence clock, severity x class, the Triage Owner rule, three rotations, a priced opt-out, an action-item cap and a regulator clock matrix. It also made its own targets more realistic.
Proposal 3 (improved): P3 grew from 12 to 20 steps by adopting P2's skeleton — charter, baseline, detection SLOs, compensation, control mapping, dry run, culture — which fills most of the prompt's requirements it previously skipped. It remains the thinnest plan in mechanics, dropped its own escalation automation, and its dependency order pushes drills and metrics past full rollout.
Round 2: P1 absorbed almost the entire distinctive vocabulary of P2's round-1 plan (Triage Owner, severity×class, three rotations, golden incident file, capped action items), producing two near-twin heavyweight plans; P2 added the genuinely new ideas of the round (a two-week minimum viable process, a funded ledger-resilience track, a cost-of-downtime model). P3 went the other way and collapsed into 22 bare titles with no content, losing everything that made it assessable.
differences: **Bridging the gap before tooling exists**: P2 S5 defines ten day-one rules, a manual duty-IC rotation drawn from the 12 teams that already have on-call, and a daily 15-minute stand-up for month one. P1 has no interim process between the charter (S1) and platform selection (S11); P3 has none either.
differences: **Prevention as a funded track**: P2 S23 is a standalone engineering roadmap with concrete bets (ledger read-only tripwire, connection-pool isolation, PITR restore tests with published timings, rollback on SLO burn). P1 keeps resilience as two bullets inside S9 and S22; P3 does not mention it.
differences: **Communication clocks**: P1 S13 demands a status-page update within 3 minutes of SEV-1 and internal updates every 5 minutes; P2 S13 sets 15 minutes internal first, 30 minutes to status page, with a mandatory no-news update. P1's 3-minute rule contradicts its own success metric of 30 minutes for 95% of SEV-1s.
differences: **Level of specification**: P1 and P2 give thresholds, timers, dollar ranges and dates throughout; P3 R2 gives only step titles and a one-line rationale each — no severity definitions, no timings, no compensation mechanics, no dates on any metric.
influences: P1 took nearly all of P2's round-1 signature ideas: Triage Owner (P2 S5→P1 S5/S6), severity×class with "class can raise, never lower" (P2 S4→P1 S4), the three-rotation model (P2 S8→P1 S7), priced opt-out and amnesty (P2 S9→P1 S8), golden incident file and month-one control mapping (P2 S1/S3→P1 S1/S3), three capped action items and the repeat-incident design review (P2 S15→P1 S16).
influences: P2 took P1's weekly synthetic-page testing of escalation ladders (P1 R1 S9→P2 S7) and P1's status-page-component-to-customer-journey mapping plus a named status-page owner (P1 R1 S14→P2 S14).
influences: P2 took P3's culture and change-management step (P3 R1 S18) and turned it into S24: on-the-spot correction of blame language, pager-fatigue monitoring, public recognition for deleted alerts.
influences: P3 took P2's "evidence clock" framing into its S1 title but nothing else of substance; it adopted no new mechanisms this round.
influences: Nobody adopted P3's error budgets triggering feature freezes (P3 R1 S8) or its rule that a SEV-1 cannot close until high-priority actions are done (P3 R1 S11) — P1 S16 and P2 S16 instead cap actions at three and track them in a separate reliability backlog.
Proposal 1 (improved): P1 rewrote itself around P2's round-1 mechanisms while keeping its own depth. New steps 4, 5, 7, 12, 14, 15, 20, 21 replace vaguer round-1 equivalents, and detection, escalation and postmortem policy are now far more concrete. A truncated step 11 and a few internal contradictions are the cost.
Proposal 2 (improved): P2 kept its round-1 architecture intact and added the three things it was missing: an immediately usable interim process, a funded prevention track, and an explicit culture step. Compliance sequencing, metrics and communications are essentially unchanged and were already strong.
Proposal 3 (worsened): P3 discarded all step content and submitted 22 titles with a single sentence of rationale each. Every operational detail it had in round 1 — severity thresholds, status-page timings, compensation mechanics, wave schedule, training curriculum — is gone, and several metrics were loosened or stripped of dates.
THE PROPOSALS OF THE FINAL ROUND (2), to assess:
--- PROPOSAL 1 (agent claudeHaiku4.5_refine_1, anthropic/claude-haiku-4-5) ---
Estimated complexity: high
Success metrics: - Median Time to Detect from 22 minutes to <5 minutes for internal detection and <10 minutes for customer journeys by month 6; <2% of customer-impacting incidents first detected by customers (vs. 40% currently) by month 9.
- Median Time to Mitigate for SEV-1 from 3h 10min to <45 minutes by month 9; SEV-2 to <90 minutes by month 9.
- Monthly alert volume from 3,400 to <400 with >95% signal-to-noise ratio by month 5; all alerts satisfy paging contract (symptom-based, owned, runbook-linked, severity-mapped).
- Zero incidents with command ambiguity lasting >15 minutes from month 2 onwards; 100% of SEV-1 and SEV-2 incidents have named IC logged in timeline within 5 minutes; 100% of Triage Owners assigned within acknowledgement.
- 100% of the 28 teams have a documented Service On-Call rotation or a signed service-ownership transfer by month 4; Platform Duty rotation and IC roster both staffed and live by month 3.
- On-call compensation policy published and effective by month 1; on-call satisfaction score ≥7/10 by month 6; zero on-call-attributed voluntary attrition by month 6.
- SLA credits paid from $1.3M annually to <$100K by month 12; credit avoidance (prevented credits) tracked and reported monthly.
- 100% of mandatory postmortems (SEV-0, SEV-1, SEV-2, and repeat incidents) published internally within 15 business days by month 4.
- Postmortem action item completion rate from 17% (11 of 64) to >90% within 60 days by month 6; median action age <30 days; zero repeat incidents caused by the same contributing factor without a design review.
- 100% of the 180 services have a named owner, a detection contract, and at least one symptom-based alert by month 6.
- Status-page first update published within 30 minutes for ≥95% of SEV-1 incidents by month 4.
- 100% of new engineers complete incident-response onboarding within 30 days of joining; IC certification includes written exam and live simulation; ≥2 certified ICs per team group; zero uncovered hours in 24x7 IC roster.
- Weekly operational review held in ≥90% of weeks; 12 of 12 monthly reliability reviews; 4 of 4 quarterly executive reviews; all reviews end with documented decisions and owners.
- SOC 2 Type II audit passes all incident-response controls (CC7.1–7.5, CC2.2–2.3, CC4.1, CC3.x) with zero findings by month 8.
- All pilot and rollout incidents captured with complete golden incident files (timeline, roles, communications, postmortem, actions, closure evidence) by month 3 onwards; audit dry run identifies zero critical gaps by month 7.
- All 28 teams transitioned to new process by month 5; all legacy alert tools decommissioned; single source of truth for incidents established and sustained.
Steps (22):
1. Executive charter, governance structure, and evidence clock
Turn the CEO email into a funded, authorized program with clear ownership and documented evidence collection for SOC 2, starting today.
- Appoint a Director of Incident Management reporting to CTO, with dotted line to COO (customer impact) and Head of Compliance (audit readiness).
- Publish a one-page charter: scope (all customer-impacting, payment-path, data-integrity incidents across 28 teams and 2 regions), decision rights (IC may freeze changes, override team preferences during incidents), and authority to mandate process participation.
- Secure annual budget for tooling, training, on-call compensation ($500–800K estimated), and resilience work. Connect funding to avoided SLA credits ($1.3M baseline).
- Establish standing Incident Management Steering Group: CTO, VP Engineering, Head of Support, Head of Compliance, one engineering manager per region. Meet monthly.
- **Start the SOC 2 evidence clock on day 1.** An audit in eight months means operating-period evidence begins now; design the process to capture evidence continuously, not retroactively.
- Make incident-response participation a documented performance expectation for all engineering managers and team leads.
- Publish timeline: working process in month 2, all 28 teams in month 5, audit-ready in month 7.
2. Baseline measurement, incident register, and evidence pack (depends on: 1)
Establish defensible baseline metrics and identify structural gaps that explain the 40% customer-detected rate and 22-minute detection time.
- Build a 12-month incident register with all 31 customer-impacting incidents: date, detection source, detection time, mitigation time, customer count, services involved, SLA credits paid, root cause class.
- Audit the current alert estate: total volume per tool, volume per team, volume per service, page-to-action ratio, top 50 noisiest rules, off-hours interruptions per engineer.
- Construct a **silent-failure register**: incidents with no internal alert fired at all. This explains the 40% customer-detected rate.
- Reconstruct the two command-ambiguity incidents minute by minute: exactly when did ownership become unclear, how long, what was the decision bottleneck.
- Survey on-call engineers (target >70% response): burden, fairness, pay expectations, escalation clarity, willingness to stay.
- Interview Support and Account Management: how do customers discover incidents, what do they complain about, how do they contact you.
- Publish the problem statement internally; retain all artifacts for SOC 2 audit evidence. This is the baseline against which all improvements are measured.
3. Control mapping and evidence architecture (depends on: 1, 2)
Design the process to generate SOC 2-compliant evidence automatically, from the first real incident, so the audit clock ticks in your favour.
- Map the new process to Trust Services Criteria CC7.1–7.5 (incident identification, response, evaluation, containment, communication), CC2.2–2.3 (authorization), CC4.1 (change management), CC3.x (information availability).
- For each control, write a one-paragraph plain-language statement, name a single owner, and specify the evidence artifact (incident record, timeline, communications log, postmortem, action tracker, training record).
- Define the **golden incident file**: one single-click export per incident containing severity, timeline, roles assigned, decisions made, communications sent, postmortem, and action items. This is the audit unit.
- Specify data retention, immutability, access control, and storage location (not laptops, not chat history that expires). Ensure evidence is searchable and organized by incident date.
- Keep a gap register with owners and dates; review fortnightly in the steering group. Identify which controls are blocked by incomplete rollout and when they unblock.
- Run an early design walkthrough with an experienced SOC 2 readiness partner inside month 1 to stress-test control design before building on it.
4. Severity and response class taxonomy (depends on: 2)
Define four severity levels and four response classes so every decision—paging, communications, postmortem, compensation—keys off a defensible rule, not a judgment call.
- **Severity by impact scope**: SEV-1 (total payment-path loss, data corruption, or >5% transaction failure for >5 min); SEV-2 (significant degradation or single region loss); SEV-3 (limited impact with workaround available); SEV-4 (internal issue or cosmetic); SEV-0 (reserved for security/regulatory/privacy events).
- **Response class** (orthogonal to severity): Availability, Performance, Data Integrity & Ledger, Security & Privacy. **Key rule: class can raise severity, never lower it.** A SEV-3 data-integrity incident gets SEV-1 response posture because integrity is not recoverable by moving faster.
- Automatic triggers: loss of one AWS region → SEV-1 or SEV-2 (class-dependent); ledger write failures → SEV-1; replication lag >10s → escalation review; payment success rate <99% for >5 min → SEV-1/2; missed settlement window → SEV-1; total external API unavailability → SEV-1.
- Who may declare: any engineer, Support, account manager (based on observed customer impact). Who may downgrade: IC only, after investigation.
- Map each level to SLA credit exposure and to customer-facing status-page state.
- Include worked examples from the last 12 months so all 28 teams recognize their own incidents in the taxonomy. Re-validate quarterly against real declarations.
5. Incident lifecycle, Triage Owner rule, and escalation policy (depends on: 4)
Eliminate the "nobody was in charge for over an hour" problem by assigning ownership the moment a page is acknowledged.
- Define lifecycle states with clear entry/exit criteria: Detected (alert fired) → Triaged (is this real and customer-impacting?) → Declared (severity assigned) → Mitigated (core issue resolved) → Resolved (all verifications done) → Postmortem (review scheduled) → Closed (action items tracked or dismissed).
- **Introduce the Triage Owner rule**: the person who acknowledges the page owns the incident until an IC is assigned or the incident is stood down. There is never an unowned gap between first page and declaration. Triage Owner's sole job: decide within 15 minutes whether this requires an IC or a direct stand-down.
- Set aggressive acknowledgement and declaration targets: page acknowledged in 5 min; triage decision (is this real?) in 15 min; severity declaration in 30 min for any customer-facing incident.
- Implement automatic escalation ladders with no human judgment required: if responder does not acknowledge in 5 min, escalate to service owner; if no ack in 10 min, escalate to team manager; if no ack in 15 min, escalate to IC on-call. Escalation is never criticized.
- Define unresponsive-team path: if a service's on-call is unreachable for 30 min, IC may direct any available engineer from any team to engage.
- For SEV-1 and SEV-2: change freeze until IC declares mitigation confirmed; IC unfreezes changes explicitly.
- Enforce one incident, one record. Timeline auto-captured from Slack channel and bridge; never written from memory later.
6. Incident roles, command structure, and decision rights (depends on: 5)
Define clear roles with one-page responsibility cards published and linked from every paging notification.
- **Incident Commander**: owns incident outcome, not the fix. Declares severity, decides escalation, approves all customer communications, freezes changes, calls responders, hands off in shifts. Non-technical ICs are acceptable; technical depth is not required.
- **Deputy IC**: assigned within 15 min of declaration; shadows IC; takes over if IC unavailable or after 4-hour shift on SEV-1. Maximum IC shift: 4 hours on SEV-1, 6 hours on SEV-2.
- **Triage Owner** (new role): owns incident from first page acknowledgement until IC takes over or stand-down decision is made. Required for all incidents.
- **Communications Lead**: owns internal Slack updates and status-page messaging; shields IC from customer contact and interruptions.
- **Scribe**: records real-time timeline with decisions, actions, and key timestamps; not responsible for fixing.
- **Subject-Matter Responders**: engineers with service context; take IC direction; report only to IC; no side channels or parallel debugging.
- **Operations Lead** (SEV-1 only): coordinates multiple responders, manages incident bridge, maintains escalation list.
- Minimum viable staffing: SEV-1 requires all roles; SEV-2 requires IC, Deputy, Comms, Scribe, SMEs; SEV-3 requires Triage Owner and IC.
- Create laminated role cards for every on-call shift location (office, home, printed in pockets).
7. Three on-call rotations: Service, Platform, and Incident Commander (depends on: 5, 6)
Directly address the "carrying a pager for another team's code" objection by making it structurally impossible.
- **Service On-Call rotation** (federated): each of the 28 teams maintains a rotation for their own services only. No engineer is paged for code their team does not own. The answer to "why am I carrying a pager?" is now simply: "for your team's code."
- **Platform Duty rotation** (centrally staffed): shared PostgreSQL cluster, Kubernetes, networking, CI/CD, observability, and incident management tooling. Nobody's product code, so it gets its own dedicated rotation. Staffed from platform teams plus volunteers from other teams; paid at premium rate.
- **Incident Commander roster** (24x7): 12–16 certified senior engineers from across all 28 teams, on one-week primary shifts with secondary backup. Covers every hour with no single point of failure and no uncovered holiday week.
- **Consequences and gates**: 16 of 28 teams have no on-call today. Each must either (a) build a Service On-Call rotation of at least 6 engineers, or (b) formally transfer service ownership to a team that will, with transfer documented and dated. No exceptions, no waivers. Unowned services are decommissioned or transferred by end of month 2.
- Merge small or low-traffic teams into shared rotations where service ownership is unclear (e.g., shared analytics, testing infrastructure).
- Enforce scheduling limits in the tooling: no engineer on-call more than 2 weeks per quarter, automatically enforced by configuration, not negotiation.
- Publish a coverage matrix for all 28 teams showing services, owners, rotation size, gaps, and monthly status.
8. On-call compensation, rest policy, and sustainability (depends on: 7)
Settle compensation before rollout, not during negotiations. Make on-call sustainable and valued.
- **Paid on-call**: effective immediately upon joining a rotation. Weekly stipend while on shift (benchmark to New York market: $600–1,000 per week per engineer), regardless of incident volume.
- **Event-based compensation**: 1.5× hourly rate for time spent mitigating out-of-hours incidents, minimum one-hour block per callout. Tracked by incident record (auto-capture from timeline).
- **Compensatory rest**: no engineer works a normal 8-hour business day after a night incident requiring >2 hours mitigation. Rest day is documented policy, not a favour granted by manager.
- **Intrusion cap**: maximum 3 unscheduled pages per week per engineer. Exceed the cap in a week and trigger an immediate review; exceed in a month and escalate to VP Engineering. Breaches are structural signal that alert quality or service stability has a problem.
- **Voluntary opt-out**: an engineer may exit a rotation; their team must hire or buy replacement coverage from paid pool at published internal rate ($X per shift). This converts culture debate into visible budget decision.
- **Amnesty policy**: incident records, near-miss reports, and false declarations are never used in performance reviews or compensation discussion. Only failure to report is a performance issue.
- **Policy publication**: publish compensation structure and effective date before any team is asked to join a rotation, and include on-call expectations in job descriptions and hiring conversations.
- **Semi-annual review**: reassess compensation and caps every six months against actual page volumes, attrition rates, and survey feedback.
9. Detection strategy: SLOs, synthetic monitoring, and customer-report intake (depends on: 4, 7)
Close the 40% customer-detected gap by monitoring customer journeys instead of infrastructure metrics.
- **SLO-based alerting**: Define SLIs and SLOs for the top 20 customer journeys (payment initiation, authorization, settlement, ledger read/write, API availability, webhook delivery, payout). Measure per region. Alert on SLO breach, not on infrastructure metric (e.g., alert on "payment success rate <99%" not "database CPU >80%").
- **Synthetic transaction monitoring**: deploy synthetic transactions from outside AWS in both regions plus a third geographic location, one-minute cadence, for all money-moving paths. These are your first alarm bell.
- **Ledger-critical signals**: PostgreSQL replication lag (target: <1s, alert >5s), connection saturation, write latency (p95), lock-wait time, transaction ID exhaustion proximity, checkpoint pressure, table bloat. These are separate alerts on shared-database health.
- **Customer-report intake** (new detection channel): Support and Account Managers can raise an incident directly in the platform. Every customer report creates an incident record automatically, and the "customer report" detection source is counted in all metrics. This is a legitimate detection method, not a failure.
- **Detection-gap rule**: whenever a customer reports an incident before internal monitoring fires, auto-create a ticket in the owning service's backlog with root cause: "Monitoring gap on [journey]."
- **Detection contract per service**: every one of the 180 services needs a named owner, at least one symptom-based alert mapped to a SLO, and a documented expected detect time (target: <5 min for payment path, <10 min for others). Published on wiki and reviewed monthly.
- **Detection drills**: run a quarterly drill per team: simulate a broken service in staging and verify it triggers a page before a human notices.
- **Resilience roadmap separation**: detection improvements do not protect against ledger corruption or multi-region failure. Fund a separate resilience roadmap to reduce shared-database blast radius and improve failover safety.
10. Alert quality standards and noise-reduction program (depends on: 9)
Cut the 3,400 monthly alerts (85% noise) to <600 with 95% signal. This is the price of admission for on-call buy-in.
- **Paging contract**: every page must satisfy all of (1) symptom-based (customer impact, not infrastructure cause), (2) actionable (linked runbook with immediate next step), (3) owned (named team responsible), (4) severity-mapped (SEV-1/2/3/4), (5) SLO-linked where applicable. **No runbook, no page.** Enforce with CI check on alert definition.
- **Separation rule**: only customer-impacting or imminent-impact signals trigger pages. Everything else becomes a dashboard entry, a ticket, or a log line. Noisy infrastructure metrics go to dashboards, not pagers.
- **Page budget per service**: no service may exceed 2 pages per on-call shift per month. Exceeding budget auto-opens a remediation ticket in the owning team's backlog (with alert-quality review assigned to tech lead).
- **Automatic suppression rules**: (1) silence alerts if service auto-recovered within 30s, (2) suppress known maintenance windows, (3) group flapping alerts (>5 in 2 min) into one page, (4) rate-limit noisy services (max 1 page per 5 min until condition clears). All suppression rules must have an expiry date; no permanent silence without a ticket.
- **Probation for new alerts**: new alert rules run as tickets only and alert to a Slack channel; after two weeks of proving actionability (every alert resulted in human action), they graduate to pager.
- **Noise sprint**: run a focused 90-day program with a public burn-down of the top 100 noisiest rules. Assign each to a named manager. Default action: fix root cause, tune threshold, or delete within 10 working days. Deletion is a legitimate successful outcome (celebrate it).
- **Correlation and deduplication**: consolidate alert sources at ingest pipeline so one outage triggering 40 alerts produces one page, not 40.
- **Alert ownership**: every alert must have an owning team and a maintenance contact. Update monthly.
11. Incident tooling consolidation and integration (depends on: 5, 10)
Replace six alert tools and ad-hoc incident records with a single source of truth that unifies paging, escalation, timeline, and audit evidence.
- **Tool selection**: choose an incident-management platform (e.g., PagerDuty, Incident.io, Opsgenie) that integrates paging schedules, escalation policies, incident records, postmortem workflow, and status-page APIs. Decision gate: month 1.
- **Event pipeline consolidation**: route all alerts from the six legacy tools into a single event pipeline that feeds the incident platform. Apply deduplication, correlation, severity/class mapping, and rate-limiting at ingest.
- **Observability integration**: connect the incident platform to your Kubernetes dashboards, PostgreSQL monitoring, distributed tracing, and logs so responders see context in one pane. Link runbooks directly into incident records.
- **Slack and bridge integration**: auto-create incident Slack channels, auto-invite roles, auto-capture timeline from channel transcript and voice-bridge recording. Timeline is not written from memory; it is auto-captured.
- **Golden incident file**: implement the export defined in S3. One click produces a complete, immutable, audit-ready PDF: severity, timeline, roles, decisions, communications, postmortem, action items, and closure evidence.
- **Dual-run period**: run both legacy and new platform in parallel for two weeks. Define rollback criteria (e.g.,
12. Escalation automation and incident lifecycle enforcement (depends on: 5, 11)
Eliminate judgment calls from the worst moments. Escalation is automatic, mechanical, and blameless.
- **Automatic escalation ladders**: page responder → if no ack in 5 min, page service owner → if no ack in 10 min, page team manager → if no ack in 15 min, page IC on-call + call them immediately (phone + SMS + Slack). No human decides to escalate; timers fire escalations.
- **Severity-based escalation tempo**: SEV-1 uses faster timers (2 min for IC on-call), SEV-2 uses moderate timers (5–10 min), SEV-3 uses slower timers (15–30 min). Configured in tooling, reviewed quarterly.
- **Dual IC rule**: if a second SEV-1 incident is detected while the first is active, immediately page and assign a separate IC. ICs never run two incidents in parallel.
- **Change freeze and rollback authority**: SEV-1 and SEV-2 trigger automatic deploy freeze. Only the IC (with CTO/VP Eng notification) may unfreeze. Freeze lifts only when IC explicitly declares mitigation confirmed and verifies no new incident symptoms for 5 min.
- **Unresponsive team escalation**: if service's on-call does not acknowledge in 30 min, IC may direct any engineer from any team (volunteers first, then rotated) to engage. This is documented and reported in monthly review (escalation = signal of rotation problem).
- **One incident, one record**: all decisions logged in the incident platform. Auto-capture from Slack, bridge, status-page updates. Timeline is the source of truth; postmortem is written from timeline, never constructed after the fact.
- **Ambiguity rule**: if two responders disagree about whether an incident should be declared, it is declared. False declarations (stand-downs within 30 min of declaration) are tracked as metrics and closed without blame.
- **Watch state**: an unconfirmed incident can live in "Watch" state for max 30 min; after that, either declare it or stand it down explicitly.
13. Internal, customer, and regulatory communications workflows (depends on: 6, 12)
Define who informs whom, in what order, via what channel, with explicit timings and pre-approved templates.
- **Internal cadence**: first update to #incidents Slack channel within 3 min of declaration (even if "Investigating"). Then updates every 5 min (SEV-1), 15 min (SEV-2), or 30 min (SEV-3), or immediately on material change (e.g., mitigation achieved, scope widened). **Comms Lead owns the update; IC must not be interrupted.**
- **Executive notification**: IC calls CTO and VP Eng within 1 min of SEV-1 declaration (not email, not Slack, call). Incident declared in Slack with severity label, IC name, and affected service. Escalation channel lead auto-pinged.
- **Customer communication channels**: status page (all 2,100 customers), direct email to affected customers (top-tier accounts and customers affected by SEV-1), account-manager calls (top 50 accounts on SEV-1).
- **Status page timings**: update within 3 min of SEV-1 declaration, 10 min of SEV-2, 30 min of SEV-3 (even if root cause unknown; use "Investigating" with next-update ETA). Updates every 5–30 min depending on severity. Always include next-update time.
- **Pre-approved templates**: draft customer-facing language for each severity and class in advance with Legal and Compliance. Templates specify impact language ("some of your transactions are delayed" not "our database failed"), workarounds if available, and next-update commitment. Never speculate on cause in customer communication.
- **Regulatory notification path**: identify incidents requiring regulator notification (NYDFS Part 500, money-transmitter rules, payment-card-network rules, securities disclosure). Build a clock matrix: event type → regulator → notification window → signer. Compliance owns all regulatory notifications (never Engineering). Pre-clear templates. Flag incidents to Compliance immediately upon declaration.
- **Account manager playbook**: contact tree for top 50 accounts, templated talking points (facts only, never speculation), escalation path if customer escalates, what to offer (service credit, technical deep-dive call).
- **Closing communication**: resolution notice, SLA credit impact, commitment date for written root-cause report, customer action required (none, or security update, etc.).
14. Status page infrastructure and customer-impact ledger (depends on: 13)
Make the status page reliable, customer-centric, and audit-ready. Track customer impact in a single durable record.
- **Status page decoupling**: host status page outside production failure domain (separate cloud, separate infrastructure, separate database). Integrate incident platform with status page so incident record drives all public updates. Status page survives total platform outage.
- **Component-to-journey mapping**: status page components map to customer journeys ("Payments", "Settlements", "Payouts", "Ledger API") not to internal services. Allow customers to subscribe to components; notify by email or webhook.
- **One-click update templates**: pre-fill status-page template with severity, impact language, next-update time, and estimated resolution. Comms Lead types minimal new info ("Root cause identified" or "Workaround available"), and updates auto-post.
- **Customer-impact ledger** (one record per incident): which customer accounts affected, which journey(s) impacted, exact start and end time of impact, estimated SLA-credit exposure. Use this single record for customer communications, credit computation, regulatory reporting, and annual review. No reconciliation of two versions of the same outage.
- **SLA credit automation**: compute credit based on duration × severity × customer tier → auto-generate customer notification → auto-post to finance system. Reconcile accrued vs. paid credits monthly and report in executive review.
- **Testing during game days**: simulate status-page outage and verify alerts continue to fire; test total region loss and confirm status page remains updated; drill runbook for manually updating status page if platform is down.
15. Postmortem policy: mandatory, blameless, three-level framework (depends on: 6, 13)
Turn postmortems from a writing exercise (11 of 64 action items closed) into the learning engine of the system.
- **Mandatory postmortems**: all SEV-0, SEV-1, and SEV-2 incidents; all SEV-3 with customer impact or repeat pattern; any near-miss IC flags; any incident where the process itself failed (IC unreachable, Comms Lead unavailable, false declaration, missed update SLA).
- **Three-level framework** (proportionate to weight): (1) lightweight async review for SEV-4 and low-impact SEV-3 (10 min template in shared doc, owner + IC review), (2) standard facilitated postmortem for SEV-2 and impactful SEV-3 (full template, facilitated by trained neutral party, published within 10 days), (3) full executive postmortem for every SEV-1 and every security incident (executive sponsor assigned, full investigation, published within 15 days, customer-facing variant prepared).
- **Fixed timeline**: draft postmortem within 5 business days, blameless review within 10 days, internal publication within 15 days.
- **Single template**: impact (who, how many, how long, financial exposure), timeline (detection through resolution), root cause (not "human error" but system condition that enabled error; what was the gap?), contributing factors (tooling, process, organization, knowledge, monitoring), what went well, what went badly, action items (≤3, rest go to reliability backlog).
- **Blameless facilitation**: train a pool of blameless postmortem facilitators (target: 10+ engineers). Require a trained, neutral facilitator for every SEV-1 and SEV-2 review. Prohibit counterfactual language ("if the engineer had"), blame language, and the phrase "human error" as a root cause.
- **Publication rule**: publish all postmortems internally by default; security review only for genuinely sensitive material (e.g., unpatched vulnerability details or customer PII in logs). Create a customer-facing root-cause report for every SEV-1, especially for regulated customers, with legal and compliance sign-off.
- **Searchability**: store postmortems in a searchable wiki or issue tracker with tags (service, class, root cause category) so teams can learn from similar incidents without repeating them.
16. Action item tracking, reliability backlog, and repeat-incident design rule (depends on: 15)
Close the loop on incident learning by enforcing verifiable, tracked action items and breaking cycles of repeat incidents.
- **Action item capping**: each postmortem generates a maximum of 3 action items. Anything beyond 3 goes into a ranked reliability backlog, not into the postmortem, to prevent overwhelming teams.
- **Action item requirements**: each item must have (1) a named human owner (not a team), (2) a due date (≤60 days, target ≤30 days), (3) a definition of done (merged code change, tested alert, audit evidence, architectural decision, new runbook, training completed) not self-reported status.
- **Single reliability backlog**: create one backlog in your engineering tracker (Jira, Linear, etc.) with mandatory label (e.g., `incident-action`), link to originating incident, and link to postmortem. Track progress weekly.
- **Closure sign-off**: Incident Commander or postmortem facilitator must sign off on closure, verifying artifact exists (code merged, alert tested in drill, runbook verified).
- **Repeat-incident rule**: if the same service or component has a second incident with the same contributing factor, **do not create another action item**. Instead, escalate immediately to an architect or tech lead and trigger a design review (not a task, a review). This breaks the cycle of repeated patches; the system needs a structure change.
- **Capacity protection**: reserve a fixed percentage of each team's sprint capacity (10–15%) for reliability work. Track unspent capacity and report to VP Engineering monthly; if a team is not spending it, work with them to identify and fix blockers.
- **Ageing and escalation**: run a weekly review of open actions; escalate anything >30 days overdue to team lead and VP Engineering. Monthly report: completion rate (target >90% within 60 days) and median action age (target <30 days).
17. Training, certification, and exercise program (depends on: 6, 13, 15, 16)
Build skills before deploying the process. Run ongoing drills so the system is tested, not guessed at.
- **Curriculum**: (1) All engineers (30-min async video): severity taxonomy, communication expectations, postmortem format, when to declare an incident, where to find runbooks. (2) On-call responders (2-hr workshop): alert tool walkthrough, playbooks by severity, escalation paths and timers, when to call manager, mobile app walkthrough, case studies from the last 12 months. (3) Incident Commanders (4-hr bootcamp + test): leadership under pressure, decision-making (severity, escalation, rollback), communicating with executives, status-page discipline, postmortem facilitation, handling ambiguity, live simulated incident (pass/fail certification). (4) Communications Leads (2-hr training): templates per severity and class, customer-communication rules (no speculation, no blame), update timings, how to shield IC, regulatory triggers.
- **IC certification**: written assessment (75% pass required) plus live simulated incident (role-play with facilitator, graded on severity declaration, escalation decisions, communication, handover). Certification valid for 12 months; recertify via annual refresher or another live sim.
- **Depth across teams**: certify at least 2 ICs per team or team group so central roster is not siloed in one group; no holiday week is uncovered.
- **Async content**: record all training videos so async teams can learn on their schedule. Create quick-reference cards (laminated, pocket-sized) for roles and playbooks; distribute to on-call locations (office, home).
- **Monthly tabletop exercises**: drawn from real incidents from the last 12 months (region loss, ledger write failure, missed settlement window, cascading failures). Facilitator describes scenario; 3–4 responders play out response (Triage Owner, IC, Comms) as if real. Run 30 min; retro for 15 min afterward.
- **Quarterly game days**: deliberately inject failures into production (database failover, status-page outage, alerting-pipeline outage, dual SEV-1 incidents). All on-call roles engage. Run 2–3 hours; measure response times, decision quality, and communication. Document findings and create action items for identified gaps.
- **Drill the process's own failure modes**: IC unreachable (on-call unavailable, phone broken), Comms Lead on PTO, two simultaneous SEV-1s, paging storm (100+ alerts), false alarm that consumes an hour. Test escalation paths, deputy takeover, and recovery.
- **New-engineer onboarding**: add incident-response module to all engineering onboarding (completion tracked, audit-ready). All engineers must complete within 30 days of joining or transferring in.
18. Metrics, dashboards, and review cadence (depends on: 2, 12, 16, 17)
Measure to prove the system works. Publish live dashboards so every engineer sees the scoreboard and the system is transparent.
- **Outcome metrics**: Median Time to Detect by source (target: <5 min internally detected, <10 min customer journeys); Median Time to Mitigate for SEV-1/2 (target: <60 min SEV-1); customer-detected incidents as % of total (target: <5%); incidents by severity (should be mostly SEV-3/4, few SEV-1); SLA credits paid (target: <$100K/year by month 12); annual credit avoidance vs. program cost.
- **Process metrics**: IC assigned within 5 min (target: >95% of incidents); page acknowledgement rate (target: >98% within 5 min); first-update timeliness (target: >95% within SLA); postmortem timeliness (target: 100% of mandatory postmortems published on time); IC roster coverage (zero uncovered hours, monitored weekly).
- **Health metrics**: alert volume and signal-to-noise ratio per team (trending toward target); off-hours pages per engineer per month (trend, cap enforcement); on-call satisfaction survey (target: >7/10); training completion by team (target: 100% within 30 days); % of services with active detection contract (target: 100%).
- **Never publish incident count as a team metric.** Reward hiding. Instead publish detection metrics (near-misses reported per team, detection gaps closed, false declarations made).
- **Live dashboards**: build dashboards visible to all engineers (not just managers) showing outcome, process, and health metrics. Auto-populate from incident platform and alert tool. Update daily. Link from Slack and internal wiki.
- **Baseline all metrics against S2 evidence pack.** Set 90-day and 12-month targets for each metric. Publish targets and progress monthly.
- **Review cadence**: (1) weekly 30-min operational review (incident by incident from prior week: what went well, what hurt, actions); (2) monthly 60-min reliability review (trends, top causes, action-item aging, alert quality per team); (3) quarterly 60-min executive review (CEO's office: customer impact, SLA credits, top five systemic causes, program ROI).
- **Quarterly process review**: what in the process wasted responder time, what confused people, what should be deleted. Solicit feedback from ICs, Comms Leads, and responders. Document changes and reasoning.
19. Pilot program with 3–4 volunteer teams (depends on: 6, 8, 11, 12, 13, 14, 15, 16, 17, 18)
Do not roll out untested to 28 teams. Run the entire process end-to-end with a small cohort using real incidents as the primary training material.
- **Team selection**: recruit 3–4 volunteers spanning criticality: one payment-path team, one ledger-adjacent team, one shared infrastructure team (platform or Kubernetes), one low-traffic team. Volunteers see early adoption and influence.
- **Full process in pilot**: new severity and class taxonomy (S4), consolidated tooling (S11), roles and Triage Owner (S5–6), three rotations (S7), escalation automation (S12), communications (S13–14), postmortems (S15), action tracking (S16), paid on-call (S8), training (S17), metrics (S18). This is not a partial test; it is the complete system.
- **Real incidents are the training**: hold a retro within 48 hours of each pilot incident (while memory is fresh). Process Owner facilitates. Discuss: what worked, what hurt, how is the runbook, is the alert tuned, did Comms template work, did roles work, was timeline auto-captured correctly. Document feedback and iterate.
- **Instrument against baseline**: measure all S18 metrics in pilot; compare results with S2 baseline. If MTTD dropped from 22 min to 8 min, that is the number you carry into the next conversation.
- **Weekly retros with pilot teams**: structured 30-min sessions reviewing policies, tooling, training, templates, and process changes. Change the written policy or tooling in response; document what changed and why so later teams see the evolution, not just the finished state.
- **Explicit exit criteria**: (1) rotation coverage sustained (≥6 engineers per rotation, zero unacknowledged pages over 3 weeks), (2) postmortems delivered on time (100% of mandatory postmortems published within 15 days), (3) action tracking established (100% of action items in backlog with owner and date), (4) metrics live (dashboards updated daily, first weekly review completed).
- **Pilot report**: document before/after numbers (MTTD, MTTR, alert noise, action completion rate, on-call satisfaction) and key process learnings. This report is the foundation for every conversation in the rollout.
20. Phased rollout sequenced by cost of failure (depends on: 16, 18, 19)
Rollout is a staged migration with readiness gates and sequencing that prioritizes visible impact, not ease.
- **Wave sequencing**: divide 28 teams into 4 waves of ~7 teams each, **ordered by incident density and customer-journey ownership** (highest-cost-of-failure teams first). Teams with the most SLA credits at stake go first; their improvement is the proof.
- **Wave spacing**: three weeks between waves. This gives each wave time to stabilize and find problems before the next cohort joins.
- **Readiness checklist per team**: (1) all services mapped and owned (no unowned services), (2) alerts cleaned to paging contract (runbook linked, severity mapped), (3) playbooks updated and tested in staging, (4) rotation staffed to ≥6 engineers, (5) team completes training module, (6) manager briefed on policy, (7) on-call compensation in effect.
- **Gate review before each wave**: Process Owner holds gate review with target teams. Move unready teams to next wave with a dated remediation plan. No exceptions, no waivers; readiness is non-negotiable.
- **Wave champion**: assign a named engineer per wave to champion the rollout, answer questions, escalate issues to Process Owner. Champions are not representatives; they are advocates and feedback collectors.
- **Communication cadence**: weekly all-hands or newsletter for 4 weeks before each wave. Explain why (owned-code-owned-pager rule, paid on-call, SLA credit savings). Use pilot numbers. Answer FAQs. Announce champion and escalation path.
- **First incident under new process**: hold a retro within 48 hours. Feed accepted process changes back through change control.
- **Retire legacy tools and processes**: at end of each wave, retire legacy alert tools, informal escalation lists, ad-hoc status-page process. No parallel processes running for >3 weeks; this prevents confusion and half-learning.
- **Sequence to avoid audit collision**: ensure no team is rolling out in the same week as the audit dry run (S21).
21. SOC 2 dry run and evidence review (depends on: 3, 20)
Convert a good working process into a provable one. Test control evidence a few months before auditors arrive, when you can still fix gaps.
- **Dry run timing**: run 6 weeks before audit window (around month 7 of this program).
- **Scope**: sample 10–15 real incidents from pilot and early rollout waves. For each incident, verify evidence artifact exists and is complete: incident record, timeline (auto-captured), severity and class declaration, roles assigned and logged, communications log (Slack + status page), postmortem (if mandatory), action items in tracker with due dates, action completion evidence (code, alert test, training record).
- **Control walkthrough**: walk through each control statement from S3 with a checklist. Is the evidence artifact present? Is it immutable? Is it searchable? Is retention adequate? Is access logged?
- **Gap remediation**: for every gap found, estimate time to fix and prioritize by audit risk. Anything risking a qualified opinion (e.g., missing postmortem, no timeline evidence) must be fixed before the audit. Test the remediation against a new incident or a resample.
- **Interview readiness**: brief 10–15 engineers who may be interviewed by auditors (ICs, Comms Leads, Process Owner, team managers). Ask them to describe the process as they actually practice it, not as written. Listen for confusion or gaps in understanding. Correct them.
- **Auditor package preparation**: assemble process documentation, sample incident records (5–10 complete golden files), training records, on-call schedules, alert quality metrics, action tracker register, and status-page archive. Organize by control. Create a table of contents and index.
- **Single audit liaison**: designate Process Owner or a small dedicated compliance person as sole point of contact for audit requests. Prevents requests scattering across 28 teams.
- **Rehearsal**: conduct mock interview with an IC and a Comms Lead. Auditors ask tough questions under pressure ("How do you know the timeline is accurate?", "What happens when both ICs are unavailable?", "Show me how you proved the alert was actionable."). Practice answering.
22. Standing governance and continuous improvement (depends on: 20, 21)
Lock in durable improvement. The classic post-audit failure is the process freezing and then decaying. This step prevents that.
- **Standing Incident Management Council**: chaired by Process Owner, monthly meetings, attendees: CTO, VP Engineering, Head of Support, Head of Compliance, one engineering manager per region, one IC, one team manager from recent wave. Agenda: metrics review, policy changes, gaps from recent incidents, escalation for contentious issues.
- **Change control mandate**: give Process Owner documented authority to change severity taxonomy, response classes, roles, communications timings, and compensation policy. Any change requires: written justification, steering group approval (monthly cadence), and documented effective date before implementation. This prevents silent drift and ensures changes are deliberate.
- **Quarterly validation**: re-validate severity and response class taxonomy against all incidents from the prior quarter. Ask: did our taxonomy correctly predict response posture? Did we misclassify? Update taxonomy if patterns emerge.
- **Annual metric re-baselining**: every 12 months, re-run measurements from S2 (alert census, incident register) to reset targets. System should improve; targets should tighten.
- **Resilience roadmap separation**: fund a distinct architectural or platform roadmap for incident prevention (reduce shared-database blast radius, multi-region failover, deploy safety, observability investments). Better incident response does not protect a single-ledger corruption or unplanned failover. These are separate problems.
- **Quarterly executive report**: CTO and VP Eng report to CEO/CFO on metric set (detection time, mitigation time, SLA credits, customer-detected %), top five systemic causes of incidents, program cost vs. credit avoidance, and strategic architecture changes in flight.
- **Public backlog of improvement ideas**: teams and engineers propose process improvements via Slack or wiki. Process Owner reviews quarterly and implements accepted ideas (e.g., "add a dashboard for detection gaps", "update postmortem template"). Publish what changed and why.
- **Celebration and learning**: share wins publicly each quarter ("We reduced MTTD from 22 min to 5 min", "Customer-detected incidents down 80%", "$800K SLA credits avoided"). Refresh training and tabletop program annually and immediately after any SEV-1 to keep the system sharp and responsive to new scenarios.
--- PROPOSAL 2 (agent deepseek-flash_refine_2, deepseek/deepseek-flash) ---
Estimated complexity: high
Success metrics: - Median time to detect falls from 22 minutes to under 8 minutes by month 6, and under 5 minutes for money-moving journeys by month 9.
- Share of customer-impacting incidents first detected by customers falls from 40% to below 10% by month 9.
- Median time to mitigate for SEV1 falls from 3 h 10 min to under 60 minutes by month 9.
- 100% of SEV1 and SEV2 incidents have a named Triage Owner within 5 minutes and a named Incident Commander within 15 minutes, from month 4.
- Zero incidents with unclear command authority lasting more than 15 minutes, measured monthly from month 4.
- Monthly page volume falls from 3,400 to under 600, with a false-positive rate below 15%, by month 5.
- No service exceeds two pages per on-call shift for three consecutive months by month 6.
- 100% of the 180 services have a named owner, a detection contract and at least one symptom-based alert by month 7.
- All 28 teams have a documented rotation, a service ownership map and at least one trained on-call engineer by month 6.
- 100% of on-call shifts are paid under a published policy from month 2, with zero on-call-attributed voluntary attrition by month 6.
- On-call satisfaction scores 7 out of 10 or better in quarterly surveys from month 6.
- The IC roster holds at least 12 certified ICs covering 24x7 with no uncovered week from month 4.
- 100% of SEV0, SEV1 and SEV2 postmortems are published internally within 15 business days from month 5.
- Action items closed within 60 days rise from 17% to above 90%, with a median action age under 30 days, by month 6.
- Status page first update is posted within 30 minutes on at least 95% of SEV1 incidents from month 4.
- Zero missed regulatory notification windows on any incident requiring notification.
- The customer-impact record is complete for 100% of customer-impacting incidents from month 5, with SLA credits reconciled monthly.
- Annual SLA credits paid fall at least 50% from the $1.3M baseline within 12 months.
- Customer-impacting incidents fall at least 40% year over year from the 31-incident baseline.
- Review cadence is sustained: weekly operational review in at least 90% of weeks, 12 of 12 monthly reliability reviews, 4 of 4 quarterly executive reviews.
- Every page has a recorded disposition — fixed, tuned or deleted — within 10 working days, from month 4.
- SOC 2 Type II is passed with zero findings related to incident response.
- 100% of new engineers complete the incident-response onboarding module within 30 days of joining.
Steps (24):
1. Mandate, one owner, and the evidence clock
This step turns the CEO's email into a funded programme with a single accountable owner, and it starts the SOC 2 clock on day one.
- Appoint a Director of Incident Management reporting to the CTO, with a dotted line to the COO for customer and SLA matters and to the Head of Compliance for audit readiness.
- Publish a one-page charter: scope covers every customer-impacting, money-moving, security and data-integrity event across all 28 teams and both regions.
- Grant explicit authority to declare an incident, set severity, freeze deploys, page any engineer in the company, and approve customer messaging.
- Fund the envelope up front: tooling, training and drill time, on-call compensation, and a small programme team, roughly $500–700K a year against $1.3M in credits paid.
- State the return plainly to the steering group: credits avoided, churn avoided, and audit findings avoided.
**Start the evidence clock now.** A SOC 2 Type II report covers a period of operating effectiveness, so evidence cannot be backfilled; an audit in eight months means the operating period begins in the first weeks of this programme.
- Stand up a fortnightly steering group of the CTO, VP Engineering, Head of Support, Head of Compliance and two engineering managers.
- Open a programme risk register with the top risks, owners and review dates, and revisit it at every steering group.
- Make participation in the incident process a documented performance expectation for every engineering manager, not an optional extra.
2. Baseline evidence pack and cost-of-downtime model (depends on: 1)
You cannot demonstrate improvement without a defensible starting point, and you cannot win the argument about noise without numbers.
- Build a 12-month incident register and re-classify all 31 customer-impacting incidents by detection source, duration, customers affected and credits paid.
- Build the silent-failure register: incidents where no internal alert fired at all, which is the number that explains the 40% customer-detected rate.
- Run an alert census per tool, per team and per service: volume, page-to-action ratio, off-hours interruptions, and the 50 noisiest rules with a named owner each.
- Survey every on-call engineer and manager on burden, fairness, escalation clarity and pay expectations, targeting above 70% response.
- Interview Support, Account Management and Sales about how customers learn of incidents and what they are promised.
- Reconstruct the two command-ambiguity incidents minute by minute to find exactly where ownership lapsed.
- Build a cost-of-downtime model: dollars per minute of impact per customer journey, used later to sequence teams and justify funding.
- Publish the pack internally as the problem statement, and retain every artifact as management-review evidence for the audit.
3. SOC 2 control mapping and evidence architecture (depends on: 1)
Most programmes leave compliance to the end; this one maps controls in month one, because the mapping decides what the process must capture from day one.
- Map the process to the Trust Services Criteria for incident identification, evaluation, response, containment, recovery and communication (CC7.1–CC7.5, CC2.2, CC2.3, CC4.1, CC3.x).
- Write each control in plain language, with one named owner and its evidence artifact.
- Define the golden incident file: one auto-assembled, single-click export per incident containing timeline, roles, severity, communications, postmortem, actions and closure proof.
- Set retention, storage location and immutability so no control depends on a laptop, a private channel, or chat history that expires.
- Keep a gap register with owners and dates, reviewed fortnightly by the steering group.
- Meet the auditor's readiness team inside the first 90 days to test the control design before anything is built on top of it.
4. Severity and class taxonomy with the trigger matrix (depends on: 2)
Severity decides how loud the response is; class decides which playbook runs and what must never be traded away.
- Define severity in customer and money terms: customers affected, dollars at risk, regions lost, ledger integrity, settlement delay, data loss and SLA exposure.
- Use SEV0 for security, privacy or regulatory events; SEV1 for total or material loss of a payment path; SEV2 for degradation or single-region loss; SEV3 for limited impact with a workaround; SEV4 for internal-only issues and near-misses.
- Define response classes: Availability, Performance, Data integrity and Ledger, Security and Privacy, Third-party dependency, Settlement and Financial, Process failure.
**Class can raise the response but never lower it.** A SEV2 data-integrity incident gets SEV1 posture, because integrity failures are not recoverable by moving faster.
- Specify automatic triggers: loss of one AWS region, ledger write failures, replication lag above threshold, a missed settlement cut-off, payment success rate below threshold for five minutes.
- State who may declare — any engineer, Support agent or account manager — and who may downgrade: the Incident Commander alone.
- Map every level to its SLA credit exposure and to the customer-visible status page state.
- Include worked examples from the last 12 months so each of the 28 teams recognises its own incidents, and revalidate the taxonomy quarterly against real declarations.
5. Day-one operating rules and the minimum viable process (depends on: 1, 4)
The full process will take months; the first useful version must be live in two weeks using the tools that already exist.
- Publish ten day-one rules that need no procurement: a named owner within five minutes, one channel per incident, one register entry per incident, one person speaking to customers.
- Apply the ambiguity rule: if two responders disagree about whether this is an incident, it is an incident.
- Make declaring free: a false alarm is closed as a false declaration, tracked as a metric, and never criticised.
- Require a register entry within 24 hours for every customer-impacting incident, even a minimal one.
- Ban silent incidents: if we know, the customer hears it from us rather than from their own reconciliation.
- Run the first 30 days on manual command, with a rotating duty Incident Commander drawn from the 12 teams that already have on-call.
- Hold a 15-minute daily incident stand-up during month one to catch friction while it is still fresh.
6. Roles, command structure and the no-unowned-minute rule (depends on: 4)
The two hour-long command failures did not happen at declaration; they happened in the gap before it, when an alert had fired and nobody owned it.
- Publish one-page role cards: Triage Owner, Incident Commander, Deputy IC, Operations Lead, Internal Comms Lead, Customer Comms Lead, Scribe, Subject-Matter Responders, and Executive Sponsor for SEV1 only.
- Introduce the Triage Owner rule: whoever acknowledges the page owns the incident until an IC takes over or the incident is stood down.
**The IC owns the incident, not the fix, and does not debug.** An IC who starts troubleshooting has abandoned command.
- Give the IC explicit decision rights: declare and escalate severity, freeze changes, halt deploys, approve customer messaging, and pull in any engineer.
- Define minimum viable staffing per severity: SEV1 fills every role; SEV2 staffs IC, scribe, comms and responders; SEV3 staffs an IC and a scribe.
- Set handover discipline: four-hour maximum IC shifts on SEV1 with a written handover, and a deputy named within 15 minutes of declaration.
- Set responder behaviour: one channel, one bridge, no side channels, and every request phrased with a named owner and a time.
- Link the role cards from every paging notification so they are one tap away at 3 AM.
7. Lifecycle, declaration and escalation policy (depends on: 4, 6)
This step defines the mechanical path from an alert to a declared incident and back to normal service, removing judgment calls from the worst moments.
- Define states with entry and exit criteria: Detected, Triaged, Declared, Mitigated, Resolved, Postmortem, Closed, plus a Watch state with a hard 30-minute timer.
- Set targets: page acknowledged within five minutes, triage decision within fifteen, declaration within thirty for anything customer-facing.
- Build automatic escalation ladders with timers at every layer: responder, service owner, team manager, IC on-call, VP Engineering.
- Define the unresponsive-team path: fifteen minutes escalates to the team manager, thirty minutes to the IC, who may then direct any available engineer.
- Define change freeze and rollback authority during SEV1 and SEV2, with the single named condition that lifts the freeze.
- Enforce one incident, one record, with the timeline captured automatically from the channel and bridge rather than written from memory afterwards.
- Add the two-simultaneous-incidents rule: a second IC is designated, and no IC ever runs two incidents at once.
- Test every escalation path weekly with synthetic pages, and adjust the timings after the first month of real operation.
8. On-call architecture across 28 teams (depends on: 4, 6)
The objection is that engineers will not carry a pager for another team's code; the answer is to build rotations that make the objection structurally impossible.
- Run a Service On-Call rotation per team, covering only that team's own services.
- Run a Platform Duty rotation for genuinely shared infrastructure: the PostgreSQL ledger cluster, Kubernetes, networking, CI/CD and observability.
- Run a central Incident Commander roster of 12–16 certified senior engineers on one-week shifts with a primary and a secondary.
**State the consequence honestly.** Sixteen of 28 teams have no rotation today; each must build one or formally transfer service ownership to a team that has one, with the transfer dated and recorded.
- Merge small or low-traffic teams into shared rotations so no rotation has fewer than six engineers.
- Cap load in the scheduling tool: no engineer on call more than two weeks per quarter, enforced by configuration rather than negotiation.
- Publish a coverage matrix of all 28 teams showing services, owners, rotation size and gaps, reviewed monthly.
- Write on-call participation into engineering job levels and hiring criteria, so the commitment is set before someone joins.
9. Compensation, rest and the economics of opting out (depends on: 8)
Unpaid on-call is the most cited reason for resistance, so settle compensation before rollout, not during it.
- Move to paid on-call: a per-shift stipend benchmarked to the New York market, agreed with HR and Finance, published with an effective date before any team is asked to join a rotation.
- Pay for callouts at 1.5× the hourly rate for time actually spent mitigating, with a minimum block per interruption.
- Provide documented compensatory rest: no normal working day after a night incident, and the rest day is policy rather than a favour granted by a manager.
- Cap intrusion with a maximum number of off-hours pages per shift, and trigger a mandatory review whenever the cap is exceeded.
**Allow opt-out, but put a price on it.** An engineer may step out of a rotation, and their team buys coverage from the paid pool at a published internal rate, which turns a cultural argument into a visible budget decision.
- Publish amnesty: incident records, near-misses and false declarations are never used in performance reviews; only failure to report is a performance issue.
- Protect against gaming: pay attaches to shifts and callouts, never to incident counts.
- Check the New York labour, overtime and tax treatment with Legal and Finance before announcing, and review the policy every six months against real page volumes, attrition and survey results.
10. Detection strategy: journeys, synthetic signals and customer-report intake (depends on: 4)
Customers detected 40% of incidents first, which makes detection the highest-leverage business problem in this programme.
- Define SLIs and SLOs for the top 20 customer journeys, measured per region: payment initiation, authorisation, settlement, ledger read and write, API availability, webhook delivery and payout.
- Alert on symptoms against those SLOs, not on cause-based infrastructure metrics.
- Add synthetic transaction monitoring from outside AWS in both regions plus a third location, on a one-minute cadence for money-moving paths.
- Add ledger and shared-PostgreSQL signals: replication lag, connection saturation, write latency, lock waits, transaction ID exhaustion, checkpoint pressure and bloat.
- Open a customer-report intake so Support and account managers can raise an incident directly, and count that path as a detection source in all reporting.
- Apply the detection-gap rule: whenever a customer reports first, a detection-gap ticket is opened automatically and owned by the responsible team.
- Publish a detection contract per service — owner, at least one symptom alert, documented expected detect time — for all 180 services.
- Run a detection drill per team: break something in staging and see whether it pages before a human notices.
11. Paging contract and the noise-reduction programme (depends on: 2, 10)
3,400 alerts a month at 85% noise is the reason engineers resent the pager, and fixing it is the price of admission for everything else in this plan.
- Publish the paging contract: every page must be symptom-based, actionable, owned, mapped to a severity and class, and linked to a runbook.
**No runbook, no page**, enforced by a CI check on the alert definition itself.
- Separate paging from ticketing: only customer-impacting or imminent-impact signals may wake a human.
- Set a page budget per team and per service, with a remediation ticket opened automatically, owned by the engineering manager, when the budget is breached.
- Put new alerts on two-week probation as ticket-only until they have proved actionable.
- Define the default action for a noisy alert: fix, tune or delete within ten working days, with deletion treated as a legitimate and celebrated outcome and the deleted count published.
- Deduplicate and correlate at the ingest pipeline so one root cause produces one page instead of forty.
- Require an expiry date on every silencing rule and temporary threshold, so suppression cannot quietly become permanent blindness.
- Run a 90-day noise sprint with a public burn-down of the top 100 noisiest rules, each owned by a named manager.
12. Incident tooling consolidation and the golden incident file (depends on: 3, 6, 10, 11)
Six alerting tools and no single incident record are structural causes of the 22-minute detection and the three-hour mitigation.
- Choose one incident platform for paging, schedules, escalation policies, incident records, postmortem workflow and status-page integration, and choose it within four weeks.
- Consolidate the six alert sources into a single event pipeline feeding that platform, with deduplication, correlation and severity and class mapping applied at ingest.
- Integrate observability, runbooks, chat and voice bridges so responders see dashboards and runbooks inside the incident record, and the timeline is captured automatically.
- Implement the golden incident file export defined in S3, so audit evidence is one click rather than a reconstruction.
- Run a dual-run period alongside the legacy tools with a defined rollback, then switch the legacy tools off on a published cutover date.
- Host the status page outside the production failure domain so it survives a total platform outage, and prove that in a game day.
- Make the platform usable from a phone: an Incident Commander must be able to run a SEV1 from a mobile device at 3 AM.
13. Internal, customer and regulator communications (depends on: 6, 7, 12)
Today the status page is written by whoever is around; this step replaces improvisation with a clock, a named owner and pre-cleared templates.
- Set the internal cadence: first update within 15 minutes of declaration, then every 30 minutes for SEV1 and hourly for SEV2, whether or not there is progress.
- Never let an employee learn of an incident from the status page: internal communication leads, external follows.
- Set the customer cadence: status page updated within 30 minutes of a SEV1 and 60 minutes of a SEV2, with a no-new-information update still mandatory.
- Define the channel hierarchy: status page for all 2,100 customers, direct email to affected customers on SEV1, and named account-manager calls for the top 50 accounts.
- Pre-approve templates per severity and class with Legal and Compliance, each carrying its own next-update time.
- Forbid speculation: customer messages never guess at cause, never assign blame, and never commit to a root cause before the postmortem.
- Define the closing communication: resolution notice, how SLA credits will be handled, and the committed date for the written report.
- Build a regulator clock matrix covering event type, regulator, notification window, signer and the shortest applicable clock, including NYDFS Part 500, money-transmitter and banking notifications, breach notification, card-network rules and public-company disclosure.
**The regulatory clock starts at awareness, not at root cause.** Route every notification through Compliance, never Engineering, and pre-clear the templates.
14. Customer trust workstream and the SLA credit ledger (depends on: 13)
The $1.3M in credits is a symptom of having no single record of customer impact, and the CEO's inbox is a symptom of customers learning things late.
- Maintain one durable customer-impact record per incident: which customers, which journeys, from when to when, and the estimated credit.
- Use that one record for communications, credit computation, regulatory reporting and the annual review, so nobody reconciles two versions of the same outage.
- Automate the credit path end to end — computation, approval, customer notification and finance treatment — so credits stop being a manual scramble three weeks later.
- Reconcile accrued against paid credits monthly and report the result in the executive review.
- Give the status page a named product owner and map its components to customer journeys, not to internal services.
- Send a CTO-signed reliability note to the top 50 accounts and publish a quarterly reliability report to all customers.
- Give account managers a script of the facts they may state, the speculation they may not, and a path for customer escalations.
- Track credit avoidance against programme cost, so the funding case stays a number rather than an argument.
15. Postmortems: mandatory set, three levels, blameless by design (depends on: 6)
Postmortems currently happen for some incidents, in various formats; this step makes them mandatory where they matter and light where they do not.
- Make postmortems mandatory for every SEV0, SEV1 and SEV2, every SEV3 with customer impact or a repeat pattern, every near-miss the IC flags, and every incident where the process itself failed.
- Use three levels so the ritual matches the weight: a lightweight async review for SEV3 and SEV4, a facilitated postmortem for SEV2, and a full review with an executive sponsor for SEV1.
- Set deadlines: draft within five business days, blameless review within ten, internal publication within fifteen.
- Adopt one template: impact, timeline, detection, response, contributing factors across tooling, process, organisation and human factors, what went well, what went badly, and actions.
- Train a pool of blameless facilitators and require a trained one for every SEV1 review, never the IC.
**Ban blame language in the template and ban "human error" as a root cause.** The question is always what system condition made the error possible.
- Publish all postmortems internally by default, with security review only for genuinely sensitive material.
- Produce a customer-facing root-cause report for SEV1 incidents, especially for regulated and top-tier accounts.
16. Action items: capped, verifiable, with the repeat-incident rule (depends on: 15)
Eleven of 64 action items closed is not a tracking problem; it is a generation problem, because the process produces more actions than the organisation can absorb.
- Cap each postmortem at three action items, with anything beyond that going into a ranked reliability backlog.
- Require every action to have a named human owner, a due date, and a definition of done that is an artifact: a merged change, an alert that fired in a drill, or a test that fails without the fix.
- Prohibit self-reported closure; closure requires the artifact and sign-off by the process owner or the Incident Commander.
- Keep one reliability backlog in the engineering tracker with a mandatory label and a link back to the originating incident.
- Reserve a fixed share of each team's sprint capacity for reliability work, with unspent capacity visible to the vice presidents.
- Apply the repeat-incident rule: a second incident with the same contributing factor triggers a design review owned by an architect, not another action item.
- Run a weekly ageing review of open actions and escalate anything more than 30 days overdue to the VP of Engineering.
- Report action completion rate and median action age monthly, by team.
17. Metrics, dashboards and the review cadence (depends on: 2, 4, 16)
Define what good looks like, then measure it in a way that rewards reporting incidents rather than hiding them.
- Define outcome metrics: time to detect by source, time to mitigate by severity and class, share of incidents detected by customers, incidents by severity and class, SLA credits paid, and customer-impact minutes.
- Define process metrics, because the process itself has SLOs: IC assigned within five minutes, page acknowledgement rate, first-update timeliness, update-cadence adherence, postmortem timeliness, action closure rate, action median age, and IC roster coverage.
- Define health metrics: alert volume and noise ratio per team, off-hours pages per engineer, on-call satisfaction, training completion, and share of services with a detection contract.
**Never publish incident count as a team metric.** It rewards hiding incidents; publish reporting metrics instead — near-misses filed, detection gaps closed, false declarations made.
- Publish live dashboards visible to every engineer, refreshed daily, with every metric baselined against the S2 evidence pack and given a 90-day and a 12-month target.
- Run a 30-minute weekly operational review incident by incident, a monthly reliability review of trends and actions, and a quarterly executive review with the CEO's office.
- Hold a quarterly review of the process itself: what wasted time, what confused responders, and what should be deleted.
- End every review with decisions and named owners, never with numbers alone.
18. Training, certification and the drill programme (depends on: 6, 7, 13, 15)
A process that lives only on a wiki page fails on the first real page, so skills are built and tested before they are needed.
- Build a practical curriculum: how to be on call, how to triage and declare, how to run an incident as IC, how to communicate, and how to write a postmortem.
- Require certification before joining the IC roster: a written assessment plus a live simulated incident.
- Certify at least two ICs per team group so the central roster has depth across all 28 teams and no holiday week is left uncovered.
- Train communications leads separately on templates, cadences, customer language and the regulatory rules.
- Run monthly tabletops drawn from the last 12 months of real incidents, including region loss, ledger write failure and a missed settlement cut-off.
- Run quarterly game days with deliberately injected failure, including PostgreSQL failover, status-page outage and alerting-pipeline outage.
- Drill the process's own failure modes, not just technical ones: IC unreachable, comms lead on leave, two simultaneous SEV1s, a paging storm, and a false alarm that burns an hour.
- Audit the process for single points of failure: who alone can perform each critical task, and what happens in their holiday week.
- Keep a mandatory onboarding module for every engineer joining or transferring in, with audit-ready completion records.
19. Pilot with three to four teams, using real incidents (depends on: 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 18)
Do not roll out to 28 teams untested, and do not rely on drills alone when real incidents are better training material.
- Recruit three to four volunteer teams spanning criticality: a payment-path team, a ledger-adjacent team, the platform team that owns the shared PostgreSQL cluster, and a low-traffic team.
- Run the entire process end to end in the pilot: severity and class, roles, escalation, communications, postmortems, action tracking, and paid on-call.
- Treat real incidents during the pilot as the primary training material, and hold a retro within 48 hours of each one, run by the process owner while the friction is fresh.
- Instrument the pilot against the baseline and publish before-and-after numbers.
- Hold weekly retrospectives with the pilot teams, and change the written policies, the tooling and the training in response, documenting what changed and why.
- Set explicit exit criteria: rotation coverage achieved, zero unacknowledged pages over two weeks, postmortems delivered on time, and actions tracked to closure.
- Produce a pilot report that every later rollout conversation starts from.
20. Phased rollout to 28 teams, sequenced by cost of failure (depends on: 16, 18, 19)
Rollout is a staged migration with readiness gates, not an email announcement, and the sequencing matters more than the schedule.
- Split the 28 teams into four waves of roughly seven, ordered by incident density and customer-journey ownership: the teams whose incidents cost the most SLA credits go first, because their improvement is the visible proof.
- Leave three weeks between waves and define a per-team readiness checklist: services mapped and owned, alerts cleaned to the paging contract, runbooks written, rotation staffed to at least six engineers, training complete, manager briefed, compensation in effect.
- Hold a gate review with the process owner before each team joins, and move an unready team to the next wave with a dated remediation plan rather than granting an exception.
- Give each wave a named champion and explain the why using the pilot's numbers.
- Handle resistance directly and early: publish the own-your-code-own-your-pager rule and the paid on-call mechanics before each wave, never after.
- Hold a retro within 48 hours of each wave's first incident under the new process, and push accepted changes through change control.
- Retire legacy tools, informal escalation lists and the ad-hoc status page process at the end of each wave on a published cutover date, so two processes never run in parallel for long.
- Sequence the waves so no team faces both a rollout and the audit dry run in the same week.
21. Audit dry run and evidence review (depends on: 3, 20)
This step converts a good process into a provable one, about six weeks before the auditors arrive.
- Sample real incidents from the pilot and the early waves against each control's evidence requirements.
- Remediate every gap found and re-test the remediated control against the same sample, prioritising anything that risks a qualified opinion.
- Brief every engineer who may be interviewed so they can describe the process as they actually practise it, not as it is written.
- Prepare the auditor package: process documentation, sample incident records, paging logs, communication logs, training records, on-call schedules, the action register and alert quality metrics.
- Designate one audit liaison and a small evidence-request team, so requests do not land on all 28 teams at once.
- Rehearse the walkthrough with an Incident Commander and a communications lead, because auditors probe realism under pressure.
- Keep the audit liaison and the process owner as close to the same person as possible, so accountability for the control is also accountability for the evidence.
22. Standing governance and process ownership (depends on: 20, 21)
The classic post-audit failure is that the process freezes and then decays, so ownership has to outlive the programme.
- Establish a standing Incident Management Council chaired by the process owner, meeting monthly with engineering, support, compliance and product representation.
- Give the process owner a documented mandate to change standards, with formal change control for any change to severity, classes, roles, communication timings or compensation.
- Revalidate the severity and class taxonomy quarterly against real declarations, and re-baseline the entire metric set annually.
- Report quarterly to the executive team on the metric set and the top five systemic causes of incidents.
- Feed incident themes into architecture review, release engineering and hiring, so the programme improves the system and not only the response.
- Keep a public backlog of process improvement ideas so any engineer can propose a change, and act on it every quarter.
23. Resilience track for the shared ledger and multi-region blast radius (depends on: 1, 17)
Better incident management cannot save a single shared ledger during a corruption event, so this runs as a funded engineering track beside the process, not inside it.
- Reduce ledger blast radius: read replicas for read paths, connection-pool isolation per service domain, statement timeouts, write throttles and per-tenant limits.
- Build a ledger protection mode: an automatic tripwire that degrades to read-only or queues writes before corruption can spread.
- Rehearse multi-region failover with the ledger as the binding constraint, and set a measured recovery time objective per region.
- Test restore and point-in-time recovery on a schedule, and publish the timings.
- Tighten deploy safety: progressive rollout, automatic rollback on SLO burn, and change-freeze windows around settlement.
- Track a small number of resilience bets with named owners and quarterly milestones, reviewed in the same executive review as the incident metrics.
24. Culture, incentives and continuous improvement (depends on: 20, 22)
Incident management is a cultural change before it is a process change, and the incentives decide whether reporting goes up or down.
- Restate the amnesty at every rollout wave: reporting, near-misses and false declarations are safe; hiding is not.
- Recognise detection improvements, deleted alerts, closed detection gaps and well-written postmortems publicly, and reward the engineers who do them.
- Correct blame language in leadership communication on the spot, including inside the executive review.
- Watch for pager fatigue with real data, and take an engineer off rotation as soon as the intrusion cap breaks.
- Run a quarterly process retro with the whole engineering organisation and publish what changed as a result.
- Refresh the training and tabletop programme annually and immediately after any SEV1.
- Share wins publicly each quarter — faster detection, fewer customer-detected incidents, credits avoided — to hold momentum after the audit.
--- PROPOSAL 3 (agent qwen3.8-flash_refine_3, alibaba/qwen3.8-flash) ---
Estimated complexity: high
Success metrics: - Median Time to Detect (MTTD) < 5 minutes.
- Median Time to Mitigate (MTTR) < 60 minutes for SEV1.
- >90% of incidents detected internally (vs. 40% previously).
- SLA credits reduced to <$100K annually.
- Monthly alert volume < 600 with >95% signal-to-noise ratio.
- Postmortem action item completion rate > 80%.
- Zero incidents with command ambiguity > 15 minutes.
- 100% of SEV1/2 postmortems published within 15 days.
- Status page updates within 30 minutes for 95% of SEV1.
- 100% of teams integrated into on-call rotations.
- On-call satisfaction score > 7/10.
- 40+ certified Incident Commanders available 24x7.
- SOC 2 Type II audit passed with zero findings.
Steps (22):
1. Executive Mandate & SOC 2 Evidence Clock
Establish clear ownership and funding to transform the CEO's concern into a governed program with immediate compliance implications.
2. Baseline Measurement & Problem Statement (depends on: 1)
Quantify the current state to build a defensible problem statement and establish the data foundation for improvement tracking.
3. Severity & Class Taxonomy (depends on: 2)
Define unambiguous incident levels based on business impact to trigger specific response protocols.
4. Incident Command Roles & Structure (depends on: 3)
Standardize roles and decision rights to eliminate command ambiguity during active incidents.
5. Tooling Consolidation & Integration (depends on: 1)
Deploy a unified incident management platform to replace fragmented alerting tools and create a single source of truth.
6. Detection Strategy & SLO Implementation (depends on: 3)
Shift monitoring focus to customer-impacting symptoms to detect issues before customers do.
7. Alert Quality Standards & Noise Reduction (depends on: 5, 6)
Implement strict alerting rules to reduce noise and restore on-call trust.
8. On-Call Architecture & Coverage Model (depends on: 4)
Design a federated, team-owned on-call structure supported by central command coverage.
9. On-Call Compensation & Wellbeing Policy (depends on: 8)
Introduce paid, sustainable on-call rotations to remove resistance and mitigate burnout.
10. Playbooks & Communication Templates (depends on: 3, 4)
Create standardized playbooks and templates for each severity level to reduce cognitive load.
11. Escalation & Incident Lifecycle Policy (depends on: 4, 5)
Define the mechanical progression from alert to resolution, including automatic escalation paths.
12. Communication Workflows (Internal & External) (depends on: 4, 10)
Establish timed protocols for internal, customer, and regulatory communications.
13. Status Page & Customer Notification Automation (depends on: 5, 12)
Integrate incident management with public status pages to ensure timely and accurate updates.
14. Postmortem Policy & Blameless Framework (depends on: 3)
Mandate and standardize postmortems to focus on system improvement rather than individual error.
15. Action Item Tracking & Governance (depends on: 14)
Ensure postmortem action items are tracked, owned, and completed to prevent recurrence.
16. Metrics, Dashboards & Review Cadence (depends on: 2, 3)
Define and monitor key performance indicators to validate the effectiveness of the incident process.
17. Training, Certification & Drills (depends on: 4, 10, 12)
Build practical skills and readiness for incident response across all roles.
18. Pilot Program with Volunteer Teams (depends on: 5, 7, 9, 11, 13, 15, 17)
Test the end-to-end process on a small scale to refine procedures before full rollout.
19. Phased Rollout to All Teams (depends on: 18)
Deploy the process to all 28 teams in manageable waves to minimize disruption and manage resistance.
20. SOC 2 Control Mapping & Evidence Framework (depends on: 1, 3)
Align the incident management process with SOC 2 requirements to ensure audit readiness.
21. SOC 2 Dry Run & Audit Support (depends on: 19, 20)
Verify evidence sufficiency and prepare the team for the external audit.
22. Governance & Continuous Improvement (depends on: 19, 21)
Establish standing governance to sustain the process and drive ongoing improvements.
You do not know which of them was voted for, and you must not guess it: judge them on their merits. Your answer has these parts:
- "summary": two or three sentences on what the final round offers, as a whole.
- "assessments": one object per final proposal, with "proposal" (its number), "fitness" ("strong", "adequate" or "weak" for the task as stated), "strengths" and "weaknesses" (lists of concrete points: steps, order, metrics, realism, risk handling).
- "ranking": the numbers of the final proposals from best to worst.
- "ranking_reasons": why that order, naming what separates each one from the next.
- "versus_initial": one object per initial proposal, comparing YOUR FIRST CHOICE with it: "proposal" (its number in round 0), "verdict" ("better" if your first choice is better than that initial proposal, "worse" or "similar"), "why" and "how" (in what concrete ways).
- "improved_over_initial": true if your first choice is better than EVERY initial proposal.
- "improvement_summary": what the deliberation added or lost with respect to the initial proposals, overall, as a Markdown list of points.
- "process_evaluation": an evaluation of the process itself, not a narration, as a Markdown list with one point per question: whether the convergence between agents was earned by better arguments or was mere imitation, whether the agents criticised each other's ideas or just copied them, whether anyone questioned the task's premises, and what was lost along the way.
- "process_issues": a list of concrete problems you observed in the process.
- "suggestions": a list of concrete changes that would make the process produce a better plan.
[VOTE COMPARISON]
[SYSTEM]
You are an expert reviewer of multi-agent planning processes.
Several LLM agents drafted plans for a task, refined them over a number of rounds while seeing each other's proposals, and finally voted for the best one.
Be exhaustive but precise: name concrete steps, ideas and metrics, never generalities. Judge plans by their fitness for the task as stated, their realism, their completeness, the soundness of their order and dependencies, how measurable their success is and how they handle things going wrong.
You are an impartial evaluator, not a chronicler: assess the proposals and the process on their merits, never rationalise what happened or assume that the outcome was right.
After your analysis, answer in the requested structure.
Every text field you write will be read by a busy person who skims. Make it easy to skim: short sentences and short paragraphs; when you name several things, prefer a list to a paragraph, with sub-items when an item has parts, but keep a single fact as a sentence; lead with the point and then the evidence; name proposals and steps by number (P2, step 4); no preamble, no repetition of the question, no closing summary; bold at most one key phrase per item or paragraph. Text fields accept Markdown: a blank line between paragraphs, "- " for lists, **bold**.
[HUMAN]
Task given to the agents: "A B2B payments platform in New York (260 engineers in 28 teams, 2,100 customers, $4B processed a month, a 99.95% SLA with service credits) runs about 180 services on Kubernetes across two AWS regions, with a shared PostgreSQL cluster for the ledger. In the last 12 months: 31 customer-impacting incidents, median time to detect 22 minutes (customers detected 40% of them first), median time to mitigate 3 h 10 min, $1.3M paid in SLA credits, two incidents in which nobody was sure who was in charge for over an hour, and a CEO email about "outages we hear about from clients". On-call exists in 12 of the 28 teams, unpaid, fed by alerts from six different tools — 3,400 alerts a month, 85% of them noise; there is no severity scale, status-page updates are written by whoever is around, and postmortems happen for some incidents, in various formats, with action items that are rarely tracked (11 of 64 closed). A SOC 2 Type II audit in eight months will test incident response. Engineers push back against "carrying a pager for other teams' code".
Define the incident management process: severity levels and what each one triggers; roles (incident commander, communications, scribe, subject-matter responders) and how they are staffed 24x7 across 28 teams; the on-call structure, rotations, compensation and the rules for alert quality; detection and escalation paths; internal and customer communications (status page, account managers, regulators when required) with their timings; postmortems (when mandatory, format, blameless review, ownership and tracking of actions); the metrics and reviews that show whether it works; and how it is introduced across the teams without waiting for the audit."
You assessed the proposals of the final round without knowing the vote, and ranked them from best to worst as: 2, 1, 3. Your reasons were:
- **P2 over P1** on three things that matter operationally: a two-week minimum viable process (S5) so the organisation is covered during the months P1 leaves uncovered; a funded resilience track (S23) that gives the 60-minute SEV-1 target a mechanism instead of a hope; and internally consistent, dated metrics. P1's 3-minute status-page rule contradicts its own 30-minute metric, its alert target (<400 vs <600) and MTTM target (45 vs 60 min) disagree with its own steps, and step 11 is cut off mid-sentence.
- P1 beats P2 on numeric thresholds (severity triggers, comp figures, escalation tempo, training pass marks) but those are easier to fill in later than the structural gaps P2 closes.
- **P1 far above P3**: P1 is executable today; P3 is a table of contents. P3 also carries an unrealistic metric (40+ certified ICs) and lost every operational detail it had in earlier rounds.
The vote is now revealed. Each voter received the same final proposals and chose one:
claudeHaiku4.5_voter_1 (anthropic/claude-haiku-4-5) voted for proposal 1. Justification: Proposal 1 delivers the most actionable and detailed implementation path, with 22 concrete, richly specified steps covering all requirements. It combines aggressive but realistic timelines (month 5 complete rollout, month 8 audit-ready), unambiguous handling of the on-call resistance problem via three explicit rotations with no exceptions, and early control mapping (step 3) that positions for SOC 2 success. While Proposal 2 is nearly equivalent and slightly stronger on resilience separation and cultural sustainability, Proposal 1's clarity, specificity, and urgency-matched pacing directly serve the objectives of a company facing an immediate audit and critical operational failures.
deepseek-flash_voter_2 (deepseek/deepseek-flash) voted for proposal 2. Justification: Proposal 2 matches Proposal 1's depth and specificity but edges ahead on the brief's hardest asks: its P5 'day-one operating rules' delivers a live minimum-viable process in two weeks, directly answering 'introduce it without waiting for the audit,' and its dedicated resilience track plus culture/incentives steps address the single shared PostgreSQL ledger and the on-call resistance that a pure process plan would miss. Its 23 concrete, time-bound success metrics and regulator clock matrix are at least as complete as Proposal 1's, while Proposal 3 is markedly thinner with one-line step descriptions and fewer metrics.
qwen3.8-flash_voter_3 (alibaba/qwen3.8-flash) voted for proposal 2. Justification: This proposal offers the most realistic and detailed roadmap by introducing a 'minimum viable process' to bridge the gap between current chaos and full implementation. It explicitly addresses the specific cultural friction regarding pager ownership through structural separation of duties and financial incentives, ensuring higher adoption rates.
The selected plan is proposal 2 (2 of 3 votes).
Your answer has these parts:
- "agrees": true if the selected plan is your first choice, false otherwise.
- "comment": two to four short sentences confronting the vote with your ranking. If they differ, say on which criteria the voters and you weigh the proposals differently, and whether the voters' justifications reveal a bias (length, confidence, style, position in the list). If they agree, say whether the voters' justifications rest on the same merits you found or on others. Do not revise your ranking.
[PROCESS]
[SYSTEM]
You are an expert reviewer of multi-agent planning processes.
Several LLM agents drafted plans for a task, refined them over a number of rounds while seeing each other's proposals, and finally voted for the best one.
Be exhaustive but precise: name concrete steps, ideas and metrics, never generalities. Judge plans by their fitness for the task as stated, their realism, their completeness, the soundness of their order and dependencies, how measurable their success is and how they handle things going wrong.
You are an impartial evaluator, not a chronicler: assess the proposals and the process on their merits, never rationalise what happened or assume that the outcome was right.
After your analysis, answer in the requested structure.
Every text field you write will be read by a busy person who skims. Make it easy to skim: short sentences and short paragraphs; when you name several things, prefer a list to a paragraph, with sub-items when an item has parts, but keep a single fact as a sentence; lead with the point and then the evidence; name proposals and steps by number (P2, step 4); no preamble, no repetition of the question, no closing summary; bold at most one key phrase per item or paragraph. Text fields accept Markdown: a blank line between paragraphs, "- " for lists, **bold**.
[HUMAN]
Task given to the agents: "A B2B payments platform in New York (260 engineers in 28 teams, 2,100 customers, $4B processed a month, a 99.95% SLA with service credits) runs about 180 services on Kubernetes across two AWS regions, with a shared PostgreSQL cluster for the ledger. In the last 12 months: 31 customer-impacting incidents, median time to detect 22 minutes (customers detected 40% of them first), median time to mitigate 3 h 10 min, $1.3M paid in SLA credits, two incidents in which nobody was sure who was in charge for over an hour, and a CEO email about "outages we hear about from clients". On-call exists in 12 of the 28 teams, unpaid, fed by alerts from six different tools — 3,400 alerts a month, 85% of them noise; there is no severity scale, status-page updates are written by whoever is around, and postmortems happen for some incidents, in various formats, with action items that are rarely tracked (11 of 64 closed). A SOC 2 Type II audit in eight months will test incident response. Engineers push back against "carrying a pager for other teams' code".
Define the incident management process: severity levels and what each one triggers; roles (incident commander, communications, scribe, subject-matter responders) and how they are staffed 24x7 across 28 teams; the on-call structure, rotations, compensation and the rules for alert quality; detection and escalation paths; internal and customer communications (status page, account managers, regulators when required) with their timings; postmortems (when mandatory, format, blameless review, ownership and tracking of actions); the metrics and reviews that show whether it works; and how it is introduced across the teams without waiting for the audit."
THE INITIAL PROPOSALS (round 0):
--- PROPOSAL 1 (agent claudeHaiku4.5_initial_1, anthropic/claude-haiku-4-5) ---
Estimated complexity: high
Success metrics: - Median time to detect reduced from 22 minutes to <8 minutes within 6 months
- Customer-detected incidents drop from 40% to <5% within 6 months
- Median time to mitigate (MTTR) reduced from 3h 10min to <45 minutes for SEV-1 incidents within 6 months
- Annual SLA credits decrease from $1.3M to <$100k within 12 months
- Alert noise reduced from 3,400 per month (85% false positive) to <400 per month (>95% signal) within 3 months
- Zero incidents with command-and-control ambiguity (>1 hour without clear IC) within 2 months
- Postmortem action item completion rate reaches >80% (from 17%) within 4 months
- On-call satisfaction score (survey) reaches >7/10 for on-call engineers within 3 months
- All 28 teams integrated into incident management system with active on-call rotations by week 20
- SOC 2 Type II audit passes incident response controls with no findings 8 months from start
- Incident commander certification: 100% of active ICs trained and drilled within 2 months
- Monthly incident review meeting established and attended by leadership; trends documented
- New incident system integration complete: single alert tool, single dashboard, all 180 services feeding in, <5 min deployment
Steps (21):
1. Define severity levels and decision criteria
Create a four-tier severity framework (SEV-1 through SEV-4) that guides all downstream decisions about response, escalation, and communications.
Each level must specify: customer impact (revenue at risk, customers affected, data loss risk); financial threshold triggering service credits; whether an incident commander is required; response time SLA (e.g., SEV-1 < 5 min notification, SEV-4 < 2 hours); and the go-live decision tree (when to declare and when to resolve).
- SEV-1: Complete service down or critical path broken for >5% of customers; every minute costs money; IC required; 99.99% uptime threatened
- SEV-2: Significant degradation, features unavailable, affecting 1–5% of customers; IC typically required
- SEV-3: Minor impact, limited customer footprint or workaround exists; escalation path but not automatic IC
- SEV-4: Observations or minor issues; alert-driven, no escalation unless pattern emerges
2. Define incident roles and responsibilities (depends on: 1)
Create the organizational roles that operate during an incident: who is in charge, who talks to customers, who writes down what happened, who fixes the system, and how decisions are made under pressure.
Each role must have a single-sentence mission, decision authority, and escalation upward.
- Incident Commander: owns decision-making and timeline; declares severity; resolves conflicts; may or may not be technical
- Deputy IC: shadow to IC and takes over if IC becomes unavailable
- Communications Lead: writes status page, notifies account teams, manages customer perception
- Scribe: records decisions, who did what, key timestamps; not responsible for fixing
- SME Responders: engineers with context on the failing service(s); take IC's direction without debate
3. Design 24x7 on-call rotation structure (depends on: 2)
Build a rotation model that covers all 28 teams with primary and backup on-call engineers every hour across weekdays, evenings, weekends, and holidays; addresses the pager-carrying resistance.
Key design decisions: Is coverage per-team (each team owns its services) or pooled (shared responder pool handles anything)? How many people per rotation? How long are shifts (one week, two weeks)? When can engineers opt out without leaving the team exposed? Which roles are on-call (IC, communications, SME)?
- Recommend: dedicated IC pool (4–6 people in fast rotation) + per-team SME on-call for each team's own services
- Recommend: two-week rotation blocks to reduce handoff friction
- Recommend: one primary, one secondary per slot; secondary handles during primary's escalation
- Provide swaps, blackout dates, and a rule that no engineer is on-call more than 2 weeks per quarter
4. Define on-call compensation and incentives (depends on: 3)
Create a pay model that makes on-call acceptable and rewards engineers who carry the pager; ties compensation to real business risk.
- Base on-call stipend: e.g., $500–1,000 per week while on-call (regardless of incidents)
- Callback pay: 1.5× hourly rate for time spent mitigating incidents during off-hours
- Incident bonus: $50–100 extra per SEV-1 or SEV-2 incident mitigated (recognition)
- Comp time: full business day off after an incident that required >2 hours mitigation during night/weekend
- Annual bonus tie-in: 10–20% bonus multiplier for flawless on-call reviews
- Communicate: position as investment in reliability, not punishment for being online
5. Consolidate alert routing infrastructure
Replace six alert tools with a single ingestion and routing system; stop engineers from being woken by duplicate alerts, and make escalation automated instead of manual.
Evaluate existing tools (likely candidates: PagerDuty, Opsgenie, or Incident.io) or build a lightweight wrapper. The system must: accept alerts from all 180 services; deduplicate and correlate (same outage, different monitoring source); route to correct on-call engineer; expose an API for playbook automation; log every alert for postmortem analysis.
- Choose tool by week 2 of S5
- Migrate alerting endpoints from 6 sources to 1 by week 4
- Set up audit trail and retention
- Ensure mobile app works (on-call engineers need to engage from phone)
6. Build alert quality rules to cut noise (depends on: 1, 5)
Implement rules that automatically suppress the 85% of alerts that are noise (flapping, transient errors, auto-recovered conditions). Target: <400 actionable alerts per month.
Rules to implement: suppress alerts if service auto-recovered within 30 seconds; deduplicate same alert from multiple monitoring sources; suppress alerts for known maintenance windows; group flapping alerts (same service, >5 occurrences in 2 minutes) into a single page to on-call; rate-limit alerts from noisy services (e.g., max 1 alert per 5 minutes per service until silence clears).
- Audit existing 3,400 alerts per month: which are true signals, which are noise
- Tag each alert source with severity level (S1, S2, S3, S4 from S1)
- Create exceptions list: services known to be noisy, require different rules
- Weekly review: alert teams that trigger >50 alerts per week for reduction strategies
7. Implement automated detection and escalation paths (depends on: 1, 5, 6)
Wire the alert system to automatically escalate based on time or severity; removes the need for manual judgment calls during chaos.
Logic: SEV-1 alert arrives → IC notified instantly via phone call + SMS + Slack + mobile; if IC does not acknowledge within 2 minutes, page deputy IC; Communications Lead pinged simultaneously. SEV-2: on-call SME for that service + IC notify via Slack and mobile, escalate to IC's manager if not acknowledged in 10 min. SEV-3/4: on-call SME only, escalate after 30 min.
- Implement in alert routing system (S5)
- Test all paths weekly via synthetic page to on-call
- Track escalation metrics: how many pages reach secondary, how many hit manager
- Adjust timing based on first month of operations
8. Build incident dashboard and status tracking (depends on: 5)
Create a single source of truth during an incident that every responder sees in real time: who is on-call, incident timeline, who said what, current status, next steps.
Dashboard displays: active incidents and their severity; who is the IC and communications lead; timeline of all events (alert fired, IC assigned, customer notified, mitigation started, resolved); Slack channel and mobile notification status; on-call rosters (who is on-call right now for each team); postmortem link as soon as incident closes.
- Integrate with alert tool (S5) to auto-populate incident creation and initial severity
- Push updates to status page and customer account managers automatically
- Log all timeline entries for audit and postmortem completeness
- Mobile-optimized so IC can work from any device
9. Write incident playbooks for each severity (depends on: 1, 2, 3)
Create a one-page (or one-screen) reference for the IC and SMEs during an incident; sequences the steps and removes ambiguity.
Each severity level gets its own playbook: who gets paged (roles, order); first questions to ask (is it real, how big, who knows); what the IC should declare in first message (status page text, account manager notification, regulatory trigger); how long before escalating to executive team; decision rules for going dark vs. continuing to update customers.
- SEV-1 playbook: immediate IC + comms + CTO notification; customer status every 5 minutes
- SEV-2 playbook: IC + comms + tech lead notification; status every 15 minutes
- SEV-3 playbook: on-call SME + comms if customer-visible; status every 30 min or as resolved
- SEV-4 playbook: on-call SME only; update customers only if promised SLA is at risk
- Include decision trees: is this SEV-1 or SEV-2? Is it our code or dependency? Escalate or containment?
10. Define internal communication workflows (depends on: 2, 3)
Specify who informs whom, in what order, via what channel (call, Slack, email) during an incident; prevents gaps like "nobody knew who was in charge for an hour."
Workflow for SEV-1: IC assigned → IC calls CTO/VP Eng and incident channel lead within 1 minute; incident declared in #incidents Slack channel with severity, IC name, service affected; SME on-call for that service joins call automatically; IC pushes updates to #incidents every 5 minutes or when material change occurs. For SEV-2: IC notifies team leads via Slack, updates #incidents every 15 min. Define escalation: if IC is unreachable, deputy IC takes over and announces it.
- Create a phone tree or on-call list accessible to responders
- Set expectations: "If you don't hear from IC in 2 minutes, call them"
- Use a single incident Slack channel per incident (auto-created by incident tool)
- Log all comms in the incident dashboard for postmortem review
11. Design customer communication and status page process (depends on: 1, 2)
Plan when and how to inform customers, account managers, and regulators; ensure 2,100 customers are not learning about outages from Twitter before you tell them.
Rules by severity: SEV-1 detected → status page updated within 3 minutes (even if root cause unknown; post "Investigating"); account managers of affected customers called within 5 minutes; regulatory notification (if payment processing down) queued for approval; customer email within 10 minutes with ETA for next update. SEV-2: status page within 10 min, account managers called within 15 min, email if affecting >10 customers. SEV-3/4: no customer communication unless SLA at risk.
- Empower Communications Lead to update status page without IC approval if delay >3 min
- Prepare templated messages for common scenarios (database failover, data pipeline stuck, service crashed)
- Route regulatory notifications through legal/compliance; don't wait for perfect root cause
- Track customer impact in real time: how many customers affected by severity
12. Establish blameless postmortem process and format (depends on: 1, 2)
Build a systematic way to learn from incidents so the same failure does not happen twice; counter the fear that admitting a mistake leads to being blamed.
Mandatory postmortems: all SEV-1 and SEV-2 incidents, within 48 hours of resolution. Optional but encouraged: SEV-3 if interesting or if >3 of same type in 30 days. Format: what was the user-visible impact and for how long; what was the root cause (not "human error" but the system condition that made error possible); timeline of discovery and response; action items with owner and deadline; blameless tone (focus on process and system design, not individual mistakes).
- Assign a facilitator (not the on-call IC) to run postmortem
- Attendees: IC, comms lead, SMEs involved, team lead, customer success if customer-facing
- Write postmortem in shared doc; make it findable (searchable, linked from incident)
- No discussion of "who screwed up"; only "why did the system allow this to happen"
13. Build action item tracking and accountability (depends on: 12)
Create a system that tracks postmortem action items so they are not forgotten; currently 11 of 64 (17%) are being tracked, leaving 53 unfinished improvements.
System: each postmortem generates action items (e.g., "add monitoring for X," "update runbook for Y," "write test for Z"). Each item gets: clear description, owner (engineer's name), due date (1–4 weeks based on priority), severity (critical = must do before similar incident happens again; important = improve next month; nice-to-have = backlog). Action items live in a dedicated Jira project visible to all teams; owners are accountable (their manager reviews quarterly). Weekly: incident commander reviews open items due that week. Monthly: each team's postmortem items reviewed in their standup.
- Export action items from postmortem document to tracking system automatically
- Require IC to sign off that an action is complete before closing
- Report on completion rate as a metric (target: >80% by month 3)
14. Define incident metrics and KPIs
Establish what "good" looks like; measure so you can improve. Target metrics for 12 months out: mean time to detect 8 minutes (vs. 22 now), customers detect first <5% of incidents (vs. 40%), MTTR 45 minutes for SEV-1 (vs. 190), SLA credits <$100k/year.
Metrics to track: (1) MTTD = time from incident start to first alert/report; disaggregate: external report vs. internal detection. (2) MTTR = time from first report to full mitigation; track by severity and by service. (3) Customer-reported incidents per month (should drop to <2 per month). (4) Alert signal-to-noise ratio (goal: <5% false positive after S6 rules). (5) On-call satisfaction (survey: would you do this again?). (6) Postmortem action completion rate. (7) Incident commander and responder utilization (hours per week per person).
- Dashboard: auto-populated from incident tool, updated daily
- Disaggregate by team and service: which teams have bad MTTR? Which service is most incident-prone?
15. Create review cadence and governance process (depends on: 14)
Establish regular rhythm to inspect the metrics, spot trends, and adjust the process itself; prevent the system from calcifying.
Weekly: incident commander and on-call lead review prior week—number of incidents, any escalations, any communication gaps. Monthly: director-level incident review—trends by service, top causes of incidents, action item status, whether severity classification is working. Quarterly: full leadership review—MTTD, MTTR, customer impact, SLA credit spend, on-call satisfaction score, any systemic changes needed. Annually: audit the entire process for SOC 2 compliance.
- Assign meeting owners: weekly = on-call lead; monthly = director of reliability; quarterly = VP Eng + CFO (SLA cost) + customer success
- Use same data dashboard (S14) for all reviews
- Publish a monthly "incident newsletter" to all engineers: what happened, what we learned, what's improving
16. Prepare SOC 2 Type II audit checklist (depends on: 1, 2, 9, 12, 13, 14, 15)
Document that the incident management system meets the control requirements for a SOC 2 audit; audit happens in 8 months, so this work builds confidence in coverage.
Audit will test: (1) Is there a defined incident response process? (2) Are roles and responsibilities clear? (3) Are incidents logged and tracked? (4) Is root cause analysis performed? (5) Are action items tracked and completed? (6) Is on-call staffing adequate? (7) Are communications timely? (8) Are postmortems documented and blameless? Create a control mapping document that links each SOC 2 requirement to your process (S1–S15). Collect evidence: incident logs, postmortem documents, action item tickets, metrics reports, training records.
- Designate a compliance owner (often a reliability lead or security engineer)
- Run a mock audit at month 6 to identify gaps
- Ensure all postmortems and incidents are retained and searchable for auditor review
17. Develop implementation and rollout plan (depends on: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16)
Create a phased timeline to roll out the incident management system across all 28 teams; avoids big-bang failure and builds credibility.
Recommended structure: Phase 1 (weeks 1–4): build and test infrastructure (S5, S8, alerting); deploy severity levels and roles (S1, S2); pick pilot teams (2–3 high-traffic teams). Phase 2 (weeks 5–12): train pilot teams, run incident drills, refine playbooks based on learning; expand to half of remaining teams. Phase 3 (weeks 13–20): full rollout to all 28 teams; continue drills; track metrics. Phase 4 (weeks 21–28): stabilize, iterate on metrics, prepare for audit.
- Assign a release manager to coordinate across teams
- Create a detailed Gantt chart with swim lanes (infra, process, training, rollout)
- Identify risks: competing priorities, engineers worried about pager burden, tool adoption friction
- Plan stakeholder engagement: weekly updates to eng leadership, monthly town halls for all engineers
18. Build training and documentation (depends on: 17)
Create role-specific education so engineers understand the new system and are confident executing during an incident.
Training tracks: (1) All engineers: 30-min video on severity levels, communication expectations, postmortem format, action items (mandatory); (2) On-call responders: 2-hour workshop on playbooks, alert tool, escalation paths, when to call manager, case studies of real incidents, mobile app walkthrough; (3) Incident commanders: 4-hour IC bootcamp covering leadership under pressure, decision-making, communicating with executives, status page discipline, postmortem facilitation, practiced drills; (4) Communications leads: templates, when to update, how to talk to customers, regulatory notification rules.
- Record videos so async teams can learn on their schedule
- Create runbooks and quick-reference cards for each role (print + digital)
- Pair new on-call engineers with experienced responder for first week
- Require IC certification before anyone joins IC rotation (pass a practical drill)
19. Execute staged rollout across teams (depends on: 18)
Progressively activate the incident management system with feedback loops at each stage; reduces risk of system-wide failure.
Wave 1 (week 6–8): 3–4 pilot teams begin on-call rotations and incident response using new system; capture feedback daily. Wave 2 (week 10–14): 8–10 additional teams, incorporating lessons from Wave 1; ensure diversity of team types (payment processing, monitoring, data pipeline, auth, etc.). Wave 3 (week 15–20): remaining teams; by now, the system is proven and less hand-holding needed.
- Daily retros with Wave 1 teams: what worked, what was confusing, what broke
- Each wave produces a "lessons learned" document that informs the next
- Track adoption metrics: how many incidents reported per team, alert quality, MTTD/MTTR
- Address resistance: engineers who are skeptical of the system, on-call burden, tool friction; assign a "change champion" in each team
20. Run incident response drills and simulations (depends on: 19)
Practice incidents in a controlled setting so responders gain confidence and gaps are found before a real crisis.
Drill 1 (week 8, pilot teams only): simulated database failover; test alert routing, IC assignment, on-call notification, status page update, postmortem. Drill 2 (week 12, all teams): full incident spanning multiple services; test escalation, communications, decision-making. Drill 3 (week 18): chaos drill with intentional confusion—what if IC is unreachable, what if comms lead is on PTO, what if an alert tool goes down? Quarterly thereafter: repeat drills with new on-call engineers.
- Schedule drills at a time when key people are available (not Friday afternoon)
- Run drill with live observers (reliability lead, manager) who provide feedback
- Measure drill success: all steps completed on time, communications clear, postmortem written
- Debrief immediately: what did we learn, what do we change?
21. Establish continuous improvement process (depends on: 15, 20)
Build a feedback loop so the incident management system evolves based on what you learn.
Monthly improvement cycle: in the monthly incident review (S15), leadership identifies one or two friction points from recent incidents (e.g., "alert tool response is slow," "postmortem meetings are boring," "action items keep slipping"). Assign an owner to investigate and propose a fix (change to playbook, tweak to tool config, additional training, process change). Test the fix with pilot team, gather feedback, roll out if successful. Annual review: can we reduce on-call burden, simplify communication, improve MTTD further? Do any teams need special handling (new services, difficult customers)?
- Keep a public backlog of improvement ideas (Jira, shared doc) so any engineer can suggest
- Quarterly all-hands: share wins ("MTTD improved 50%," "on-call satisfaction up," "zero customer-detected incidents") to build momentum
- Link improvements back to metrics (S14): show that changes actually worked
--- PROPOSAL 2 (agent deepseek-flash_initial_2, deepseek/deepseek-flash) ---
Estimated complexity: high
Success metrics: - Median time to detect reduced from 22 minutes to 5 minutes or less by month 6.
- Customer-detected incidents reduced from 40 percent to 10 percent or less of all customer-impacting incidents by month 9.
- Median time to mitigate reduced from 3 h 10 min to 60 minutes or less for SEV1 and SEV2 by month 9.
- Customer-impacting incidents reduced by at least 40 percent year over year from the 31 incident baseline.
- SLA credits paid reduced by at least 50 percent year over year from the $1.3M baseline.
- 100 percent of SEV1 incidents have a named Incident Commander within 5 minutes and a scribe within 15 minutes.
- Zero incidents in which command authority is unclear for more than 15 minutes.
- 28 of 28 teams have a documented on-call rotation, an owning service map and at least one trained on-call engineer by month 6.
- Central Incident Commander roster holds at least 40 certified ICs covering 24x7 with no single point of failure.
- 100 percent of on-call rotations are paid under a published policy by month 5.
- Monthly alert volume reduced from 3,400 to below 700, with a false-positive rate below 20 percent.
- No service exceeds 2 pages per on-call shift, measured monthly for three consecutive months.
- 100 percent of SEV1 and SEV2 postmortems published internally within 15 business days.
- At least 90 percent of postmortem action items closed within 60 days, up from 17 percent (11 of 64).
- Status-page first update published within 30 minutes on at least 95 percent of SEV1 incidents.
- Zero missed regulatory notification windows on any incident requiring notification.
- SOC 2 Type II audit passed with no findings related to incident response.
- Review cadence sustained: weekly operational review in at least 90 percent of weeks, 12 of 12 monthly reliability reviews, 4 of 4 quarterly executive reviews.
- On-call satisfaction at 70 percent or higher on the quarterly survey, with zero on-call-attributed voluntary attrition.
- 100 percent of new engineers complete the incident-response onboarding module within 30 days of joining.
Steps (21):
1. Programme charter, ownership and executive mandate
This step turns the CEO's email into a funded programme with a named owner and explicit authority. Without it, every downstream decision stalls in cross-team negotiation.
- Appoint a single accountable process owner (for example a Director of Incident Management) reporting to the CTO, with a dotted line to the COO for customer and SLA matters.
- Publish a one-page charter covering scope (all customer-impacting and money-moving incidents), decision rights, and the power to override team preferences during an active incident.
- Define the funding envelope: tooling licences, training time, exercise time and on-call compensation, with an indicative annual figure.
- Set the timeline against the SOC 2 date: a working process in four months, evidence accumulating from month two, audit-ready by month seven.
- Stand up a steering group with CTO, VP Engineering, Head of Support, Head of Compliance and one engineering manager per region.
- Agree that incident-process participation is a documented performance expectation for engineering managers, not an optional extra.
2. Baseline measurement and evidence pack (depends on: 1)
You cannot demonstrate improvement without a defensible starting point, and you cannot win the argument about noise without numbers.
- Build a 12-month incident register: date, retro-assigned severity, detection source, time to detect, time to mitigate, customer impact, services involved and SLA credits paid.
- Quantify the alert estate per tool, per team and per service; compute page-to-action ratio, list the 50 noisiest rules and count off-hours interruptions per engineer.
- Survey on-call engineers and managers on burden, fairness and clarity of escalation, targeting a response rate above 70 percent.
- Interview Support, Account Management and Sales about how customers learn of incidents and what they complain about.
- Document exactly where the current process breaks: unclear command in the two known incidents, postmortem action closure at 11 of 64, and ad-hoc status-page authorship.
- Publish the pack internally as the problem statement and retain it as management-review evidence for the SOC 2 audit.
3. Severity taxonomy and trigger matrix (depends on: 1, 2)
Severity is the keystone of the whole process. Every other rule, from paging to communications timing to postmortems, is keyed off it.
- Define four levels plus a special SEV0 for security or regulatory events: SEV1 for total or material loss of a payment path, SEV2 for degradation or single-region loss, SEV3 for limited impact with a workaround, SEV4 for internal-only issues and near-misses.
- Anchor each level in customer and money terms: customers affected, dollars at risk, regions lost, ledger integrity, settlement delay.
- Specify automatic triggers, for example loss of one AWS region, ledger write failures, a missed settlement cut-off, or payment success rate below threshold for five minutes.
- State who may declare each level (any engineer, Support or account manager may declare) and who may only recommend a downgrade (the Incident Commander alone).
- Map each level to SLA credit exposure and to the customer-visible status-page state.
- Include worked examples from the last 12 months so teams recognise their own incidents in the definitions.
- Add a review clause: the taxonomy is re-validated quarterly against real declarations.
4. Incident roles, command structure and decision rights (depends on: 2, 3)
The two incidents where nobody knew who was in charge for over an hour are the direct brief for this step.
- Define roles with one-page role cards: Incident Commander, Deputy IC, Operations Lead, Internal Communications Lead, Customer Communications Lead, Scribe, Subject-Matter Responders and an Executive Sponsor for SEV1 only.
- State the core rule plainly: the IC owns the incident, not the fix, and does not debug.
- Give the IC explicit decision rights: declaring and escalating severity, freezing changes, halting deploys, approving customer messaging and calling additional responders.
- Define minimum viable role coverage per severity: SEV1 staffs every role, SEV3 staffs an IC and a scribe only.
- Define handover discipline: maximum four-hour IC shifts on SEV1, a written handover template, and a Deputy IC nominated within 15 minutes.
- Set behaviour standards for responders: one incident channel, one bridge, no side channels, and explicit asks with a named owner and a time.
- Publish role cards on the internal wiki and link them from every paging notification.
5. Escalation, paging and incident lifecycle policy (depends on: 3, 4)
This step defines the mechanical path from an alert to a declared incident and back to normal service.
- Define lifecycle states with entry and exit criteria: Detected, Triaged, Declared, Mitigated, Resolved, Postmortem, Closed.
- Set acknowledgement targets: page acknowledged within five minutes, triage decision within fifteen, declaration within thirty for anything customer-facing.
- Define escalation ladders per layer (responder, service owner, team manager, IC on-call, VP Engineering) each with an automatic timer.
- Make escalation blameless and automatic: no responder is ever criticised for escalating, and timers fire whether or not a human asks.
- Define change freeze and rollback authority during SEV1 and SEV2, and the single condition that lifts the freeze.
- Enforce one incident, one record: the incident record is the sole source of truth for timeline, roles and communications.
- Require every SEV1 and SEV2 to produce an automatically captured timeline from channel and bridge, never one written from memory afterwards.
6. Detection strategy: SLOs, signals and customer-journey monitoring (depends on: 2, 3)
Customers detected 40 percent of incidents first. That number is the reason this step exists.
- Define SLIs and SLOs for the top 20 customer journeys, including payment initiation, settlement, ledger read and write, API availability and webhook delivery, measured per region.
- Require symptom-based alerting on those SLOs rather than cause-based alerting on infrastructure metrics.
- Add synthetic transaction monitoring from outside AWS in both regions and a third location, on a one-minute cadence for money-moving paths.
- Add ledger and shared-PostgreSQL specific signals: replication lag, connection saturation, write latency, transaction ID exhaustion and checkpoint pressure.
- Open a customer-reported path so Support and account managers can raise an incident directly, and count that path as a detection source in reporting.
- Publish a detection contract per service: every one of the 180 services needs an owner, at least one symptom alert and a documented expected detect time.
- Fund a separate resilience track to reduce shared-cluster blast radius, because better detection will not save a single shared ledger during a corruption event.
7. Alert quality standard and noise-reduction programme (depends on: 2, 3, 6)
3,400 alerts a month with 85 percent noise is the reason engineers resent the pager. Fixing it is the price of admission for everything else.
- Publish alert standards: every page must be symptom-based, actionable, owned, linked to a runbook and mapped to a severity. No runbook, no page.
- Separate paging from ticketing: only customer-impacting or imminent-impact signals may page; everything else becomes a ticket or a dashboard entry.
- Set a noise budget per team and per service, for example no service may exceed two pages per on-call shift, measured monthly.
- Define the default action for a noisy alert: fix, tune or delete within ten working days, with deletion treated as a legitimate and successful outcome.
- Introduce correlation and deduplication at the event pipeline so a single root cause produces one page instead of forty.
- Require expiry dates on every silencing rule and temporary threshold so suppression cannot become permanent blindness.
- Run a 90-day noise sprint with a visible burn-down of the top 100 noisiest rules, each owned by a named manager.
- Report page-to-action ratio per team in the monthly reliability review.
8. On-call architecture and 24x7 coverage model across 28 teams (depends on: 3, 4)
This is the hardest political step. The answer to carrying a pager for another team's code is that every team carries its own, and the platform carries the shared risk.
- Adopt a federated model: every service has exactly one owning team, and that team's primary on-call carries its own pager. No team is paged for code it does not own.
- State the consequence honestly: 16 of 28 teams currently have no on-call. They must build one or formally transfer ownership of their services to a team that will.
- Merge small or low-traffic teams into shared rotations so no rotation has fewer than six engineers, below which coverage is not sustainable.
- Create a separate, centrally staffed Incident Commander on-call roster drawn from trained senior engineers across all teams, covering 24x7.
- Define primary and secondary per rotation, with the secondary engaged only on a no-acknowledge or an explicit request.
- Define coverage across the two AWS regions and New York business hours: one global IC rotation, service on-call aligned to their service's users.
- Define the unresponsive-team path: fifteen minutes without acknowledgement escalates to the team manager, thirty minutes to the IC, who may then direct any available engineer.
- Publish a coverage matrix of all 28 teams showing services, rotation size and gaps, reviewed monthly.
- Make on-call participation an explicit expectation in engineering job levels and hiring criteria.
9. On-call compensation, wellbeing and sustainability policy (depends on: 8)
Unpaid on-call is the most cited reason for resistance. The policy must be settled before rollout, not negotiated during it.
- Move to paid on-call: a per-shift stipend or salary uplift agreed with HR and Finance and benchmarked to the New York market.
- Pay event-based compensation for incident callouts outside business hours, with a minimum call-out block.
- Provide compensatory rest: no engineer works a normal day after a night incident, and the rest day is documented, not granted as a favour.
- Cap intrusion by defining a maximum number of off-hours pages per shift, with a mandatory review triggered whenever it is exceeded.
- Define a voluntary opt-out path for engineers with genuine constraints, balanced by an explicit obligation that someone else is paid to take the shift.
- Include on-call expectation and compensation in offers and job descriptions so the commitment is set before hiring.
- Publish the policy with an effective date before any team is asked to join a new rotation.
- Review the policy every six months against actual page volumes, attrition and survey results.
10. Internal and customer communications policy with timing SLAs (depends on: 3, 4)
Today the status page is written by whoever is around. This step replaces improvisation with a clock and a named owner.
- Set internal cadence: first update within 15 minutes of declaration, then every 30 minutes for SEV1 and 60 minutes for SEV2, regardless of whether there is progress.
- Set customer cadence: status page updated within 30 minutes of a SEV1 and 60 minutes of a SEV2, and a no-new-information update is still mandatory.
- Define the channel hierarchy: status page for everyone, direct email to affected customers on SEV1, named account-manager calls for the top 50 accounts.
- Prepare templates per severity in advance with legal and compliance pre-approval, covering detection, impact, workaround, mitigation and next-update time.
- Define regulatory obligations explicitly: money transmitter and banking regulator notification windows, security breach notification, and who signs off (Compliance, not Engineering).
- Prohibit speculation: customer communications never guess at cause or blame and never commit to a root cause before the postmortem.
- Define the closing communication: resolution notice, expected SLA credit handling and the committed date for a written report.
- Assign a named Customer Communications Lead per incident with a trained deputy on every SEV1.
11. Status page, notification tooling and account-manager playbook (depends on: 10)
Policy without tooling collapses at three in the morning. This step makes publishing a five-minute action.
- Upgrade or replace the status page so components map to customer journeys rather than internal services, with subscriber control per component.
- Integrate the incident tool with the status page so the incident record drives the update and the public timeline.
- Provide one-click templates pre-filled with severity, impact language and next-update time.
- Give account managers a playbook: contact tree, what they may say, what they must not say, and how to escalate a customer question into the incident channel within minutes.
- Define the SLA credit process end to end, covering computation, approval, customer notification and finance treatment, so credits stop being a manual scramble.
- Host the status page outside the production failure domain so it survives a total platform outage.
- Test publishing during game days, including a simulated status-page outage and a simulated loss of the primary region.
12. Postmortem policy, template and blameless review process (depends on: 3, 4)
Only 11 of 64 action items closed means the postmortem ritual is currently a writing exercise. This step rebuilds it around learning and tracking.
- Make postmortems mandatory for every SEV1 and SEV2, every SEV3 with customer impact or a repeat pattern, and any near-miss the IC flags.
- Set deadlines: draft within five business days, blameless review within ten, internal publication within fifteen.
- Adopt a single template: impact, timeline, detection, response, contributing factors, what went well, what went badly, action items.
- Train a pool of blameless facilitators and require a trained facilitator for every SEV1 review.
- Prohibit counterfactual and blame language in the template, and require contributing factors across tooling, process, organisation and human factors.
- Limit action items to a small number of concrete, verifiable, owned changes with dates.
- Publish all postmortems internally by default, with security review only for genuinely sensitive material.
- Produce a customer-facing root cause report variant for SEV1 incidents, especially those affecting regulated or top-tier accounts.
13. Corrective action tracking and reliability backlog governance (depends on: 12)
A postmortem without durable action tracking is a complaint, not a control.
- Create a single reliability backlog in the engineering tracker with a mandatory label, owner, due date and link to the originating incident.
- Define closure criteria that require evidence: a merged change, a tested alert or a verified drill, never a self-reported status change.
- Protect capacity by reserving a fixed percentage of each team's sprint for reliability work, with unspent capacity visible to vice presidents.
- Run a weekly ageing review of open actions and escalate anything overdue by more than 30 days to the VP Engineering.
- Report closure rate and median age monthly, targeting more than 90 percent closed within 60 days.
- Require a repeat incident in the same area to trigger a design review rather than another action item.
14. Incident tooling consolidation and integration (depends on: 3, 5, 7, 11)
Six alerting tools and no single incident record is a structural cause of the 22-minute detection and the three-hour mitigation.
- Select one incident management platform for paging, on-call schedules, escalation policies, incident records and postmortem workflow.
- Consolidate the six alerting sources into a single event pipeline feeding that platform, with deduplication and severity mapping applied at ingest.
- Integrate with platform and ledger observability so responders see dashboards and runbooks inside the incident record.
- Integrate chat and bridge: incident channel auto-created, timeline auto-captured, decisions logged as they happen.
- Define the data model and retention required for SOC 2 evidence: who did what, when, and under whose authority.
- Run a dual-run period alongside the old tools with a defined rollback, then switch off the legacy tools on a published date.
- Budget for licences, migration effort and a two-week hardening period after cutover.
15. Training, certification and exercise programme (depends on: 4, 5, 10, 12)
A process that exists only on a wiki page fails on the first real page. Skills have to be built and tested before they are needed.
- Build a short practical curriculum: how to be on call, how to declare an incident, how to run an incident as IC, how to communicate and how to write a postmortem.
- Require certification before joining the IC on-call roster: a written assessment plus a live simulated incident.
- Train at least two certified ICs per team group so the central roster has depth across all 28 teams.
- Run monthly tabletops on realistic scenarios drawn from the last 12 months, including region loss and ledger corruption.
- Run quarterly game days with deliberately injected failure, including shared-PostgreSQL failover and status-page outage.
- Add a mandatory onboarding module for every engineer joining or transferring in, with completion records suitable for audit.
- Track training completion by team and publish it in the monthly reliability review.
16. Metrics, dashboards and review cadence (depends on: 2, 3)
The programme needs a public scoreboard, or it will quietly rot after the audit.
- Define the outcome metrics: time to detect by source, time to mitigate, percentage of incidents detected by customers (target below ten), incidents by severity and SLA credits paid.
- Define the process metrics: declaration latency, page acknowledgement rate, IC roster coverage, first-update timeliness and update-cadence adherence.
- Define the health metrics: alert volume and noise ratio per team, off-hours pages per engineer, postmortem timeliness, action closure rate and action age.
- Publish live dashboards visible to every engineer, not only to managers, refreshed daily.
- Institute a weekly operational review of 30 minutes going incident by incident, a monthly reliability review of trends and actions, and a quarterly executive review with the CEO's office.
- Baseline every metric against the S2 evidence pack and set 90-day and 12-month targets.
- Require every review to end with decisions and owners, not just numbers.
17. Pilot with volunteer teams (depends on: 5, 7, 9, 11, 12, 13, 14, 15, 16)
Do not roll out to 28 teams untested. Run the entire process end to end with a small cohort first.
- Recruit three to four volunteer teams covering a mix of criticality: one payment-path team, one ledger-adjacent team, one platform team and one low-traffic team.
- Run the complete process in the pilot: new severity scale, roles, escalation, communications, postmortems, action tracking and paid on-call.
- Instrument the pilot against the S16 metrics and compare results with the S2 baseline.
- Hold weekly retrospectives with pilot teams and iterate on the written policies, the tooling and the training.
- Fix the top issues found before any wider rollout and document what changed and why.
- Produce a pilot report with before-and-after numbers to carry into every rollout conversation.
- Set explicit pilot exit criteria: rotation coverage achieved, no unacknowledged pages over a defined period, postmortems delivered on time and actions tracked.
18. Phased rollout to all 28 teams (depends on: 13, 16, 17)
Rollout is a staged migration with readiness gates, not an email announcement.
- Sequence the 28 teams into four waves of roughly seven teams, ordered by customer impact, with three weeks between waves.
- Define a per-team readiness checklist: services mapped and owned, alerts cleaned to standard, runbooks written, rotation staffed, training complete and manager briefed.
- Hold a gate review with the process owner before each team joins, and move unready teams to the next wave with a dated remediation plan.
- Give each wave a named champion and run an internal communications cadence that explains the why using pilot numbers.
- Handle resistance directly by publishing the own-your-code-own-your-pager rule and the paid on-call mechanics before each wave, not after.
- Retire legacy tools, escalation lists and the informal status-page process at the end of each wave on a published cutover date.
- Harvest feedback formally at each wave and push accepted changes back into the policy documents through change control.
19. SOC 2 incident-response control mapping and evidence framework (depends on: 1, 3, 10, 12)
The audit will test incident response. Mapping controls now means evidence accumulates from the first real incident under the new process.
- Map the process to the relevant Trust Services Criteria for incident identification, response, evaluation of incidents and communication of security events.
- Write control statements in auditor language and name a single owner for each control.
- Specify the evidence artifact for each control: incident record, severity classification, paging log, communications log, postmortem, action tracker entry and training record.
- Set evidence retention and storage location so nothing depends on a laptop or on chat history that expires.
- Run an early walkthrough with an experienced compliance partner or the auditor's readiness team to test the design before the audit window.
- Flag controls that are blocked until rollout reaches the relevant teams, and keep a gap register with owners and dates reviewed fortnightly by the steering group.
20. SOC 2 dry run, gap remediation and audit support (depends on: 18, 19)
This step converts a good process into a provable one, a few months before the auditors arrive.
- Run an internal dry run six weeks before the audit window, sampling real incidents from the pilot and early waves against each control's evidence requirements.
- Remediate every gap found, prioritising anything that risks a qualified opinion.
- Brief every engineer who may be interviewed so they can describe the process as they actually practise it.
- Prepare the auditor package: process documentation, sample incident records, training records, on-call schedules and the action tracking register.
- Designate a single audit liaison and a small evidence-request team so requests do not land on 28 teams at once.
- Rehearse the walkthrough with the incident commander and communications roles, since auditors probe realism under pressure.
- Schedule the dry run and the audit alongside rollout waves so no team is hit with both in the same week.
21. Standing governance, process ownership and continuous improvement (depends on: 18, 20)
The classic post-audit failure is that the process freezes and then decays. This step keeps it alive.
- Establish a standing Incident Management Council chaired by the process owner, meeting monthly with engineering, support, compliance and product representation.
- Give the process owner a documented mandate to change standards, and require formal change control for any change to severity, roles, communications timings or compensation.
- Re-validate the severity taxonomy quarterly against real declarations and re-baseline the metric set annually.
- Fund a resilience roadmap separate from incident response: reducing shared-database blast radius, multi-region failover readiness and deploy safety.
- Report quarterly to the executive team on the metric set and on the top five systemic causes of incidents.
- Feed incident themes into architecture review, release engineering and hiring so the programme improves the system, not just the response.
- Refresh the training and tabletop programme annually and after any SEV1.
--- PROPOSAL 3 (agent qwen3.8-flash_initial_3, alibaba/qwen3.8-flash) ---
Estimated complexity: high
Success metrics: - MTTD reduced to under 10 minutes with >90% detection occurring before customer reporting.
- MTTR for Sev1/Sev2 incidents reduced to under 60 minutes.
- SLA credits paid reduced by 50% within 12 months.
- On-call alert volume reduced by 80% (from 3,400 to <600 actionable alerts/month).
- 100% of Sev1/Sev2 incidents have a blameless postmortem completed within 5 days.
- >90% completion rate of postmortem action items within 30 days of assignment.
- Zero ambiguity in incident command: 100% of major incidents have a designated IC logged in the timeline.
- SOC 2 Type II audit passes incident response controls with zero critical findings.
Steps (12):
1. Establish Governance and Executive Sponsorship
Secure explicit CEO and CTO backing to mandate the program, framing it as a revenue protection and compliance necessity rather than an engineering burden. Form a cross-functional Incident Management Task Force.
- Identify a single executive owner (e.g., VP of Engineering or Head of Platform) with authority over budget and staffing.
- Create a charter defining the scope: severity definitions, on-call compensation, and tool consolidation.
- Communicate the strategic priority to all 28 team leads to preempt pushback regarding pager responsibility.
2. Define Severity Matrix and Response SLAs (depends on: 1)
Create a rigorous, binary severity scale to eliminate ambiguity in escalation and communication speed. Align internal response times with the 99.95% SLA.
- **Sev1 (Critical)**: System-wide outage, data corruption, or >5% transaction failure rate. Immediate executive page.
- **Sev2 (Major)**: Single customer segment impacted or >1% transaction failure rate. High-priority team response.
- **Sev3 (Minor)**: Degraded performance or isolated bug. Standard business-hours follow-up.
- **Sev4 (Cosmetic)**: No functional impact. Backlog item.
- Define maximum detection and mitigation times for each level.
3. Restructure On-Call Organization and Compensation (depends on: 1)
Transform the current unpaid, fragmented on-call system into a professional, compensated, and centralized rotation to ensure 24x7 coverage across all 28 teams. Address the 'pushback' by making it mandatory and paid.
- Consolidate the 12 current on-call teams into a unified, rotating pool covering all 28 teams.
- Implement a fair compensation model: base stipend + hourly rate for active hours + incident bonus for Sev1/2.
- Define a strict 'Buddy System' or 'Follow-the-Sun' rotation if geographically feasible, otherwise ensure adequate team size to rotate fairly.
- Mandate that every team must participate in the rotation; no exemptions.
4. Define Incident Roles and RACI Model (depends on: 2)
Eliminate the 'who is in charge' ambiguity by pre-defining roles and establishing a clear chain of command for every incident. Roles are assigned per incident, not permanently held.
- **Incident Commander (IC)**: Single point of authority. Owns the timeline, delegates tasks, and approves communications. Must be senior.
- **Scribe**: Maintains the real-time timeline log of actions, timestamps, and decisions.
- **Communications Lead**: Owns external/status page updates and internal Slack broadcasts. Shields the IC from interruptions.
- **Technical Leads (SMEs)**: Deployed by IC to troubleshoot specific services. They do *not* coordinate with each other; they report to IC.
- Create a 'Rapid Response Team' (RRT) for high-severity incidents involving the shared PostgreSQL cluster.
5. Consolidate Monitoring and Enforce Alert Hygiene (depends on: 3)
Reduce the 3,400 monthly alerts (85% noise) to a signal-to-noise ratio that allows engineers to trust the pager. Standardize tooling.
- Consolidate the six disparate alerting tools into a single platform (e.g., PagerDuty, Opsgenie, or Datadog).
- Implement 'Actionable Alerting': If an alert pages, there must be a documented, immediate runbook action. No 'informational' pages.
- Set a hard cap on alert volume per engineer/night. Breaching the cap triggers a mandatory alert quality review.
- Deprecate legacy tools that cannot map to the severity matrix defined in S2.
6. Build Detection and Escalation Automation (depends on: 5)
Reduce the 22-minute mean time to detect (MTTD) by automating the path from metric breach to page. Ensure customers never find an issue before the internal pager does.
- Implement synthetic transactions for core payment flows to detect failures before real users do.
- Configure escalation policies: If not acknowledged in 5 minutes, auto-escalate to the team lead, then to the Incident Commander pool.
- Map specific Kubernetes and PostgreSQL metrics to the Severity Matrix (S2) to trigger automated pages.
- Create a 'Major Incident Channel' in Slack that is automatically opened and pinned when a Sev1/2 is triggered.
7. Standardize Communication Templates and Timing (depends on: 4, 6)
Codify the rules for external and internal communications to stop the 'whoever is around' approach. Link these timings to the Severity Matrix.
- **Status Page Updates**:
- Sev1: First update within 5 minutes of detection. Updates every 15 minutes.
- Sev2: First update within 30 minutes. Updates every 60 minutes.
- **Customer Direct Mail**: For Sev1 impacting specific accounts, Account Managers send pre-approved templates within 1 hour of mitigation.
- **Regulator Notification**: Define legal triggers for NY financial regulations.
- Create a 'Comms War Room' Slack channel restricted to the IC and Comms Lead.
8. Formalize the Postmortem Process (depends on: 4, 7)
Transform postmortems from optional/various formats into a mandatory, blameless, and tracked process. This directly addresses the SOC 2 audit requirement for continuous improvement.
- Mandate a postmortem for every Sev1 and Sev2 incident within 5 business days.
- Enforce a 'Blameless' charter: Focus on systemic/process failures, not individual errors.
- Adopt a standardized template: Timeline, Root Cause Analysis (5 Whys or Fishbone), Impact Analysis, and Action Items.
- Action Items must have: Owner, Due Date, and Priority. They are tracked in Jira/Linear, not the doc.
9. Implement Action Item Tracking and Governance (depends on: 8)
Solve the '11 of 64 closed' problem by integrating postmortem actions into the core engineering workflow and executive reviews. Ensure fixes are shipped.
- Integrate Jira with the Postmortem tool (e.g., OpsLevel, Rootly, or Notion API). Auto-create tickets for every AI.
- Block 'Sev1 Closure' until all high-priority Action Items are closed or explicitly deferred with executive approval.
- Weekly Engineering Leadership review of open postmortem actions older than 30 days.
- Tie action item completion rates to team performance reviews.
10. Deploy Training and Runbook Library (depends on: 6, 7, 8)
Address the skill gap and the 'pushback' by equipping engineers with the tools and knowledge to succeed on-call. Reduce anxiety through clarity.
- Audit all 180 services. Require every service to have a 'Page-Worthy' runbook.
- Run 'Game Day' (Chaos Engineering) exercises quarterly to practice role assignments (IC, Scribe, Comms) in a simulated Sev1.
- Conduct mandatory training for the new on-call rotation and compensation model.
- Provide a 'Playbook' for the Incident Commander role, including scripts for communication.
11. Phased Rollout and Change Management (depends on: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10)
Introduce the new process in phases to manage cultural resistance and technical debt, ensuring adoption before the SOC 2 audit. Do not 'big bang' the change.
- **Month 1-2**: Launch Severity Matrix and Consolidated Tooling (S2, S5). Pilot on 5 high-traffic teams.
- **Month 3-4**: Roll out new On-Call Compensation and Roles (S3, S4) to all 28 teams. Enforce new Communication Templates.
- **Month 5-6**: Enforce strict Postmortem and Action Item Tracking (S8, S9).
- **Month 7**: Final audit readiness check. Simulate SOC 2 auditor questions on incident response.
12. Define Metrics for Success and Continuous Review (depends on: 2, 6, 8, 9)
Establish a dashboard to prove the new system works, linking directly to the $1.3M savings in SLA credits. Monitor the health of the incident process itself.
- Track MTTD (Target: <5 min), MTTR (Target: <30 min for Sev2), and Detection Ratio (Target: >95% internal detection).
- Monitor 'Alert Noise' (target: <5% of alerts result in a non-actionable page).
- Track 'Action Item Completion Rate' (Target: 100% on-time closure for Sev1/2).
- Calculate 'SLA Credit Avoidance' vs. On-Call Program Cost to prove ROI to leadership.
HOW THE ROUNDS WENT (from the round analyses):
Round 0: All three agents converge on the same skeleton — severity matrix → roles → paid on-call → tool consolidation → comms SLAs → blameless postmortems with tracked actions → metrics → phased rollout — but differ sharply in depth and in the on-call model. P2 is the most operationally concrete (baseline evidence pack, SLO-based detection, federated ownership, SOC 2 control mapping plus dry run); P1 is exhaustive but serializes training and drills after full rollout; P3 is the compact version and leaves audit evidence and detection engineering thin.
shared: **Four-tier severity scale as the keystone**: every plan keys paging, comms cadence and postmortem obligation off SEV1–SEV4 (P1 S1, P2 S3, P3 S2).
shared: Same role set — IC who commands but does not debug, comms lead, scribe, SME responders — with explicit decision rights (P1 S2, P2 S4, P3 S4).
shared: Collapse the six alerting tools into one platform and attack the 85% noise with actionable-alert standards, dedup and per-team caps (P1 S5–S6, P2 S7/S14, P3 S5).
shared: Paid on-call plus mandatory blameless postmortems for SEV1/SEV2 with action items tracked in Jira and reviewed by leadership; pilot first, then waves, not a big bang (P1 S4/S12/S13/S19, P2 S9/S12/S13/S18, P3 S3/S8/S9/S11).
differences: **On-call architecture**: P2 S8 is federated — one owning team per service, "own your code, own your pager", minimum rotation of six, 16 teams must build a rotation or transfer ownership, plus a central 24x7 IC roster. P3 S3 does the opposite, merging 12 rotations into one unified pool covering all 28 teams, which directly recreates the "pager for other teams' code" grievance it claims to solve. P1 S3 sits in between: dedicated IC pool of 4–6 plus per-team SME on-call.
differences: **Detection engineering**: only P2 S6 defines SLIs/SLOs for the top 20 customer journeys, symptom-based alerting, external synthetics in three locations, ledger-specific PostgreSQL signals (replication lag, TXID exhaustion) and a separate blast-radius resilience track. P3 S6 has synthetics only; P1 treats detection mostly as alert routing and never defines SLOs.
differences: **SOC 2 rigour**: P2 devotes S19–S20 to Trust Services Criteria mapping, a named owner and evidence artifact per control, retention rules and a dry run six weeks out. P1 S16 has a control mapping plus a month-6 mock audit. P3 has one bullet in S11 ("simulate auditor questions") — far too light for an eight-month Type II window.
differences: **Ordering flaws differ**: P1 S17 depends on all sixteen prior steps and pushes training (S18), rollout (S19) and the first drill (S20) to the end, so nobody practises before going live. P2 front-loads a 12-month baseline register (S2) that every metric later hangs on. P3 S11 defers comp and role rollout to months 3–4 while alert cleanup starts month 1, and P3 S9 ties action-item completion to performance reviews, which cuts against its own blameless charter (S8).
Proposal 1: 21 steps covering severity tiers, roles, a hybrid IC-pool/per-team SME rotation, concrete comp numbers ($500–1,000/week stipend, 1.5x callback, comp day), single alert tool with suppression rules, per-severity playbooks, status-page timings (SEV1 in 3 minutes), postmortem and action-item tracking, metrics and governance cadence. Rollout is a single mega-step (S17) depending on all sixteen predecessors, followed by training, waves and drills. Targets are the most aggressive of the three: MTTD <8 min, credits <$100k, 0 findings.
Proposal 2: Starts with a funded charter and a named process owner (S1), then a 12-month incident and alert baseline used as both problem statement and audit evidence (S2). Builds severity with a SEV0 for security/regulatory events, federated per-team on-call with a central IC roster, SLO- and synthetic-based detection, a 90-day noise sprint with a two-pages-per-shift budget, comms timing SLAs with pre-approved legal templates and a status page outside the failure domain, evidence-based action closure, then pilot, four gated waves and a SOC 2 dry run. Ends with a standing council and a separate resilience roadmap.
Proposal 3: Twelve steps: executive charter, severity matrix tied to transaction failure rates, a single mandatory paid on-call pool for all 28 teams, RACI roles plus a Rapid Response Team for the shared PostgreSQL cluster, tool consolidation with runbook-or-no-page hygiene, synthetic transactions and auto-escalation, tight status-page timings (SEV1 first update in 5 minutes), mandatory 5-day postmortems, Jira-integrated action tracking, game days, and a month-by-month rollout to month 7. Metrics target MTTD <5 min and 95% internal detection.
Round 1: P2's round-0 architecture became the de facto template: P1 and P3 both rebuilt their plans on it step-for-step, while P2 itself deepened its version with genuinely new mechanics (Triage Owner, severity x class, priced opt-out, three-action cap, regulator clock matrix, SOC 2 evidence clock). The round converged strongly on structure; what still separates the plans is depth of incident mechanics, realism of comms timings and where audit work sits in the sequence.
differences: **Command mechanics between alert and declaration.** P2 alone closes the gap that caused the two hour-long ownership failures: Triage Owner rule (step 5), Watch state with a 30-minute timer, "declaring is free", the ambiguity rule and the two-simultaneous-SEV-1 rule (step 6). P1 has an escalation ladder (step 9) but no owner-from-first-ack rule; P3 has no escalation or lifecycle step at all and dropped its round-0 5-minute auto-escalation.
differences: **On-call shape.** P2 builds three rotations including a paid Platform Duty for the shared PostgreSQL/Kubernetes estate (step 8) and prices opting out against a paid pool (step 9). P1 (steps 10–11) and P3 (step 5) stop at federated team rotations plus a central IC roster, leaving shared infrastructure ownership implicit.
differences: **Customer communication timings and money.** P1 demands a status-page update within 3 minutes and SEV-1 updates every 5 minutes (step 13) — contradicting its own success metric of 30 minutes. P2 uses 30/60 minutes (step 12) plus a regulator clock matrix naming NYDFS Part 500 and a customer-impact ledger driving SLA credit automation (step 13). P3 uses 15/30 minutes (step 9) with no credit process at all.
differences: **Where audit work sits.** P2 maps controls and defines the "golden incident file" in month one (step 3) on the argument that Type II evidence cannot be backfilled. P1 places control mapping at step 21, P3 at step 12 after postmortems. P3 also schedules game days (step 16) and the metrics dashboard (step 17) only after full rollout, so nothing is drilled or measured during the pilot.
influences: P1 and P3 rebuilt on P2's round-0 skeleton almost wholesale: charter (P2 s1), baseline evidence pack (s2), severity trigger matrix (s3), role cards (s4), SLO/synthetic detection (s6), "no runbook, no page" and the 90-day noise sprint (s7), federated on-call with a six-engineer floor (s8), paid on-call with compensatory rest (s9), pilot then four waves with readiness gates (s17, s18), control mapping and dry run (s19, s20).
influences: P1 kept only one structural idea of its own: severity playbooks with decision trees (its round-0 s9, now step 12); everything else in its 23 steps mirrors P2's ordering.
influences: P2 took P3's ROI framing (P3 s12, credit avoidance vs programme cost) into its step 13, and P1's mobile-pager requirement (P1 s5/s8) into step 11 ("run a SEV-1 from a phone at 3am").
influences: Nobody adopted P1's round-0 per-incident bonus (s4) — P2 explicitly bans pay attached to incident counts (step 9) — and nobody adopted P3's tying of action-item completion to performance reviews (s9), which P2 contradicts with a published amnesty.
influences: P1 carried over P2's round-0 metric of 40+ certified ICs, which P2 itself cut to 12–16 this round as more realistic for 260 engineers.
Proposal 1 (improved): P1 abandoned its generic round-0 framework and adopted P2's structure nearly step-for-step, gaining a charter, baseline evidence pack, SLO-based detection, a paging contract, a federated on-call model and a pilot-then-waves rollout. It kept its own useful severity playbooks step. Residual weaknesses are internal inconsistencies in timings and severity definitions.
Proposal 2 (improved): P2 kept its 21-step shape but added several mechanisms that close real gaps rather than restating policy: the SOC 2 evidence clock, severity x class, the Triage Owner rule, three rotations, a priced opt-out, an action-item cap and a regulator clock matrix. It also made its own targets more realistic.
Proposal 3 (improved): P3 grew from 12 to 20 steps by adopting P2's skeleton — charter, baseline, detection SLOs, compensation, control mapping, dry run, culture — which fills most of the prompt's requirements it previously skipped. It remains the thinnest plan in mechanics, dropped its own escalation automation, and its dependency order pushes drills and metrics past full rollout.
Round 2: P1 absorbed almost the entire distinctive vocabulary of P2's round-1 plan (Triage Owner, severity×class, three rotations, golden incident file, capped action items), producing two near-twin heavyweight plans; P2 added the genuinely new ideas of the round (a two-week minimum viable process, a funded ledger-resilience track, a cost-of-downtime model). P3 went the other way and collapsed into 22 bare titles with no content, losing everything that made it assessable.
differences: **Bridging the gap before tooling exists**: P2 S5 defines ten day-one rules, a manual duty-IC rotation drawn from the 12 teams that already have on-call, and a daily 15-minute stand-up for month one. P1 has no interim process between the charter (S1) and platform selection (S11); P3 has none either.
differences: **Prevention as a funded track**: P2 S23 is a standalone engineering roadmap with concrete bets (ledger read-only tripwire, connection-pool isolation, PITR restore tests with published timings, rollback on SLO burn). P1 keeps resilience as two bullets inside S9 and S22; P3 does not mention it.
differences: **Communication clocks**: P1 S13 demands a status-page update within 3 minutes of SEV-1 and internal updates every 5 minutes; P2 S13 sets 15 minutes internal first, 30 minutes to status page, with a mandatory no-news update. P1's 3-minute rule contradicts its own success metric of 30 minutes for 95% of SEV-1s.
differences: **Level of specification**: P1 and P2 give thresholds, timers, dollar ranges and dates throughout; P3 R2 gives only step titles and a one-line rationale each — no severity definitions, no timings, no compensation mechanics, no dates on any metric.
influences: P1 took nearly all of P2's round-1 signature ideas: Triage Owner (P2 S5→P1 S5/S6), severity×class with "class can raise, never lower" (P2 S4→P1 S4), the three-rotation model (P2 S8→P1 S7), priced opt-out and amnesty (P2 S9→P1 S8), golden incident file and month-one control mapping (P2 S1/S3→P1 S1/S3), three capped action items and the repeat-incident design review (P2 S15→P1 S16).
influences: P2 took P1's weekly synthetic-page testing of escalation ladders (P1 R1 S9→P2 S7) and P1's status-page-component-to-customer-journey mapping plus a named status-page owner (P1 R1 S14→P2 S14).
influences: P2 took P3's culture and change-management step (P3 R1 S18) and turned it into S24: on-the-spot correction of blame language, pager-fatigue monitoring, public recognition for deleted alerts.
influences: P3 took P2's "evidence clock" framing into its S1 title but nothing else of substance; it adopted no new mechanisms this round.
influences: Nobody adopted P3's error budgets triggering feature freezes (P3 R1 S8) or its rule that a SEV-1 cannot close until high-priority actions are done (P3 R1 S11) — P1 S16 and P2 S16 instead cap actions at three and track them in a separate reliability backlog.
Proposal 1 (improved): P1 rewrote itself around P2's round-1 mechanisms while keeping its own depth. New steps 4, 5, 7, 12, 14, 15, 20, 21 replace vaguer round-1 equivalents, and detection, escalation and postmortem policy are now far more concrete. A truncated step 11 and a few internal contradictions are the cost.
Proposal 2 (improved): P2 kept its round-1 architecture intact and added the three things it was missing: an immediately usable interim process, a funded prevention track, and an explicit culture step. Compliance sequencing, metrics and communications are essentially unchanged and were already strong.
Proposal 3 (worsened): P3 discarded all step content and submitted 22 titles with a single sentence of rationale each. Every operational detail it had in round 1 — severity thresholds, status-page timings, compensation mechanics, wave schedule, training curriculum — is gone, and several metrics were loosened or stripped of dates.
THE PROPOSALS OF THE FINAL ROUND (2):
--- PROPOSAL 1 (agent claudeHaiku4.5_refine_1, anthropic/claude-haiku-4-5) ---
Estimated complexity: high
Success metrics: - Median Time to Detect from 22 minutes to <5 minutes for internal detection and <10 minutes for customer journeys by month 6; <2% of customer-impacting incidents first detected by customers (vs. 40% currently) by month 9.
- Median Time to Mitigate for SEV-1 from 3h 10min to <45 minutes by month 9; SEV-2 to <90 minutes by month 9.
- Monthly alert volume from 3,400 to <400 with >95% signal-to-noise ratio by month 5; all alerts satisfy paging contract (symptom-based, owned, runbook-linked, severity-mapped).
- Zero incidents with command ambiguity lasting >15 minutes from month 2 onwards; 100% of SEV-1 and SEV-2 incidents have named IC logged in timeline within 5 minutes; 100% of Triage Owners assigned within acknowledgement.
- 100% of the 28 teams have a documented Service On-Call rotation or a signed service-ownership transfer by month 4; Platform Duty rotation and IC roster both staffed and live by month 3.
- On-call compensation policy published and effective by month 1; on-call satisfaction score ≥7/10 by month 6; zero on-call-attributed voluntary attrition by month 6.
- SLA credits paid from $1.3M annually to <$100K by month 12; credit avoidance (prevented credits) tracked and reported monthly.
- 100% of mandatory postmortems (SEV-0, SEV-1, SEV-2, and repeat incidents) published internally within 15 business days by month 4.
- Postmortem action item completion rate from 17% (11 of 64) to >90% within 60 days by month 6; median action age <30 days; zero repeat incidents caused by the same contributing factor without a design review.
- 100% of the 180 services have a named owner, a detection contract, and at least one symptom-based alert by month 6.
- Status-page first update published within 30 minutes for ≥95% of SEV-1 incidents by month 4.
- 100% of new engineers complete incident-response onboarding within 30 days of joining; IC certification includes written exam and live simulation; ≥2 certified ICs per team group; zero uncovered hours in 24x7 IC roster.
- Weekly operational review held in ≥90% of weeks; 12 of 12 monthly reliability reviews; 4 of 4 quarterly executive reviews; all reviews end with documented decisions and owners.
- SOC 2 Type II audit passes all incident-response controls (CC7.1–7.5, CC2.2–2.3, CC4.1, CC3.x) with zero findings by month 8.
- All pilot and rollout incidents captured with complete golden incident files (timeline, roles, communications, postmortem, actions, closure evidence) by month 3 onwards; audit dry run identifies zero critical gaps by month 7.
- All 28 teams transitioned to new process by month 5; all legacy alert tools decommissioned; single source of truth for incidents established and sustained.
Steps (22):
1. Executive charter, governance structure, and evidence clock
Turn the CEO email into a funded, authorized program with clear ownership and documented evidence collection for SOC 2, starting today.
- Appoint a Director of Incident Management reporting to CTO, with dotted line to COO (customer impact) and Head of Compliance (audit readiness).
- Publish a one-page charter: scope (all customer-impacting, payment-path, data-integrity incidents across 28 teams and 2 regions), decision rights (IC may freeze changes, override team preferences during incidents), and authority to mandate process participation.
- Secure annual budget for tooling, training, on-call compensation ($500–800K estimated), and resilience work. Connect funding to avoided SLA credits ($1.3M baseline).
- Establish standing Incident Management Steering Group: CTO, VP Engineering, Head of Support, Head of Compliance, one engineering manager per region. Meet monthly.
- **Start the SOC 2 evidence clock on day 1.** An audit in eight months means operating-period evidence begins now; design the process to capture evidence continuously, not retroactively.
- Make incident-response participation a documented performance expectation for all engineering managers and team leads.
- Publish timeline: working process in month 2, all 28 teams in month 5, audit-ready in month 7.
2. Baseline measurement, incident register, and evidence pack (depends on: 1)
Establish defensible baseline metrics and identify structural gaps that explain the 40% customer-detected rate and 22-minute detection time.
- Build a 12-month incident register with all 31 customer-impacting incidents: date, detection source, detection time, mitigation time, customer count, services involved, SLA credits paid, root cause class.
- Audit the current alert estate: total volume per tool, volume per team, volume per service, page-to-action ratio, top 50 noisiest rules, off-hours interruptions per engineer.
- Construct a **silent-failure register**: incidents with no internal alert fired at all. This explains the 40% customer-detected rate.
- Reconstruct the two command-ambiguity incidents minute by minute: exactly when did ownership become unclear, how long, what was the decision bottleneck.
- Survey on-call engineers (target >70% response): burden, fairness, pay expectations, escalation clarity, willingness to stay.
- Interview Support and Account Management: how do customers discover incidents, what do they complain about, how do they contact you.
- Publish the problem statement internally; retain all artifacts for SOC 2 audit evidence. This is the baseline against which all improvements are measured.
3. Control mapping and evidence architecture (depends on: 1, 2)
Design the process to generate SOC 2-compliant evidence automatically, from the first real incident, so the audit clock ticks in your favour.
- Map the new process to Trust Services Criteria CC7.1–7.5 (incident identification, response, evaluation, containment, communication), CC2.2–2.3 (authorization), CC4.1 (change management), CC3.x (information availability).
- For each control, write a one-paragraph plain-language statement, name a single owner, and specify the evidence artifact (incident record, timeline, communications log, postmortem, action tracker, training record).
- Define the **golden incident file**: one single-click export per incident containing severity, timeline, roles assigned, decisions made, communications sent, postmortem, and action items. This is the audit unit.
- Specify data retention, immutability, access control, and storage location (not laptops, not chat history that expires). Ensure evidence is searchable and organized by incident date.
- Keep a gap register with owners and dates; review fortnightly in the steering group. Identify which controls are blocked by incomplete rollout and when they unblock.
- Run an early design walkthrough with an experienced SOC 2 readiness partner inside month 1 to stress-test control design before building on it.
4. Severity and response class taxonomy (depends on: 2)
Define four severity levels and four response classes so every decision—paging, communications, postmortem, compensation—keys off a defensible rule, not a judgment call.
- **Severity by impact scope**: SEV-1 (total payment-path loss, data corruption, or >5% transaction failure for >5 min); SEV-2 (significant degradation or single region loss); SEV-3 (limited impact with workaround available); SEV-4 (internal issue or cosmetic); SEV-0 (reserved for security/regulatory/privacy events).
- **Response class** (orthogonal to severity): Availability, Performance, Data Integrity & Ledger, Security & Privacy. **Key rule: class can raise severity, never lower it.** A SEV-3 data-integrity incident gets SEV-1 response posture because integrity is not recoverable by moving faster.
- Automatic triggers: loss of one AWS region → SEV-1 or SEV-2 (class-dependent); ledger write failures → SEV-1; replication lag >10s → escalation review; payment success rate <99% for >5 min → SEV-1/2; missed settlement window → SEV-1; total external API unavailability → SEV-1.
- Who may declare: any engineer, Support, account manager (based on observed customer impact). Who may downgrade: IC only, after investigation.
- Map each level to SLA credit exposure and to customer-facing status-page state.
- Include worked examples from the last 12 months so all 28 teams recognize their own incidents in the taxonomy. Re-validate quarterly against real declarations.
5. Incident lifecycle, Triage Owner rule, and escalation policy (depends on: 4)
Eliminate the "nobody was in charge for over an hour" problem by assigning ownership the moment a page is acknowledged.
- Define lifecycle states with clear entry/exit criteria: Detected (alert fired) → Triaged (is this real and customer-impacting?) → Declared (severity assigned) → Mitigated (core issue resolved) → Resolved (all verifications done) → Postmortem (review scheduled) → Closed (action items tracked or dismissed).
- **Introduce the Triage Owner rule**: the person who acknowledges the page owns the incident until an IC is assigned or the incident is stood down. There is never an unowned gap between first page and declaration. Triage Owner's sole job: decide within 15 minutes whether this requires an IC or a direct stand-down.
- Set aggressive acknowledgement and declaration targets: page acknowledged in 5 min; triage decision (is this real?) in 15 min; severity declaration in 30 min for any customer-facing incident.
- Implement automatic escalation ladders with no human judgment required: if responder does not acknowledge in 5 min, escalate to service owner; if no ack in 10 min, escalate to team manager; if no ack in 15 min, escalate to IC on-call. Escalation is never criticized.
- Define unresponsive-team path: if a service's on-call is unreachable for 30 min, IC may direct any available engineer from any team to engage.
- For SEV-1 and SEV-2: change freeze until IC declares mitigation confirmed; IC unfreezes changes explicitly.
- Enforce one incident, one record. Timeline auto-captured from Slack channel and bridge; never written from memory later.
6. Incident roles, command structure, and decision rights (depends on: 5)
Define clear roles with one-page responsibility cards published and linked from every paging notification.
- **Incident Commander**: owns incident outcome, not the fix. Declares severity, decides escalation, approves all customer communications, freezes changes, calls responders, hands off in shifts. Non-technical ICs are acceptable; technical depth is not required.
- **Deputy IC**: assigned within 15 min of declaration; shadows IC; takes over if IC unavailable or after 4-hour shift on SEV-1. Maximum IC shift: 4 hours on SEV-1, 6 hours on SEV-2.
- **Triage Owner** (new role): owns incident from first page acknowledgement until IC takes over or stand-down decision is made. Required for all incidents.
- **Communications Lead**: owns internal Slack updates and status-page messaging; shields IC from customer contact and interruptions.
- **Scribe**: records real-time timeline with decisions, actions, and key timestamps; not responsible for fixing.
- **Subject-Matter Responders**: engineers with service context; take IC direction; report only to IC; no side channels or parallel debugging.
- **Operations Lead** (SEV-1 only): coordinates multiple responders, manages incident bridge, maintains escalation list.
- Minimum viable staffing: SEV-1 requires all roles; SEV-2 requires IC, Deputy, Comms, Scribe, SMEs; SEV-3 requires Triage Owner and IC.
- Create laminated role cards for every on-call shift location (office, home, printed in pockets).
7. Three on-call rotations: Service, Platform, and Incident Commander (depends on: 5, 6)
Directly address the "carrying a pager for another team's code" objection by making it structurally impossible.
- **Service On-Call rotation** (federated): each of the 28 teams maintains a rotation for their own services only. No engineer is paged for code their team does not own. The answer to "why am I carrying a pager?" is now simply: "for your team's code."
- **Platform Duty rotation** (centrally staffed): shared PostgreSQL cluster, Kubernetes, networking, CI/CD, observability, and incident management tooling. Nobody's product code, so it gets its own dedicated rotation. Staffed from platform teams plus volunteers from other teams; paid at premium rate.
- **Incident Commander roster** (24x7): 12–16 certified senior engineers from across all 28 teams, on one-week primary shifts with secondary backup. Covers every hour with no single point of failure and no uncovered holiday week.
- **Consequences and gates**: 16 of 28 teams have no on-call today. Each must either (a) build a Service On-Call rotation of at least 6 engineers, or (b) formally transfer service ownership to a team that will, with transfer documented and dated. No exceptions, no waivers. Unowned services are decommissioned or transferred by end of month 2.
- Merge small or low-traffic teams into shared rotations where service ownership is unclear (e.g., shared analytics, testing infrastructure).
- Enforce scheduling limits in the tooling: no engineer on-call more than 2 weeks per quarter, automatically enforced by configuration, not negotiation.
- Publish a coverage matrix for all 28 teams showing services, owners, rotation size, gaps, and monthly status.
8. On-call compensation, rest policy, and sustainability (depends on: 7)
Settle compensation before rollout, not during negotiations. Make on-call sustainable and valued.
- **Paid on-call**: effective immediately upon joining a rotation. Weekly stipend while on shift (benchmark to New York market: $600–1,000 per week per engineer), regardless of incident volume.
- **Event-based compensation**: 1.5× hourly rate for time spent mitigating out-of-hours incidents, minimum one-hour block per callout. Tracked by incident record (auto-capture from timeline).
- **Compensatory rest**: no engineer works a normal 8-hour business day after a night incident requiring >2 hours mitigation. Rest day is documented policy, not a favour granted by manager.
- **Intrusion cap**: maximum 3 unscheduled pages per week per engineer. Exceed the cap in a week and trigger an immediate review; exceed in a month and escalate to VP Engineering. Breaches are structural signal that alert quality or service stability has a problem.
- **Voluntary opt-out**: an engineer may exit a rotation; their team must hire or buy replacement coverage from paid pool at published internal rate ($X per shift). This converts culture debate into visible budget decision.
- **Amnesty policy**: incident records, near-miss reports, and false declarations are never used in performance reviews or compensation discussion. Only failure to report is a performance issue.
- **Policy publication**: publish compensation structure and effective date before any team is asked to join a rotation, and include on-call expectations in job descriptions and hiring conversations.
- **Semi-annual review**: reassess compensation and caps every six months against actual page volumes, attrition rates, and survey feedback.
9. Detection strategy: SLOs, synthetic monitoring, and customer-report intake (depends on: 4, 7)
Close the 40% customer-detected gap by monitoring customer journeys instead of infrastructure metrics.
- **SLO-based alerting**: Define SLIs and SLOs for the top 20 customer journeys (payment initiation, authorization, settlement, ledger read/write, API availability, webhook delivery, payout). Measure per region. Alert on SLO breach, not on infrastructure metric (e.g., alert on "payment success rate <99%" not "database CPU >80%").
- **Synthetic transaction monitoring**: deploy synthetic transactions from outside AWS in both regions plus a third geographic location, one-minute cadence, for all money-moving paths. These are your first alarm bell.
- **Ledger-critical signals**: PostgreSQL replication lag (target: <1s, alert >5s), connection saturation, write latency (p95), lock-wait time, transaction ID exhaustion proximity, checkpoint pressure, table bloat. These are separate alerts on shared-database health.
- **Customer-report intake** (new detection channel): Support and Account Managers can raise an incident directly in the platform. Every customer report creates an incident record automatically, and the "customer report" detection source is counted in all metrics. This is a legitimate detection method, not a failure.
- **Detection-gap rule**: whenever a customer reports an incident before internal monitoring fires, auto-create a ticket in the owning service's backlog with root cause: "Monitoring gap on [journey]."
- **Detection contract per service**: every one of the 180 services needs a named owner, at least one symptom-based alert mapped to a SLO, and a documented expected detect time (target: <5 min for payment path, <10 min for others). Published on wiki and reviewed monthly.
- **Detection drills**: run a quarterly drill per team: simulate a broken service in staging and verify it triggers a page before a human notices.
- **Resilience roadmap separation**: detection improvements do not protect against ledger corruption or multi-region failure. Fund a separate resilience roadmap to reduce shared-database blast radius and improve failover safety.
10. Alert quality standards and noise-reduction program (depends on: 9)
Cut the 3,400 monthly alerts (85% noise) to <600 with 95% signal. This is the price of admission for on-call buy-in.
- **Paging contract**: every page must satisfy all of (1) symptom-based (customer impact, not infrastructure cause), (2) actionable (linked runbook with immediate next step), (3) owned (named team responsible), (4) severity-mapped (SEV-1/2/3/4), (5) SLO-linked where applicable. **No runbook, no page.** Enforce with CI check on alert definition.
- **Separation rule**: only customer-impacting or imminent-impact signals trigger pages. Everything else becomes a dashboard entry, a ticket, or a log line. Noisy infrastructure metrics go to dashboards, not pagers.
- **Page budget per service**: no service may exceed 2 pages per on-call shift per month. Exceeding budget auto-opens a remediation ticket in the owning team's backlog (with alert-quality review assigned to tech lead).
- **Automatic suppression rules**: (1) silence alerts if service auto-recovered within 30s, (2) suppress known maintenance windows, (3) group flapping alerts (>5 in 2 min) into one page, (4) rate-limit noisy services (max 1 page per 5 min until condition clears). All suppression rules must have an expiry date; no permanent silence without a ticket.
- **Probation for new alerts**: new alert rules run as tickets only and alert to a Slack channel; after two weeks of proving actionability (every alert resulted in human action), they graduate to pager.
- **Noise sprint**: run a focused 90-day program with a public burn-down of the top 100 noisiest rules. Assign each to a named manager. Default action: fix root cause, tune threshold, or delete within 10 working days. Deletion is a legitimate successful outcome (celebrate it).
- **Correlation and deduplication**: consolidate alert sources at ingest pipeline so one outage triggering 40 alerts produces one page, not 40.
- **Alert ownership**: every alert must have an owning team and a maintenance contact. Update monthly.
11. Incident tooling consolidation and integration (depends on: 5, 10)
Replace six alert tools and ad-hoc incident records with a single source of truth that unifies paging, escalation, timeline, and audit evidence.
- **Tool selection**: choose an incident-management platform (e.g., PagerDuty, Incident.io, Opsgenie) that integrates paging schedules, escalation policies, incident records, postmortem workflow, and status-page APIs. Decision gate: month 1.
- **Event pipeline consolidation**: route all alerts from the six legacy tools into a single event pipeline that feeds the incident platform. Apply deduplication, correlation, severity/class mapping, and rate-limiting at ingest.
- **Observability integration**: connect the incident platform to your Kubernetes dashboards, PostgreSQL monitoring, distributed tracing, and logs so responders see context in one pane. Link runbooks directly into incident records.
- **Slack and bridge integration**: auto-create incident Slack channels, auto-invite roles, auto-capture timeline from channel transcript and voice-bridge recording. Timeline is not written from memory; it is auto-captured.
- **Golden incident file**: implement the export defined in S3. One click produces a complete, immutable, audit-ready PDF: severity, timeline, roles, decisions, communications, postmortem, action items, and closure evidence.
- **Dual-run period**: run both legacy and new platform in parallel for two weeks. Define rollback criteria (e.g.,
12. Escalation automation and incident lifecycle enforcement (depends on: 5, 11)
Eliminate judgment calls from the worst moments. Escalation is automatic, mechanical, and blameless.
- **Automatic escalation ladders**: page responder → if no ack in 5 min, page service owner → if no ack in 10 min, page team manager → if no ack in 15 min, page IC on-call + call them immediately (phone + SMS + Slack). No human decides to escalate; timers fire escalations.
- **Severity-based escalation tempo**: SEV-1 uses faster timers (2 min for IC on-call), SEV-2 uses moderate timers (5–10 min), SEV-3 uses slower timers (15–30 min). Configured in tooling, reviewed quarterly.
- **Dual IC rule**: if a second SEV-1 incident is detected while the first is active, immediately page and assign a separate IC. ICs never run two incidents in parallel.
- **Change freeze and rollback authority**: SEV-1 and SEV-2 trigger automatic deploy freeze. Only the IC (with CTO/VP Eng notification) may unfreeze. Freeze lifts only when IC explicitly declares mitigation confirmed and verifies no new incident symptoms for 5 min.
- **Unresponsive team escalation**: if service's on-call does not acknowledge in 30 min, IC may direct any engineer from any team (volunteers first, then rotated) to engage. This is documented and reported in monthly review (escalation = signal of rotation problem).
- **One incident, one record**: all decisions logged in the incident platform. Auto-capture from Slack, bridge, status-page updates. Timeline is the source of truth; postmortem is written from timeline, never constructed after the fact.
- **Ambiguity rule**: if two responders disagree about whether an incident should be declared, it is declared. False declarations (stand-downs within 30 min of declaration) are tracked as metrics and closed without blame.
- **Watch state**: an unconfirmed incident can live in "Watch" state for max 30 min; after that, either declare it or stand it down explicitly.
13. Internal, customer, and regulatory communications workflows (depends on: 6, 12)
Define who informs whom, in what order, via what channel, with explicit timings and pre-approved templates.
- **Internal cadence**: first update to #incidents Slack channel within 3 min of declaration (even if "Investigating"). Then updates every 5 min (SEV-1), 15 min (SEV-2), or 30 min (SEV-3), or immediately on material change (e.g., mitigation achieved, scope widened). **Comms Lead owns the update; IC must not be interrupted.**
- **Executive notification**: IC calls CTO and VP Eng within 1 min of SEV-1 declaration (not email, not Slack, call). Incident declared in Slack with severity label, IC name, and affected service. Escalation channel lead auto-pinged.
- **Customer communication channels**: status page (all 2,100 customers), direct email to affected customers (top-tier accounts and customers affected by SEV-1), account-manager calls (top 50 accounts on SEV-1).
- **Status page timings**: update within 3 min of SEV-1 declaration, 10 min of SEV-2, 30 min of SEV-3 (even if root cause unknown; use "Investigating" with next-update ETA). Updates every 5–30 min depending on severity. Always include next-update time.
- **Pre-approved templates**: draft customer-facing language for each severity and class in advance with Legal and Compliance. Templates specify impact language ("some of your transactions are delayed" not "our database failed"), workarounds if available, and next-update commitment. Never speculate on cause in customer communication.
- **Regulatory notification path**: identify incidents requiring regulator notification (NYDFS Part 500, money-transmitter rules, payment-card-network rules, securities disclosure). Build a clock matrix: event type → regulator → notification window → signer. Compliance owns all regulatory notifications (never Engineering). Pre-clear templates. Flag incidents to Compliance immediately upon declaration.
- **Account manager playbook**: contact tree for top 50 accounts, templated talking points (facts only, never speculation), escalation path if customer escalates, what to offer (service credit, technical deep-dive call).
- **Closing communication**: resolution notice, SLA credit impact, commitment date for written root-cause report, customer action required (none, or security update, etc.).
14. Status page infrastructure and customer-impact ledger (depends on: 13)
Make the status page reliable, customer-centric, and audit-ready. Track customer impact in a single durable record.
- **Status page decoupling**: host status page outside production failure domain (separate cloud, separate infrastructure, separate database). Integrate incident platform with status page so incident record drives all public updates. Status page survives total platform outage.
- **Component-to-journey mapping**: status page components map to customer journeys ("Payments", "Settlements", "Payouts", "Ledger API") not to internal services. Allow customers to subscribe to components; notify by email or webhook.
- **One-click update templates**: pre-fill status-page template with severity, impact language, next-update time, and estimated resolution. Comms Lead types minimal new info ("Root cause identified" or "Workaround available"), and updates auto-post.
- **Customer-impact ledger** (one record per incident): which customer accounts affected, which journey(s) impacted, exact start and end time of impact, estimated SLA-credit exposure. Use this single record for customer communications, credit computation, regulatory reporting, and annual review. No reconciliation of two versions of the same outage.
- **SLA credit automation**: compute credit based on duration × severity × customer tier → auto-generate customer notification → auto-post to finance system. Reconcile accrued vs. paid credits monthly and report in executive review.
- **Testing during game days**: simulate status-page outage and verify alerts continue to fire; test total region loss and confirm status page remains updated; drill runbook for manually updating status page if platform is down.
15. Postmortem policy: mandatory, blameless, three-level framework (depends on: 6, 13)
Turn postmortems from a writing exercise (11 of 64 action items closed) into the learning engine of the system.
- **Mandatory postmortems**: all SEV-0, SEV-1, and SEV-2 incidents; all SEV-3 with customer impact or repeat pattern; any near-miss IC flags; any incident where the process itself failed (IC unreachable, Comms Lead unavailable, false declaration, missed update SLA).
- **Three-level framework** (proportionate to weight): (1) lightweight async review for SEV-4 and low-impact SEV-3 (10 min template in shared doc, owner + IC review), (2) standard facilitated postmortem for SEV-2 and impactful SEV-3 (full template, facilitated by trained neutral party, published within 10 days), (3) full executive postmortem for every SEV-1 and every security incident (executive sponsor assigned, full investigation, published within 15 days, customer-facing variant prepared).
- **Fixed timeline**: draft postmortem within 5 business days, blameless review within 10 days, internal publication within 15 days.
- **Single template**: impact (who, how many, how long, financial exposure), timeline (detection through resolution), root cause (not "human error" but system condition that enabled error; what was the gap?), contributing factors (tooling, process, organization, knowledge, monitoring), what went well, what went badly, action items (≤3, rest go to reliability backlog).
- **Blameless facilitation**: train a pool of blameless postmortem facilitators (target: 10+ engineers). Require a trained, neutral facilitator for every SEV-1 and SEV-2 review. Prohibit counterfactual language ("if the engineer had"), blame language, and the phrase "human error" as a root cause.
- **Publication rule**: publish all postmortems internally by default; security review only for genuinely sensitive material (e.g., unpatched vulnerability details or customer PII in logs). Create a customer-facing root-cause report for every SEV-1, especially for regulated customers, with legal and compliance sign-off.
- **Searchability**: store postmortems in a searchable wiki or issue tracker with tags (service, class, root cause category) so teams can learn from similar incidents without repeating them.
16. Action item tracking, reliability backlog, and repeat-incident design rule (depends on: 15)
Close the loop on incident learning by enforcing verifiable, tracked action items and breaking cycles of repeat incidents.
- **Action item capping**: each postmortem generates a maximum of 3 action items. Anything beyond 3 goes into a ranked reliability backlog, not into the postmortem, to prevent overwhelming teams.
- **Action item requirements**: each item must have (1) a named human owner (not a team), (2) a due date (≤60 days, target ≤30 days), (3) a definition of done (merged code change, tested alert, audit evidence, architectural decision, new runbook, training completed) not self-reported status.
- **Single reliability backlog**: create one backlog in your engineering tracker (Jira, Linear, etc.) with mandatory label (e.g., `incident-action`), link to originating incident, and link to postmortem. Track progress weekly.
- **Closure sign-off**: Incident Commander or postmortem facilitator must sign off on closure, verifying artifact exists (code merged, alert tested in drill, runbook verified).
- **Repeat-incident rule**: if the same service or component has a second incident with the same contributing factor, **do not create another action item**. Instead, escalate immediately to an architect or tech lead and trigger a design review (not a task, a review). This breaks the cycle of repeated patches; the system needs a structure change.
- **Capacity protection**: reserve a fixed percentage of each team's sprint capacity (10–15%) for reliability work. Track unspent capacity and report to VP Engineering monthly; if a team is not spending it, work with them to identify and fix blockers.
- **Ageing and escalation**: run a weekly review of open actions; escalate anything >30 days overdue to team lead and VP Engineering. Monthly report: completion rate (target >90% within 60 days) and median action age (target <30 days).
17. Training, certification, and exercise program (depends on: 6, 13, 15, 16)
Build skills before deploying the process. Run ongoing drills so the system is tested, not guessed at.
- **Curriculum**: (1) All engineers (30-min async video): severity taxonomy, communication expectations, postmortem format, when to declare an incident, where to find runbooks. (2) On-call responders (2-hr workshop): alert tool walkthrough, playbooks by severity, escalation paths and timers, when to call manager, mobile app walkthrough, case studies from the last 12 months. (3) Incident Commanders (4-hr bootcamp + test): leadership under pressure, decision-making (severity, escalation, rollback), communicating with executives, status-page discipline, postmortem facilitation, handling ambiguity, live simulated incident (pass/fail certification). (4) Communications Leads (2-hr training): templates per severity and class, customer-communication rules (no speculation, no blame), update timings, how to shield IC, regulatory triggers.
- **IC certification**: written assessment (75% pass required) plus live simulated incident (role-play with facilitator, graded on severity declaration, escalation decisions, communication, handover). Certification valid for 12 months; recertify via annual refresher or another live sim.
- **Depth across teams**: certify at least 2 ICs per team or team group so central roster is not siloed in one group; no holiday week is uncovered.
- **Async content**: record all training videos so async teams can learn on their schedule. Create quick-reference cards (laminated, pocket-sized) for roles and playbooks; distribute to on-call locations (office, home).
- **Monthly tabletop exercises**: drawn from real incidents from the last 12 months (region loss, ledger write failure, missed settlement window, cascading failures). Facilitator describes scenario; 3–4 responders play out response (Triage Owner, IC, Comms) as if real. Run 30 min; retro for 15 min afterward.
- **Quarterly game days**: deliberately inject failures into production (database failover, status-page outage, alerting-pipeline outage, dual SEV-1 incidents). All on-call roles engage. Run 2–3 hours; measure response times, decision quality, and communication. Document findings and create action items for identified gaps.
- **Drill the process's own failure modes**: IC unreachable (on-call unavailable, phone broken), Comms Lead on PTO, two simultaneous SEV-1s, paging storm (100+ alerts), false alarm that consumes an hour. Test escalation paths, deputy takeover, and recovery.
- **New-engineer onboarding**: add incident-response module to all engineering onboarding (completion tracked, audit-ready). All engineers must complete within 30 days of joining or transferring in.
18. Metrics, dashboards, and review cadence (depends on: 2, 12, 16, 17)
Measure to prove the system works. Publish live dashboards so every engineer sees the scoreboard and the system is transparent.
- **Outcome metrics**: Median Time to Detect by source (target: <5 min internally detected, <10 min customer journeys); Median Time to Mitigate for SEV-1/2 (target: <60 min SEV-1); customer-detected incidents as % of total (target: <5%); incidents by severity (should be mostly SEV-3/4, few SEV-1); SLA credits paid (target: <$100K/year by month 12); annual credit avoidance vs. program cost.
- **Process metrics**: IC assigned within 5 min (target: >95% of incidents); page acknowledgement rate (target: >98% within 5 min); first-update timeliness (target: >95% within SLA); postmortem timeliness (target: 100% of mandatory postmortems published on time); IC roster coverage (zero uncovered hours, monitored weekly).
- **Health metrics**: alert volume and signal-to-noise ratio per team (trending toward target); off-hours pages per engineer per month (trend, cap enforcement); on-call satisfaction survey (target: >7/10); training completion by team (target: 100% within 30 days); % of services with active detection contract (target: 100%).
- **Never publish incident count as a team metric.** Reward hiding. Instead publish detection metrics (near-misses reported per team, detection gaps closed, false declarations made).
- **Live dashboards**: build dashboards visible to all engineers (not just managers) showing outcome, process, and health metrics. Auto-populate from incident platform and alert tool. Update daily. Link from Slack and internal wiki.
- **Baseline all metrics against S2 evidence pack.** Set 90-day and 12-month targets for each metric. Publish targets and progress monthly.
- **Review cadence**: (1) weekly 30-min operational review (incident by incident from prior week: what went well, what hurt, actions); (2) monthly 60-min reliability review (trends, top causes, action-item aging, alert quality per team); (3) quarterly 60-min executive review (CEO's office: customer impact, SLA credits, top five systemic causes, program ROI).
- **Quarterly process review**: what in the process wasted responder time, what confused people, what should be deleted. Solicit feedback from ICs, Comms Leads, and responders. Document changes and reasoning.
19. Pilot program with 3–4 volunteer teams (depends on: 6, 8, 11, 12, 13, 14, 15, 16, 17, 18)
Do not roll out untested to 28 teams. Run the entire process end-to-end with a small cohort using real incidents as the primary training material.
- **Team selection**: recruit 3–4 volunteers spanning criticality: one payment-path team, one ledger-adjacent team, one shared infrastructure team (platform or Kubernetes), one low-traffic team. Volunteers see early adoption and influence.
- **Full process in pilot**: new severity and class taxonomy (S4), consolidated tooling (S11), roles and Triage Owner (S5–6), three rotations (S7), escalation automation (S12), communications (S13–14), postmortems (S15), action tracking (S16), paid on-call (S8), training (S17), metrics (S18). This is not a partial test; it is the complete system.
- **Real incidents are the training**: hold a retro within 48 hours of each pilot incident (while memory is fresh). Process Owner facilitates. Discuss: what worked, what hurt, how is the runbook, is the alert tuned, did Comms template work, did roles work, was timeline auto-captured correctly. Document feedback and iterate.
- **Instrument against baseline**: measure all S18 metrics in pilot; compare results with S2 baseline. If MTTD dropped from 22 min to 8 min, that is the number you carry into the next conversation.
- **Weekly retros with pilot teams**: structured 30-min sessions reviewing policies, tooling, training, templates, and process changes. Change the written policy or tooling in response; document what changed and why so later teams see the evolution, not just the finished state.
- **Explicit exit criteria**: (1) rotation coverage sustained (≥6 engineers per rotation, zero unacknowledged pages over 3 weeks), (2) postmortems delivered on time (100% of mandatory postmortems published within 15 days), (3) action tracking established (100% of action items in backlog with owner and date), (4) metrics live (dashboards updated daily, first weekly review completed).
- **Pilot report**: document before/after numbers (MTTD, MTTR, alert noise, action completion rate, on-call satisfaction) and key process learnings. This report is the foundation for every conversation in the rollout.
20. Phased rollout sequenced by cost of failure (depends on: 16, 18, 19)
Rollout is a staged migration with readiness gates and sequencing that prioritizes visible impact, not ease.
- **Wave sequencing**: divide 28 teams into 4 waves of ~7 teams each, **ordered by incident density and customer-journey ownership** (highest-cost-of-failure teams first). Teams with the most SLA credits at stake go first; their improvement is the proof.
- **Wave spacing**: three weeks between waves. This gives each wave time to stabilize and find problems before the next cohort joins.
- **Readiness checklist per team**: (1) all services mapped and owned (no unowned services), (2) alerts cleaned to paging contract (runbook linked, severity mapped), (3) playbooks updated and tested in staging, (4) rotation staffed to ≥6 engineers, (5) team completes training module, (6) manager briefed on policy, (7) on-call compensation in effect.
- **Gate review before each wave**: Process Owner holds gate review with target teams. Move unready teams to next wave with a dated remediation plan. No exceptions, no waivers; readiness is non-negotiable.
- **Wave champion**: assign a named engineer per wave to champion the rollout, answer questions, escalate issues to Process Owner. Champions are not representatives; they are advocates and feedback collectors.
- **Communication cadence**: weekly all-hands or newsletter for 4 weeks before each wave. Explain why (owned-code-owned-pager rule, paid on-call, SLA credit savings). Use pilot numbers. Answer FAQs. Announce champion and escalation path.
- **First incident under new process**: hold a retro within 48 hours. Feed accepted process changes back through change control.
- **Retire legacy tools and processes**: at end of each wave, retire legacy alert tools, informal escalation lists, ad-hoc status-page process. No parallel processes running for >3 weeks; this prevents confusion and half-learning.
- **Sequence to avoid audit collision**: ensure no team is rolling out in the same week as the audit dry run (S21).
21. SOC 2 dry run and evidence review (depends on: 3, 20)
Convert a good working process into a provable one. Test control evidence a few months before auditors arrive, when you can still fix gaps.
- **Dry run timing**: run 6 weeks before audit window (around month 7 of this program).
- **Scope**: sample 10–15 real incidents from pilot and early rollout waves. For each incident, verify evidence artifact exists and is complete: incident record, timeline (auto-captured), severity and class declaration, roles assigned and logged, communications log (Slack + status page), postmortem (if mandatory), action items in tracker with due dates, action completion evidence (code, alert test, training record).
- **Control walkthrough**: walk through each control statement from S3 with a checklist. Is the evidence artifact present? Is it immutable? Is it searchable? Is retention adequate? Is access logged?
- **Gap remediation**: for every gap found, estimate time to fix and prioritize by audit risk. Anything risking a qualified opinion (e.g., missing postmortem, no timeline evidence) must be fixed before the audit. Test the remediation against a new incident or a resample.
- **Interview readiness**: brief 10–15 engineers who may be interviewed by auditors (ICs, Comms Leads, Process Owner, team managers). Ask them to describe the process as they actually practice it, not as written. Listen for confusion or gaps in understanding. Correct them.
- **Auditor package preparation**: assemble process documentation, sample incident records (5–10 complete golden files), training records, on-call schedules, alert quality metrics, action tracker register, and status-page archive. Organize by control. Create a table of contents and index.
- **Single audit liaison**: designate Process Owner or a small dedicated compliance person as sole point of contact for audit requests. Prevents requests scattering across 28 teams.
- **Rehearsal**: conduct mock interview with an IC and a Comms Lead. Auditors ask tough questions under pressure ("How do you know the timeline is accurate?", "What happens when both ICs are unavailable?", "Show me how you proved the alert was actionable."). Practice answering.
22. Standing governance and continuous improvement (depends on: 20, 21)
Lock in durable improvement. The classic post-audit failure is the process freezing and then decaying. This step prevents that.
- **Standing Incident Management Council**: chaired by Process Owner, monthly meetings, attendees: CTO, VP Engineering, Head of Support, Head of Compliance, one engineering manager per region, one IC, one team manager from recent wave. Agenda: metrics review, policy changes, gaps from recent incidents, escalation for contentious issues.
- **Change control mandate**: give Process Owner documented authority to change severity taxonomy, response classes, roles, communications timings, and compensation policy. Any change requires: written justification, steering group approval (monthly cadence), and documented effective date before implementation. This prevents silent drift and ensures changes are deliberate.
- **Quarterly validation**: re-validate severity and response class taxonomy against all incidents from the prior quarter. Ask: did our taxonomy correctly predict response posture? Did we misclassify? Update taxonomy if patterns emerge.
- **Annual metric re-baselining**: every 12 months, re-run measurements from S2 (alert census, incident register) to reset targets. System should improve; targets should tighten.
- **Resilience roadmap separation**: fund a distinct architectural or platform roadmap for incident prevention (reduce shared-database blast radius, multi-region failover, deploy safety, observability investments). Better incident response does not protect a single-ledger corruption or unplanned failover. These are separate problems.
- **Quarterly executive report**: CTO and VP Eng report to CEO/CFO on metric set (detection time, mitigation time, SLA credits, customer-detected %), top five systemic causes of incidents, program cost vs. credit avoidance, and strategic architecture changes in flight.
- **Public backlog of improvement ideas**: teams and engineers propose process improvements via Slack or wiki. Process Owner reviews quarterly and implements accepted ideas (e.g., "add a dashboard for detection gaps", "update postmortem template"). Publish what changed and why.
- **Celebration and learning**: share wins publicly each quarter ("We reduced MTTD from 22 min to 5 min", "Customer-detected incidents down 80%", "$800K SLA credits avoided"). Refresh training and tabletop program annually and immediately after any SEV-1 to keep the system sharp and responsive to new scenarios.
--- PROPOSAL 2 (agent deepseek-flash_refine_2, deepseek/deepseek-flash) ---
Estimated complexity: high
Success metrics: - Median time to detect falls from 22 minutes to under 8 minutes by month 6, and under 5 minutes for money-moving journeys by month 9.
- Share of customer-impacting incidents first detected by customers falls from 40% to below 10% by month 9.
- Median time to mitigate for SEV1 falls from 3 h 10 min to under 60 minutes by month 9.
- 100% of SEV1 and SEV2 incidents have a named Triage Owner within 5 minutes and a named Incident Commander within 15 minutes, from month 4.
- Zero incidents with unclear command authority lasting more than 15 minutes, measured monthly from month 4.
- Monthly page volume falls from 3,400 to under 600, with a false-positive rate below 15%, by month 5.
- No service exceeds two pages per on-call shift for three consecutive months by month 6.
- 100% of the 180 services have a named owner, a detection contract and at least one symptom-based alert by month 7.
- All 28 teams have a documented rotation, a service ownership map and at least one trained on-call engineer by month 6.
- 100% of on-call shifts are paid under a published policy from month 2, with zero on-call-attributed voluntary attrition by month 6.
- On-call satisfaction scores 7 out of 10 or better in quarterly surveys from month 6.
- The IC roster holds at least 12 certified ICs covering 24x7 with no uncovered week from month 4.
- 100% of SEV0, SEV1 and SEV2 postmortems are published internally within 15 business days from month 5.
- Action items closed within 60 days rise from 17% to above 90%, with a median action age under 30 days, by month 6.
- Status page first update is posted within 30 minutes on at least 95% of SEV1 incidents from month 4.
- Zero missed regulatory notification windows on any incident requiring notification.
- The customer-impact record is complete for 100% of customer-impacting incidents from month 5, with SLA credits reconciled monthly.
- Annual SLA credits paid fall at least 50% from the $1.3M baseline within 12 months.
- Customer-impacting incidents fall at least 40% year over year from the 31-incident baseline.
- Review cadence is sustained: weekly operational review in at least 90% of weeks, 12 of 12 monthly reliability reviews, 4 of 4 quarterly executive reviews.
- Every page has a recorded disposition — fixed, tuned or deleted — within 10 working days, from month 4.
- SOC 2 Type II is passed with zero findings related to incident response.
- 100% of new engineers complete the incident-response onboarding module within 30 days of joining.
Steps (24):
1. Mandate, one owner, and the evidence clock
This step turns the CEO's email into a funded programme with a single accountable owner, and it starts the SOC 2 clock on day one.
- Appoint a Director of Incident Management reporting to the CTO, with a dotted line to the COO for customer and SLA matters and to the Head of Compliance for audit readiness.
- Publish a one-page charter: scope covers every customer-impacting, money-moving, security and data-integrity event across all 28 teams and both regions.
- Grant explicit authority to declare an incident, set severity, freeze deploys, page any engineer in the company, and approve customer messaging.
- Fund the envelope up front: tooling, training and drill time, on-call compensation, and a small programme team, roughly $500–700K a year against $1.3M in credits paid.
- State the return plainly to the steering group: credits avoided, churn avoided, and audit findings avoided.
**Start the evidence clock now.** A SOC 2 Type II report covers a period of operating effectiveness, so evidence cannot be backfilled; an audit in eight months means the operating period begins in the first weeks of this programme.
- Stand up a fortnightly steering group of the CTO, VP Engineering, Head of Support, Head of Compliance and two engineering managers.
- Open a programme risk register with the top risks, owners and review dates, and revisit it at every steering group.
- Make participation in the incident process a documented performance expectation for every engineering manager, not an optional extra.
2. Baseline evidence pack and cost-of-downtime model (depends on: 1)
You cannot demonstrate improvement without a defensible starting point, and you cannot win the argument about noise without numbers.
- Build a 12-month incident register and re-classify all 31 customer-impacting incidents by detection source, duration, customers affected and credits paid.
- Build the silent-failure register: incidents where no internal alert fired at all, which is the number that explains the 40% customer-detected rate.
- Run an alert census per tool, per team and per service: volume, page-to-action ratio, off-hours interruptions, and the 50 noisiest rules with a named owner each.
- Survey every on-call engineer and manager on burden, fairness, escalation clarity and pay expectations, targeting above 70% response.
- Interview Support, Account Management and Sales about how customers learn of incidents and what they are promised.
- Reconstruct the two command-ambiguity incidents minute by minute to find exactly where ownership lapsed.
- Build a cost-of-downtime model: dollars per minute of impact per customer journey, used later to sequence teams and justify funding.
- Publish the pack internally as the problem statement, and retain every artifact as management-review evidence for the audit.
3. SOC 2 control mapping and evidence architecture (depends on: 1)
Most programmes leave compliance to the end; this one maps controls in month one, because the mapping decides what the process must capture from day one.
- Map the process to the Trust Services Criteria for incident identification, evaluation, response, containment, recovery and communication (CC7.1–CC7.5, CC2.2, CC2.3, CC4.1, CC3.x).
- Write each control in plain language, with one named owner and its evidence artifact.
- Define the golden incident file: one auto-assembled, single-click export per incident containing timeline, roles, severity, communications, postmortem, actions and closure proof.
- Set retention, storage location and immutability so no control depends on a laptop, a private channel, or chat history that expires.
- Keep a gap register with owners and dates, reviewed fortnightly by the steering group.
- Meet the auditor's readiness team inside the first 90 days to test the control design before anything is built on top of it.
4. Severity and class taxonomy with the trigger matrix (depends on: 2)
Severity decides how loud the response is; class decides which playbook runs and what must never be traded away.
- Define severity in customer and money terms: customers affected, dollars at risk, regions lost, ledger integrity, settlement delay, data loss and SLA exposure.
- Use SEV0 for security, privacy or regulatory events; SEV1 for total or material loss of a payment path; SEV2 for degradation or single-region loss; SEV3 for limited impact with a workaround; SEV4 for internal-only issues and near-misses.
- Define response classes: Availability, Performance, Data integrity and Ledger, Security and Privacy, Third-party dependency, Settlement and Financial, Process failure.
**Class can raise the response but never lower it.** A SEV2 data-integrity incident gets SEV1 posture, because integrity failures are not recoverable by moving faster.
- Specify automatic triggers: loss of one AWS region, ledger write failures, replication lag above threshold, a missed settlement cut-off, payment success rate below threshold for five minutes.
- State who may declare — any engineer, Support agent or account manager — and who may downgrade: the Incident Commander alone.
- Map every level to its SLA credit exposure and to the customer-visible status page state.
- Include worked examples from the last 12 months so each of the 28 teams recognises its own incidents, and revalidate the taxonomy quarterly against real declarations.
5. Day-one operating rules and the minimum viable process (depends on: 1, 4)
The full process will take months; the first useful version must be live in two weeks using the tools that already exist.
- Publish ten day-one rules that need no procurement: a named owner within five minutes, one channel per incident, one register entry per incident, one person speaking to customers.
- Apply the ambiguity rule: if two responders disagree about whether this is an incident, it is an incident.
- Make declaring free: a false alarm is closed as a false declaration, tracked as a metric, and never criticised.
- Require a register entry within 24 hours for every customer-impacting incident, even a minimal one.
- Ban silent incidents: if we know, the customer hears it from us rather than from their own reconciliation.
- Run the first 30 days on manual command, with a rotating duty Incident Commander drawn from the 12 teams that already have on-call.
- Hold a 15-minute daily incident stand-up during month one to catch friction while it is still fresh.
6. Roles, command structure and the no-unowned-minute rule (depends on: 4)
The two hour-long command failures did not happen at declaration; they happened in the gap before it, when an alert had fired and nobody owned it.
- Publish one-page role cards: Triage Owner, Incident Commander, Deputy IC, Operations Lead, Internal Comms Lead, Customer Comms Lead, Scribe, Subject-Matter Responders, and Executive Sponsor for SEV1 only.
- Introduce the Triage Owner rule: whoever acknowledges the page owns the incident until an IC takes over or the incident is stood down.
**The IC owns the incident, not the fix, and does not debug.** An IC who starts troubleshooting has abandoned command.
- Give the IC explicit decision rights: declare and escalate severity, freeze changes, halt deploys, approve customer messaging, and pull in any engineer.
- Define minimum viable staffing per severity: SEV1 fills every role; SEV2 staffs IC, scribe, comms and responders; SEV3 staffs an IC and a scribe.
- Set handover discipline: four-hour maximum IC shifts on SEV1 with a written handover, and a deputy named within 15 minutes of declaration.
- Set responder behaviour: one channel, one bridge, no side channels, and every request phrased with a named owner and a time.
- Link the role cards from every paging notification so they are one tap away at 3 AM.
7. Lifecycle, declaration and escalation policy (depends on: 4, 6)
This step defines the mechanical path from an alert to a declared incident and back to normal service, removing judgment calls from the worst moments.
- Define states with entry and exit criteria: Detected, Triaged, Declared, Mitigated, Resolved, Postmortem, Closed, plus a Watch state with a hard 30-minute timer.
- Set targets: page acknowledged within five minutes, triage decision within fifteen, declaration within thirty for anything customer-facing.
- Build automatic escalation ladders with timers at every layer: responder, service owner, team manager, IC on-call, VP Engineering.
- Define the unresponsive-team path: fifteen minutes escalates to the team manager, thirty minutes to the IC, who may then direct any available engineer.
- Define change freeze and rollback authority during SEV1 and SEV2, with the single named condition that lifts the freeze.
- Enforce one incident, one record, with the timeline captured automatically from the channel and bridge rather than written from memory afterwards.
- Add the two-simultaneous-incidents rule: a second IC is designated, and no IC ever runs two incidents at once.
- Test every escalation path weekly with synthetic pages, and adjust the timings after the first month of real operation.
8. On-call architecture across 28 teams (depends on: 4, 6)
The objection is that engineers will not carry a pager for another team's code; the answer is to build rotations that make the objection structurally impossible.
- Run a Service On-Call rotation per team, covering only that team's own services.
- Run a Platform Duty rotation for genuinely shared infrastructure: the PostgreSQL ledger cluster, Kubernetes, networking, CI/CD and observability.
- Run a central Incident Commander roster of 12–16 certified senior engineers on one-week shifts with a primary and a secondary.
**State the consequence honestly.** Sixteen of 28 teams have no rotation today; each must build one or formally transfer service ownership to a team that has one, with the transfer dated and recorded.
- Merge small or low-traffic teams into shared rotations so no rotation has fewer than six engineers.
- Cap load in the scheduling tool: no engineer on call more than two weeks per quarter, enforced by configuration rather than negotiation.
- Publish a coverage matrix of all 28 teams showing services, owners, rotation size and gaps, reviewed monthly.
- Write on-call participation into engineering job levels and hiring criteria, so the commitment is set before someone joins.
9. Compensation, rest and the economics of opting out (depends on: 8)
Unpaid on-call is the most cited reason for resistance, so settle compensation before rollout, not during it.
- Move to paid on-call: a per-shift stipend benchmarked to the New York market, agreed with HR and Finance, published with an effective date before any team is asked to join a rotation.
- Pay for callouts at 1.5× the hourly rate for time actually spent mitigating, with a minimum block per interruption.
- Provide documented compensatory rest: no normal working day after a night incident, and the rest day is policy rather than a favour granted by a manager.
- Cap intrusion with a maximum number of off-hours pages per shift, and trigger a mandatory review whenever the cap is exceeded.
**Allow opt-out, but put a price on it.** An engineer may step out of a rotation, and their team buys coverage from the paid pool at a published internal rate, which turns a cultural argument into a visible budget decision.
- Publish amnesty: incident records, near-misses and false declarations are never used in performance reviews; only failure to report is a performance issue.
- Protect against gaming: pay attaches to shifts and callouts, never to incident counts.
- Check the New York labour, overtime and tax treatment with Legal and Finance before announcing, and review the policy every six months against real page volumes, attrition and survey results.
10. Detection strategy: journeys, synthetic signals and customer-report intake (depends on: 4)
Customers detected 40% of incidents first, which makes detection the highest-leverage business problem in this programme.
- Define SLIs and SLOs for the top 20 customer journeys, measured per region: payment initiation, authorisation, settlement, ledger read and write, API availability, webhook delivery and payout.
- Alert on symptoms against those SLOs, not on cause-based infrastructure metrics.
- Add synthetic transaction monitoring from outside AWS in both regions plus a third location, on a one-minute cadence for money-moving paths.
- Add ledger and shared-PostgreSQL signals: replication lag, connection saturation, write latency, lock waits, transaction ID exhaustion, checkpoint pressure and bloat.
- Open a customer-report intake so Support and account managers can raise an incident directly, and count that path as a detection source in all reporting.
- Apply the detection-gap rule: whenever a customer reports first, a detection-gap ticket is opened automatically and owned by the responsible team.
- Publish a detection contract per service — owner, at least one symptom alert, documented expected detect time — for all 180 services.
- Run a detection drill per team: break something in staging and see whether it pages before a human notices.
11. Paging contract and the noise-reduction programme (depends on: 2, 10)
3,400 alerts a month at 85% noise is the reason engineers resent the pager, and fixing it is the price of admission for everything else in this plan.
- Publish the paging contract: every page must be symptom-based, actionable, owned, mapped to a severity and class, and linked to a runbook.
**No runbook, no page**, enforced by a CI check on the alert definition itself.
- Separate paging from ticketing: only customer-impacting or imminent-impact signals may wake a human.
- Set a page budget per team and per service, with a remediation ticket opened automatically, owned by the engineering manager, when the budget is breached.
- Put new alerts on two-week probation as ticket-only until they have proved actionable.
- Define the default action for a noisy alert: fix, tune or delete within ten working days, with deletion treated as a legitimate and celebrated outcome and the deleted count published.
- Deduplicate and correlate at the ingest pipeline so one root cause produces one page instead of forty.
- Require an expiry date on every silencing rule and temporary threshold, so suppression cannot quietly become permanent blindness.
- Run a 90-day noise sprint with a public burn-down of the top 100 noisiest rules, each owned by a named manager.
12. Incident tooling consolidation and the golden incident file (depends on: 3, 6, 10, 11)
Six alerting tools and no single incident record are structural causes of the 22-minute detection and the three-hour mitigation.
- Choose one incident platform for paging, schedules, escalation policies, incident records, postmortem workflow and status-page integration, and choose it within four weeks.
- Consolidate the six alert sources into a single event pipeline feeding that platform, with deduplication, correlation and severity and class mapping applied at ingest.
- Integrate observability, runbooks, chat and voice bridges so responders see dashboards and runbooks inside the incident record, and the timeline is captured automatically.
- Implement the golden incident file export defined in S3, so audit evidence is one click rather than a reconstruction.
- Run a dual-run period alongside the legacy tools with a defined rollback, then switch the legacy tools off on a published cutover date.
- Host the status page outside the production failure domain so it survives a total platform outage, and prove that in a game day.
- Make the platform usable from a phone: an Incident Commander must be able to run a SEV1 from a mobile device at 3 AM.
13. Internal, customer and regulator communications (depends on: 6, 7, 12)
Today the status page is written by whoever is around; this step replaces improvisation with a clock, a named owner and pre-cleared templates.
- Set the internal cadence: first update within 15 minutes of declaration, then every 30 minutes for SEV1 and hourly for SEV2, whether or not there is progress.
- Never let an employee learn of an incident from the status page: internal communication leads, external follows.
- Set the customer cadence: status page updated within 30 minutes of a SEV1 and 60 minutes of a SEV2, with a no-new-information update still mandatory.
- Define the channel hierarchy: status page for all 2,100 customers, direct email to affected customers on SEV1, and named account-manager calls for the top 50 accounts.
- Pre-approve templates per severity and class with Legal and Compliance, each carrying its own next-update time.
- Forbid speculation: customer messages never guess at cause, never assign blame, and never commit to a root cause before the postmortem.
- Define the closing communication: resolution notice, how SLA credits will be handled, and the committed date for the written report.
- Build a regulator clock matrix covering event type, regulator, notification window, signer and the shortest applicable clock, including NYDFS Part 500, money-transmitter and banking notifications, breach notification, card-network rules and public-company disclosure.
**The regulatory clock starts at awareness, not at root cause.** Route every notification through Compliance, never Engineering, and pre-clear the templates.
14. Customer trust workstream and the SLA credit ledger (depends on: 13)
The $1.3M in credits is a symptom of having no single record of customer impact, and the CEO's inbox is a symptom of customers learning things late.
- Maintain one durable customer-impact record per incident: which customers, which journeys, from when to when, and the estimated credit.
- Use that one record for communications, credit computation, regulatory reporting and the annual review, so nobody reconciles two versions of the same outage.
- Automate the credit path end to end — computation, approval, customer notification and finance treatment — so credits stop being a manual scramble three weeks later.
- Reconcile accrued against paid credits monthly and report the result in the executive review.
- Give the status page a named product owner and map its components to customer journeys, not to internal services.
- Send a CTO-signed reliability note to the top 50 accounts and publish a quarterly reliability report to all customers.
- Give account managers a script of the facts they may state, the speculation they may not, and a path for customer escalations.
- Track credit avoidance against programme cost, so the funding case stays a number rather than an argument.
15. Postmortems: mandatory set, three levels, blameless by design (depends on: 6)
Postmortems currently happen for some incidents, in various formats; this step makes them mandatory where they matter and light where they do not.
- Make postmortems mandatory for every SEV0, SEV1 and SEV2, every SEV3 with customer impact or a repeat pattern, every near-miss the IC flags, and every incident where the process itself failed.
- Use three levels so the ritual matches the weight: a lightweight async review for SEV3 and SEV4, a facilitated postmortem for SEV2, and a full review with an executive sponsor for SEV1.
- Set deadlines: draft within five business days, blameless review within ten, internal publication within fifteen.
- Adopt one template: impact, timeline, detection, response, contributing factors across tooling, process, organisation and human factors, what went well, what went badly, and actions.
- Train a pool of blameless facilitators and require a trained one for every SEV1 review, never the IC.
**Ban blame language in the template and ban "human error" as a root cause.** The question is always what system condition made the error possible.
- Publish all postmortems internally by default, with security review only for genuinely sensitive material.
- Produce a customer-facing root-cause report for SEV1 incidents, especially for regulated and top-tier accounts.
16. Action items: capped, verifiable, with the repeat-incident rule (depends on: 15)
Eleven of 64 action items closed is not a tracking problem; it is a generation problem, because the process produces more actions than the organisation can absorb.
- Cap each postmortem at three action items, with anything beyond that going into a ranked reliability backlog.
- Require every action to have a named human owner, a due date, and a definition of done that is an artifact: a merged change, an alert that fired in a drill, or a test that fails without the fix.
- Prohibit self-reported closure; closure requires the artifact and sign-off by the process owner or the Incident Commander.
- Keep one reliability backlog in the engineering tracker with a mandatory label and a link back to the originating incident.
- Reserve a fixed share of each team's sprint capacity for reliability work, with unspent capacity visible to the vice presidents.
- Apply the repeat-incident rule: a second incident with the same contributing factor triggers a design review owned by an architect, not another action item.
- Run a weekly ageing review of open actions and escalate anything more than 30 days overdue to the VP of Engineering.
- Report action completion rate and median action age monthly, by team.
17. Metrics, dashboards and the review cadence (depends on: 2, 4, 16)
Define what good looks like, then measure it in a way that rewards reporting incidents rather than hiding them.
- Define outcome metrics: time to detect by source, time to mitigate by severity and class, share of incidents detected by customers, incidents by severity and class, SLA credits paid, and customer-impact minutes.
- Define process metrics, because the process itself has SLOs: IC assigned within five minutes, page acknowledgement rate, first-update timeliness, update-cadence adherence, postmortem timeliness, action closure rate, action median age, and IC roster coverage.
- Define health metrics: alert volume and noise ratio per team, off-hours pages per engineer, on-call satisfaction, training completion, and share of services with a detection contract.
**Never publish incident count as a team metric.** It rewards hiding incidents; publish reporting metrics instead — near-misses filed, detection gaps closed, false declarations made.
- Publish live dashboards visible to every engineer, refreshed daily, with every metric baselined against the S2 evidence pack and given a 90-day and a 12-month target.
- Run a 30-minute weekly operational review incident by incident, a monthly reliability review of trends and actions, and a quarterly executive review with the CEO's office.
- Hold a quarterly review of the process itself: what wasted time, what confused responders, and what should be deleted.
- End every review with decisions and named owners, never with numbers alone.
18. Training, certification and the drill programme (depends on: 6, 7, 13, 15)
A process that lives only on a wiki page fails on the first real page, so skills are built and tested before they are needed.
- Build a practical curriculum: how to be on call, how to triage and declare, how to run an incident as IC, how to communicate, and how to write a postmortem.
- Require certification before joining the IC roster: a written assessment plus a live simulated incident.
- Certify at least two ICs per team group so the central roster has depth across all 28 teams and no holiday week is left uncovered.
- Train communications leads separately on templates, cadences, customer language and the regulatory rules.
- Run monthly tabletops drawn from the last 12 months of real incidents, including region loss, ledger write failure and a missed settlement cut-off.
- Run quarterly game days with deliberately injected failure, including PostgreSQL failover, status-page outage and alerting-pipeline outage.
- Drill the process's own failure modes, not just technical ones: IC unreachable, comms lead on leave, two simultaneous SEV1s, a paging storm, and a false alarm that burns an hour.
- Audit the process for single points of failure: who alone can perform each critical task, and what happens in their holiday week.
- Keep a mandatory onboarding module for every engineer joining or transferring in, with audit-ready completion records.
19. Pilot with three to four teams, using real incidents (depends on: 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 18)
Do not roll out to 28 teams untested, and do not rely on drills alone when real incidents are better training material.
- Recruit three to four volunteer teams spanning criticality: a payment-path team, a ledger-adjacent team, the platform team that owns the shared PostgreSQL cluster, and a low-traffic team.
- Run the entire process end to end in the pilot: severity and class, roles, escalation, communications, postmortems, action tracking, and paid on-call.
- Treat real incidents during the pilot as the primary training material, and hold a retro within 48 hours of each one, run by the process owner while the friction is fresh.
- Instrument the pilot against the baseline and publish before-and-after numbers.
- Hold weekly retrospectives with the pilot teams, and change the written policies, the tooling and the training in response, documenting what changed and why.
- Set explicit exit criteria: rotation coverage achieved, zero unacknowledged pages over two weeks, postmortems delivered on time, and actions tracked to closure.
- Produce a pilot report that every later rollout conversation starts from.
20. Phased rollout to 28 teams, sequenced by cost of failure (depends on: 16, 18, 19)
Rollout is a staged migration with readiness gates, not an email announcement, and the sequencing matters more than the schedule.
- Split the 28 teams into four waves of roughly seven, ordered by incident density and customer-journey ownership: the teams whose incidents cost the most SLA credits go first, because their improvement is the visible proof.
- Leave three weeks between waves and define a per-team readiness checklist: services mapped and owned, alerts cleaned to the paging contract, runbooks written, rotation staffed to at least six engineers, training complete, manager briefed, compensation in effect.
- Hold a gate review with the process owner before each team joins, and move an unready team to the next wave with a dated remediation plan rather than granting an exception.
- Give each wave a named champion and explain the why using the pilot's numbers.
- Handle resistance directly and early: publish the own-your-code-own-your-pager rule and the paid on-call mechanics before each wave, never after.
- Hold a retro within 48 hours of each wave's first incident under the new process, and push accepted changes through change control.
- Retire legacy tools, informal escalation lists and the ad-hoc status page process at the end of each wave on a published cutover date, so two processes never run in parallel for long.
- Sequence the waves so no team faces both a rollout and the audit dry run in the same week.
21. Audit dry run and evidence review (depends on: 3, 20)
This step converts a good process into a provable one, about six weeks before the auditors arrive.
- Sample real incidents from the pilot and the early waves against each control's evidence requirements.
- Remediate every gap found and re-test the remediated control against the same sample, prioritising anything that risks a qualified opinion.
- Brief every engineer who may be interviewed so they can describe the process as they actually practise it, not as it is written.
- Prepare the auditor package: process documentation, sample incident records, paging logs, communication logs, training records, on-call schedules, the action register and alert quality metrics.
- Designate one audit liaison and a small evidence-request team, so requests do not land on all 28 teams at once.
- Rehearse the walkthrough with an Incident Commander and a communications lead, because auditors probe realism under pressure.
- Keep the audit liaison and the process owner as close to the same person as possible, so accountability for the control is also accountability for the evidence.
22. Standing governance and process ownership (depends on: 20, 21)
The classic post-audit failure is that the process freezes and then decays, so ownership has to outlive the programme.
- Establish a standing Incident Management Council chaired by the process owner, meeting monthly with engineering, support, compliance and product representation.
- Give the process owner a documented mandate to change standards, with formal change control for any change to severity, classes, roles, communication timings or compensation.
- Revalidate the severity and class taxonomy quarterly against real declarations, and re-baseline the entire metric set annually.
- Report quarterly to the executive team on the metric set and the top five systemic causes of incidents.
- Feed incident themes into architecture review, release engineering and hiring, so the programme improves the system and not only the response.
- Keep a public backlog of process improvement ideas so any engineer can propose a change, and act on it every quarter.
23. Resilience track for the shared ledger and multi-region blast radius (depends on: 1, 17)
Better incident management cannot save a single shared ledger during a corruption event, so this runs as a funded engineering track beside the process, not inside it.
- Reduce ledger blast radius: read replicas for read paths, connection-pool isolation per service domain, statement timeouts, write throttles and per-tenant limits.
- Build a ledger protection mode: an automatic tripwire that degrades to read-only or queues writes before corruption can spread.
- Rehearse multi-region failover with the ledger as the binding constraint, and set a measured recovery time objective per region.
- Test restore and point-in-time recovery on a schedule, and publish the timings.
- Tighten deploy safety: progressive rollout, automatic rollback on SLO burn, and change-freeze windows around settlement.
- Track a small number of resilience bets with named owners and quarterly milestones, reviewed in the same executive review as the incident metrics.
24. Culture, incentives and continuous improvement (depends on: 20, 22)
Incident management is a cultural change before it is a process change, and the incentives decide whether reporting goes up or down.
- Restate the amnesty at every rollout wave: reporting, near-misses and false declarations are safe; hiding is not.
- Recognise detection improvements, deleted alerts, closed detection gaps and well-written postmortems publicly, and reward the engineers who do them.
- Correct blame language in leadership communication on the spot, including inside the executive review.
- Watch for pager fatigue with real data, and take an engineer off rotation as soon as the intrusion cap breaks.
- Run a quarterly process retro with the whole engineering organisation and publish what changed as a result.
- Refresh the training and tabletop programme annually and immediately after any SEV1.
- Share wins publicly each quarter — faster detection, fewer customer-detected incidents, credits avoided — to hold momentum after the audit.
--- PROPOSAL 3 (agent qwen3.8-flash_refine_3, alibaba/qwen3.8-flash) ---
Estimated complexity: high
Success metrics: - Median Time to Detect (MTTD) < 5 minutes.
- Median Time to Mitigate (MTTR) < 60 minutes for SEV1.
- >90% of incidents detected internally (vs. 40% previously).
- SLA credits reduced to <$100K annually.
- Monthly alert volume < 600 with >95% signal-to-noise ratio.
- Postmortem action item completion rate > 80%.
- Zero incidents with command ambiguity > 15 minutes.
- 100% of SEV1/2 postmortems published within 15 days.
- Status page updates within 30 minutes for 95% of SEV1.
- 100% of teams integrated into on-call rotations.
- On-call satisfaction score > 7/10.
- 40+ certified Incident Commanders available 24x7.
- SOC 2 Type II audit passed with zero findings.
Steps (22):
1. Executive Mandate & SOC 2 Evidence Clock
Establish clear ownership and funding to transform the CEO's concern into a governed program with immediate compliance implications.
2. Baseline Measurement & Problem Statement (depends on: 1)
Quantify the current state to build a defensible problem statement and establish the data foundation for improvement tracking.
3. Severity & Class Taxonomy (depends on: 2)
Define unambiguous incident levels based on business impact to trigger specific response protocols.
4. Incident Command Roles & Structure (depends on: 3)
Standardize roles and decision rights to eliminate command ambiguity during active incidents.
5. Tooling Consolidation & Integration (depends on: 1)
Deploy a unified incident management platform to replace fragmented alerting tools and create a single source of truth.
6. Detection Strategy & SLO Implementation (depends on: 3)
Shift monitoring focus to customer-impacting symptoms to detect issues before customers do.
7. Alert Quality Standards & Noise Reduction (depends on: 5, 6)
Implement strict alerting rules to reduce noise and restore on-call trust.
8. On-Call Architecture & Coverage Model (depends on: 4)
Design a federated, team-owned on-call structure supported by central command coverage.
9. On-Call Compensation & Wellbeing Policy (depends on: 8)
Introduce paid, sustainable on-call rotations to remove resistance and mitigate burnout.
10. Playbooks & Communication Templates (depends on: 3, 4)
Create standardized playbooks and templates for each severity level to reduce cognitive load.
11. Escalation & Incident Lifecycle Policy (depends on: 4, 5)
Define the mechanical progression from alert to resolution, including automatic escalation paths.
12. Communication Workflows (Internal & External) (depends on: 4, 10)
Establish timed protocols for internal, customer, and regulatory communications.
13. Status Page & Customer Notification Automation (depends on: 5, 12)
Integrate incident management with public status pages to ensure timely and accurate updates.
14. Postmortem Policy & Blameless Framework (depends on: 3)
Mandate and standardize postmortems to focus on system improvement rather than individual error.
15. Action Item Tracking & Governance (depends on: 14)
Ensure postmortem action items are tracked, owned, and completed to prevent recurrence.
16. Metrics, Dashboards & Review Cadence (depends on: 2, 3)
Define and monitor key performance indicators to validate the effectiveness of the incident process.
17. Training, Certification & Drills (depends on: 4, 10, 12)
Build practical skills and readiness for incident response across all roles.
18. Pilot Program with Volunteer Teams (depends on: 5, 7, 9, 11, 13, 15, 17)
Test the end-to-end process on a small scale to refine procedures before full rollout.
19. Phased Rollout to All Teams (depends on: 18)
Deploy the process to all 28 teams in manageable waves to minimize disruption and manage resistance.
20. SOC 2 Control Mapping & Evidence Framework (depends on: 1, 3)
Align the incident management process with SOC 2 requirements to ensure audit readiness.
21. SOC 2 Dry Run & Audit Support (depends on: 19, 20)
Verify evidence sufficiency and prepare the team for the external audit.
22. Governance & Continuous Improvement (depends on: 19, 21)
Establish standing governance to sustain the process and drive ongoing improvements.
You have just assessed the final proposals, blind to the vote; your first choice was proposal 2. Now evaluate the PROCESS itself, not the plans: a judgement, not a narration. Your answer has these parts:
- "process_evaluation": a Markdown list with one point per question: whether the convergence between agents was earned by better arguments or was mere imitation; whether the agents criticised each other's ideas or just copied them; whether anyone questioned the task's premises; what was lost along the way. Never empty.
- "process_issues": a list of concrete problems you observed in the process (empty only if you saw none).
- "suggestions": a list of concrete changes that would make the process produce a better plan.
[BRIEF]
[SYSTEM]
You are an expert reviewer of multi-agent planning processes.
Several LLM agents drafted plans for a task, refined them over a number of rounds while seeing each other's proposals, and finally voted for the best one.
Be exhaustive but precise: name concrete steps, ideas and metrics, never generalities. Judge plans by their fitness for the task as stated, their realism, their completeness, the soundness of their order and dependencies, how measurable their success is and how they handle things going wrong.
You are an impartial evaluator, not a chronicler: assess the proposals and the process on their merits, never rationalise what happened or assume that the outcome was right.
After your analysis, answer in the requested structure.
Every text field you write will be read by a busy person who skims. Make it easy to skim: short sentences and short paragraphs; when you name several things, prefer a list to a paragraph, with sub-items when an item has parts, but keep a single fact as a sentence; lead with the point and then the evidence; name proposals and steps by number (P2, step 4); no preamble, no repetition of the question, no closing summary; bold at most one key phrase per item or paragraph. Text fields accept Markdown: a blank line between paragraphs, "- " for lists, **bold**.
[HUMAN]
Task given to the agents: "A B2B payments platform in New York (260 engineers in 28 teams, 2,100 customers, $4B processed a month, a 99.95% SLA with service credits) runs about 180 services on Kubernetes across two AWS regions, with a shared PostgreSQL cluster for the ledger. In the last 12 months: 31 customer-impacting incidents, median time to detect 22 minutes (customers detected 40% of them first), median time to mitigate 3 h 10 min, $1.3M paid in SLA credits, two incidents in which nobody was sure who was in charge for over an hour, and a CEO email about "outages we hear about from clients". On-call exists in 12 of the 28 teams, unpaid, fed by alerts from six different tools — 3,400 alerts a month, 85% of them noise; there is no severity scale, status-page updates are written by whoever is around, and postmortems happen for some incidents, in various formats, with action items that are rarely tracked (11 of 64 closed). A SOC 2 Type II audit in eight months will test incident response. Engineers push back against "carrying a pager for other teams' code".
Define the incident management process: severity levels and what each one triggers; roles (incident commander, communications, scribe, subject-matter responders) and how they are staffed 24x7 across 28 teams; the on-call structure, rotations, compensation and the rules for alert quality; detection and escalation paths; internal and customer communications (status page, account managers, regulators when required) with their timings; postmortems (when mandatory, format, blameless review, ownership and tracking of actions); the metrics and reviews that show whether it works; and how it is introduced across the teams without waiting for the audit."
You have analysed a deliberation between 3 agents over 3 rounds and its vote. Here is everything you wrote, in order:
ROUNDS:
Round 0: All three agents converge on the same skeleton — severity matrix → roles → paid on-call → tool consolidation → comms SLAs → blameless postmortems with tracked actions → metrics → phased rollout — but differ sharply in depth and in the on-call model. P2 is the most operationally concrete (baseline evidence pack, SLO-based detection, federated ownership, SOC 2 control mapping plus dry run); P1 is exhaustive but serializes training and drills after full rollout; P3 is the compact version and leaves audit evidence and detection engineering thin.
shared: **Four-tier severity scale as the keystone**: every plan keys paging, comms cadence and postmortem obligation off SEV1–SEV4 (P1 S1, P2 S3, P3 S2).
shared: Same role set — IC who commands but does not debug, comms lead, scribe, SME responders — with explicit decision rights (P1 S2, P2 S4, P3 S4).
shared: Collapse the six alerting tools into one platform and attack the 85% noise with actionable-alert standards, dedup and per-team caps (P1 S5–S6, P2 S7/S14, P3 S5).
shared: Paid on-call plus mandatory blameless postmortems for SEV1/SEV2 with action items tracked in Jira and reviewed by leadership; pilot first, then waves, not a big bang (P1 S4/S12/S13/S19, P2 S9/S12/S13/S18, P3 S3/S8/S9/S11).
differences: **On-call architecture**: P2 S8 is federated — one owning team per service, "own your code, own your pager", minimum rotation of six, 16 teams must build a rotation or transfer ownership, plus a central 24x7 IC roster. P3 S3 does the opposite, merging 12 rotations into one unified pool covering all 28 teams, which directly recreates the "pager for other teams' code" grievance it claims to solve. P1 S3 sits in between: dedicated IC pool of 4–6 plus per-team SME on-call.
differences: **Detection engineering**: only P2 S6 defines SLIs/SLOs for the top 20 customer journeys, symptom-based alerting, external synthetics in three locations, ledger-specific PostgreSQL signals (replication lag, TXID exhaustion) and a separate blast-radius resilience track. P3 S6 has synthetics only; P1 treats detection mostly as alert routing and never defines SLOs.
differences: **SOC 2 rigour**: P2 devotes S19–S20 to Trust Services Criteria mapping, a named owner and evidence artifact per control, retention rules and a dry run six weeks out. P1 S16 has a control mapping plus a month-6 mock audit. P3 has one bullet in S11 ("simulate auditor questions") — far too light for an eight-month Type II window.
differences: **Ordering flaws differ**: P1 S17 depends on all sixteen prior steps and pushes training (S18), rollout (S19) and the first drill (S20) to the end, so nobody practises before going live. P2 front-loads a 12-month baseline register (S2) that every metric later hangs on. P3 S11 defers comp and role rollout to months 3–4 while alert cleanup starts month 1, and P3 S9 ties action-item completion to performance reviews, which cuts against its own blameless charter (S8).
Proposal 1: 21 steps covering severity tiers, roles, a hybrid IC-pool/per-team SME rotation, concrete comp numbers ($500–1,000/week stipend, 1.5x callback, comp day), single alert tool with suppression rules, per-severity playbooks, status-page timings (SEV1 in 3 minutes), postmortem and action-item tracking, metrics and governance cadence. Rollout is a single mega-step (S17) depending on all sixteen predecessors, followed by training, waves and drills. Targets are the most aggressive of the three: MTTD <8 min, credits <$100k, 0 findings.
Proposal 2: Starts with a funded charter and a named process owner (S1), then a 12-month incident and alert baseline used as both problem statement and audit evidence (S2). Builds severity with a SEV0 for security/regulatory events, federated per-team on-call with a central IC roster, SLO- and synthetic-based detection, a 90-day noise sprint with a two-pages-per-shift budget, comms timing SLAs with pre-approved legal templates and a status page outside the failure domain, evidence-based action closure, then pilot, four gated waves and a SOC 2 dry run. Ends with a standing council and a separate resilience roadmap.
Proposal 3: Twelve steps: executive charter, severity matrix tied to transaction failure rates, a single mandatory paid on-call pool for all 28 teams, RACI roles plus a Rapid Response Team for the shared PostgreSQL cluster, tool consolidation with runbook-or-no-page hygiene, synthetic transactions and auto-escalation, tight status-page timings (SEV1 first update in 5 minutes), mandatory 5-day postmortems, Jira-integrated action tracking, game days, and a month-by-month rollout to month 7. Metrics target MTTD <5 min and 95% internal detection.
Round 1: P2's round-0 architecture became the de facto template: P1 and P3 both rebuilt their plans on it step-for-step, while P2 itself deepened its version with genuinely new mechanics (Triage Owner, severity x class, priced opt-out, three-action cap, regulator clock matrix, SOC 2 evidence clock). The round converged strongly on structure; what still separates the plans is depth of incident mechanics, realism of comms timings and where audit work sits in the sequence.
differences: **Command mechanics between alert and declaration.** P2 alone closes the gap that caused the two hour-long ownership failures: Triage Owner rule (step 5), Watch state with a 30-minute timer, "declaring is free", the ambiguity rule and the two-simultaneous-SEV-1 rule (step 6). P1 has an escalation ladder (step 9) but no owner-from-first-ack rule; P3 has no escalation or lifecycle step at all and dropped its round-0 5-minute auto-escalation.
differences: **On-call shape.** P2 builds three rotations including a paid Platform Duty for the shared PostgreSQL/Kubernetes estate (step 8) and prices opting out against a paid pool (step 9). P1 (steps 10–11) and P3 (step 5) stop at federated team rotations plus a central IC roster, leaving shared infrastructure ownership implicit.
differences: **Customer communication timings and money.** P1 demands a status-page update within 3 minutes and SEV-1 updates every 5 minutes (step 13) — contradicting its own success metric of 30 minutes. P2 uses 30/60 minutes (step 12) plus a regulator clock matrix naming NYDFS Part 500 and a customer-impact ledger driving SLA credit automation (step 13). P3 uses 15/30 minutes (step 9) with no credit process at all.
differences: **Where audit work sits.** P2 maps controls and defines the "golden incident file" in month one (step 3) on the argument that Type II evidence cannot be backfilled. P1 places control mapping at step 21, P3 at step 12 after postmortems. P3 also schedules game days (step 16) and the metrics dashboard (step 17) only after full rollout, so nothing is drilled or measured during the pilot.
influences: P1 and P3 rebuilt on P2's round-0 skeleton almost wholesale: charter (P2 s1), baseline evidence pack (s2), severity trigger matrix (s3), role cards (s4), SLO/synthetic detection (s6), "no runbook, no page" and the 90-day noise sprint (s7), federated on-call with a six-engineer floor (s8), paid on-call with compensatory rest (s9), pilot then four waves with readiness gates (s17, s18), control mapping and dry run (s19, s20).
influences: P1 kept only one structural idea of its own: severity playbooks with decision trees (its round-0 s9, now step 12); everything else in its 23 steps mirrors P2's ordering.
influences: P2 took P3's ROI framing (P3 s12, credit avoidance vs programme cost) into its step 13, and P1's mobile-pager requirement (P1 s5/s8) into step 11 ("run a SEV-1 from a phone at 3am").
influences: Nobody adopted P1's round-0 per-incident bonus (s4) — P2 explicitly bans pay attached to incident counts (step 9) — and nobody adopted P3's tying of action-item completion to performance reviews (s9), which P2 contradicts with a published amnesty.
influences: P1 carried over P2's round-0 metric of 40+ certified ICs, which P2 itself cut to 12–16 this round as more realistic for 260 engineers.
Proposal 1 (improved): P1 abandoned its generic round-0 framework and adopted P2's structure nearly step-for-step, gaining a charter, baseline evidence pack, SLO-based detection, a paging contract, a federated on-call model and a pilot-then-waves rollout. It kept its own useful severity playbooks step. Residual weaknesses are internal inconsistencies in timings and severity definitions.
Proposal 2 (improved): P2 kept its 21-step shape but added several mechanisms that close real gaps rather than restating policy: the SOC 2 evidence clock, severity x class, the Triage Owner rule, three rotations, a priced opt-out, an action-item cap and a regulator clock matrix. It also made its own targets more realistic.
Proposal 3 (improved): P3 grew from 12 to 20 steps by adopting P2's skeleton — charter, baseline, detection SLOs, compensation, control mapping, dry run, culture — which fills most of the prompt's requirements it previously skipped. It remains the thinnest plan in mechanics, dropped its own escalation automation, and its dependency order pushes drills and metrics past full rollout.
Round 2: P1 absorbed almost the entire distinctive vocabulary of P2's round-1 plan (Triage Owner, severity×class, three rotations, golden incident file, capped action items), producing two near-twin heavyweight plans; P2 added the genuinely new ideas of the round (a two-week minimum viable process, a funded ledger-resilience track, a cost-of-downtime model). P3 went the other way and collapsed into 22 bare titles with no content, losing everything that made it assessable.
differences: **Bridging the gap before tooling exists**: P2 S5 defines ten day-one rules, a manual duty-IC rotation drawn from the 12 teams that already have on-call, and a daily 15-minute stand-up for month one. P1 has no interim process between the charter (S1) and platform selection (S11); P3 has none either.
differences: **Prevention as a funded track**: P2 S23 is a standalone engineering roadmap with concrete bets (ledger read-only tripwire, connection-pool isolation, PITR restore tests with published timings, rollback on SLO burn). P1 keeps resilience as two bullets inside S9 and S22; P3 does not mention it.
differences: **Communication clocks**: P1 S13 demands a status-page update within 3 minutes of SEV-1 and internal updates every 5 minutes; P2 S13 sets 15 minutes internal first, 30 minutes to status page, with a mandatory no-news update. P1's 3-minute rule contradicts its own success metric of 30 minutes for 95% of SEV-1s.
differences: **Level of specification**: P1 and P2 give thresholds, timers, dollar ranges and dates throughout; P3 R2 gives only step titles and a one-line rationale each — no severity definitions, no timings, no compensation mechanics, no dates on any metric.
influences: P1 took nearly all of P2's round-1 signature ideas: Triage Owner (P2 S5→P1 S5/S6), severity×class with "class can raise, never lower" (P2 S4→P1 S4), the three-rotation model (P2 S8→P1 S7), priced opt-out and amnesty (P2 S9→P1 S8), golden incident file and month-one control mapping (P2 S1/S3→P1 S1/S3), three capped action items and the repeat-incident design review (P2 S15→P1 S16).
influences: P2 took P1's weekly synthetic-page testing of escalation ladders (P1 R1 S9→P2 S7) and P1's status-page-component-to-customer-journey mapping plus a named status-page owner (P1 R1 S14→P2 S14).
influences: P2 took P3's culture and change-management step (P3 R1 S18) and turned it into S24: on-the-spot correction of blame language, pager-fatigue monitoring, public recognition for deleted alerts.
influences: P3 took P2's "evidence clock" framing into its S1 title but nothing else of substance; it adopted no new mechanisms this round.
influences: Nobody adopted P3's error budgets triggering feature freezes (P3 R1 S8) or its rule that a SEV-1 cannot close until high-priority actions are done (P3 R1 S11) — P1 S16 and P2 S16 instead cap actions at three and track them in a separate reliability backlog.
Proposal 1 (improved): P1 rewrote itself around P2's round-1 mechanisms while keeping its own depth. New steps 4, 5, 7, 12, 14, 15, 20, 21 replace vaguer round-1 equivalents, and detection, escalation and postmortem policy are now far more concrete. A truncated step 11 and a few internal contradictions are the cost.
Proposal 2 (improved): P2 kept its round-1 architecture intact and added the three things it was missing: an immediately usable interim process, a funded prevention track, and an explicit culture step. Compliance sequencing, metrics and communications are essentially unchanged and were already strong.
Proposal 3 (worsened): P3 discarded all step content and submitted 22 titles with a single sentence of rationale each. Every operational detail it had in round 1 — severity thresholds, status-page timings, compensation mechanics, wave schedule, training curriculum — is gone, and several metrics were loosened or stripped of dates.
OUTCOME (the final round, ranked blind):
The final round offers two near-identical heavyweight plans (P1 and P2) built on the same architecture — evidence clock and SOC 2 control mapping in month one, baseline register, severity×class taxonomy, Triage Owner rule, three on-call rotations (Service/Platform/IC), paid on-call with priced opt-out, SLO and synthetic detection, paging contract, capped and verifiable postmortem actions, pilot then four gated waves, dry run, standing council — plus P3, which submitted 22 step titles with no content. P2 is distinguished by a two-week minimum viable process, a funded ledger-resilience track, a cost-of-downtime model and internally consistent metrics; P1 is distinguished by sharper numeric thresholds but carries several self-contradictions and one truncated step. P3 is not executable as written.
Ranking: 2, 1, 3
- **P2 over P1** on three things that matter operationally: a two-week minimum viable process (S5) so the organisation is covered during the months P1 leaves uncovered; a funded resilience track (S23) that gives the 60-minute SEV-1 target a mechanism instead of a hope; and internally consistent, dated metrics. P1's 3-minute status-page rule contradicts its own 30-minute metric, its alert target (<400 vs <600) and MTTM target (45 vs 60 min) disagree with its own steps, and step 11 is cut off mid-sentence.
- P1 beats P2 on numeric thresholds (severity triggers, comp figures, escalation tempo, training pass marks) but those are easier to fill in later than the structural gaps P2 closes.
- **P1 far above P3**: P1 is executable today; P3 is a table of contents. P3 also carries an unrealistic metric (40+ certified ICs) and lost every operational detail it had in earlier rounds.
Better than every initial proposal: True
- **Gained: an interim process.** The best final plan makes something work in two weeks (day-one rules, manual duty IC, register entry within 24 hours) instead of waiting for procurement.
- **Gained: the ownership gap is named and closed.** The Triage Owner rule, the Watch state with a 30-minute timer, the ambiguity rule and the two-simultaneous-SEV-1 rule target the specific failure described in the brief.
- **Gained: compliance sequencing.** Control mapping, the golden incident file and retention rules moved to month one on the argument that Type II evidence covers an operating period and cannot be backfilled.
- **Gained: prevention as a funded track** (ledger tripwire, connection-pool isolation, PITR restore timings, rollback on SLO burn) — the only credible route from a 3h10 mitigation to under an hour.
- **Gained: honest economics.** Priced opt-out, a cost-of-downtime model, credit avoidance tracked against programme cost, and pay detached from incident counts.
- **Gained: metric hygiene.** Incident count banned as a team metric; targets dated and baselined against a 12-month register plus a silent-failure register.
- **Lost: per-severity playbooks with decision trees** (initial P1 S9) survive only as templates and training in the final plans.
- **Lost: plan diversity.** Two of the three finals are near-twins, and the third was hollowed out to titles; no alternative on-call architecture or staged scope was argued at the end.
- **Not gained: a budget that adds up.** Both heavyweight finals keep a $500–800K envelope that ~30 paid rotations alone would exhaust.
PROCESS:
- **Convergence was half-earned, half-imitation.** P2's round-0 plan really was the strongest on the merits (baseline evidence pack, SLO/synthetic detection, federated ownership, TSC mapping plus dry run), so migrating toward it was rational. But the migration happened by wholesale vocabulary transfer — Triage Owner, severity×class, golden incident file, three rotations, priced opt-out, three-action cap, no-runbook-no-page — with no agent ever stating *why* that mechanism beats what it replaced. P1 abandoned its own hybrid IC-pool model and per-incident bonus without a single sentence of rebuttal.
- **Criticism existed only as silent non-adoption.** The only visible judgements were implicit: P2 banning pay tied to incident counts (killing P1's bonus), nobody adopting P3's action-items-in-performance-reviews (which contradicted its own blameless charter), P2 cutting its own "40+ certified ICs" to 12–16. No agent ever named a rival's flaw and argued against it. Consequently obvious defects survived three rounds: **P1's status-page rule of 3 minutes contradicts its own success metric of 30 minutes for 95% of SEV-1s**, and P1's final step 11 is literally truncated mid-sentence ("Define rollback criteria (e.g.,").
- **Premises were barely questioned.** The one real challenge is P2 S23: incident process cannot save a single shared PostgreSQL ledger, so fund a resilience track (read-only tripwire, connection-pool isolation, PITR timings) alongside. Nobody questioned whether MTTD <5 min is attainable, whether a 99.95% SLA is compatible with one shared ledger cluster, or when the SOC 2 Type II *observation window* actually opens relative to the "audit in eight months" — P2 invokes an "evidence clock" but no agent checked the date.
- **Nobody audited the numbers, and the numbers do not hold.** 28 service rotations with primary and secondary is roughly 56 engineers on call at any hour; at the proposed $600–1,000/week stipend that alone is ~$1.8–2.9M/year, before Platform Duty and the IC roster — three to five times P1's ($500–800K) and P2's ($500–700K) stated envelopes, and above the $1.3M in credits being used to justify the programme. Similarly, 260 engineers over 28 teams is ~9 per team, so the six-engineer rotation floor plus a two-weeks-per-quarter cap is arithmetically tight; "merge small teams" was asserted, never modelled.
- **What was lost: diversity, then content.** By round 2 there was effectively one plan with two echoes, so the vote adjudicated nothing. P3 destroyed its own work — round 2 is 22 titles with all severity thresholds, comms timings, compensation mechanics, wave schedule and curriculum deleted, and it still carries the stale "40+ certified ICs" metric P2 had already retired. Dropped without argument: P3's error budgets triggering feature freezes, P3's rule that a SEV-1 cannot close with open high-priority actions, P1's per-severity playbooks with decision trees (diluted), and any serious alternative to the federated on-call model.
- **Pressure was purely additive.** P2 grew 21→24 steps, P1 21→22 with far denser bullets; no round asked what to delete. A plan that celebrates deleting alerts never pruned itself, and nobody assessed whether one Director plus four pilot teams can absorb 24 concurrent workstreams inside eight months.
Problems: No critique mechanism: agents never named a rival's flaw or defended a rejected idea; disagreement was expressed only by silently not copying.; P1's internal contradiction survived all three rounds — status-page update within 3 minutes of SEV-1 (S13) versus its own success metric of 30 minutes for 95% of SEV-1s — and so did SEV-1 internal updates every 5 minutes, which is unrunnable for a 3h-class incident.; No output quality gate: P1's final step 11 is cut off mid-sentence, and P3's round-2 submission is 22 bare titles with zero operational content, a strict regression from its round-1 plan.; Cross-step arithmetic never checked: the on-call compensation budget ($500–800K) is 3–5x too small for ~56 engineers on call at $600–1,000/week, and the six-engineer rotation floor was never reconciled with ~9 engineers per team.; Metric-to-step date consistency never validated: e.g. P2 promises 100% paid on-call shifts from month 2 while 16 teams still have no rotation and waves run to month 6.; Premature monoculture: after round 1 all three plans shared one skeleton, so rounds 2's vote had no genuine alternative to choose between — the outcome ratified round-0 seniority, not argued superiority.; No end-to-end stress test of any plan against a concrete scenario (ledger corruption during settlement, IC unreachable, status page down), even though all plans prescribe exactly such drills for the organisation.; Unbounded growth: success-metric lists reached 24 bullets (P2) and step counts 24; no round rewarded simplification, feasibility under a single programme owner, or a critical path to the audit date.; The SOC 2 premise went unexamined: nobody established when the Type II observation period starts, which determines whether "audit-ready by month 7" is a real plan or a hope.; Copy without freshness checks: P3 carried P2's abandoned "40+ certified ICs" target into the final round after P2 had corrected it to 12–16.
Suggestions: Require an explicit objections block each round: every agent names at least two concrete defects in each rival plan (quoting the step) plus the fix, and must justify in one sentence any mechanism it adopts from another plan.; Assign a numbers-auditor role for one round: recompute compensation cost from rotation headcount, rotation size against 260/28 engineers, IC roster hours, and check every success-metric date against the step that delivers it.; Force a scenario walkthrough: each plan must narrate minute-by-minute a SEV-1 ledger corruption at 02:00 on a settlement day with the IC unreachable and the status page down — gaps like the missing pre-declaration owner surface immediately.; Add a pruning round with a hard budget (e.g. max 15 steps, max 12 success metrics): agents must delete and justify what survives, countering additive drift.; Validate submissions mechanically: reject steps with no acceptance criteria, owner or date, and reject any round whose content is shorter than the agent's previous round unless the deletion is argued.; Mandate divergence for at least one round: assign one agent to defend a centralised responder pool, one to defend a minimal 90-day version aimed only at the audit and the noise, so the final vote compares real alternatives rather than near-twins.; Require one premise challenge per agent per round (e.g. is MTTD <5 min reachable on a shared ledger; is the SLA renegotiable; has the Type II window already opened), scored as part of the plan's merit.; Run an automated consistency check on each plan between success metrics and step text (timings, targets, headcounts) and publish the discrepancies to all agents before the next round.
THE VOTE, CONFRONTED WITH YOUR RANKING:
The analyst agrees with the vote. The vote lands on P2, my first choice, and the two P2 voters cite exactly the structural reasons I did: the two-week minimum viable process (S5) that covers the gap before full rollout, the funded resilience track aimed at the shared PostgreSQL ledger, and the explicit handling of the "pager for other teams' code" resistance. The dissenting vote for P1 rests on step count and specificity (22 steps, numeric thresholds, control mapping in step 3) but never tests P1's internal consistency
contradictory 3-minute vs 30-minute status-page timing, <400 vs <600 alert target, 45 vs 60 min MTTM, and step 11 truncated mid-sentence
which looks like a length-and-confidence bias. No voter ranked P3 highly, consistent with my view that it degenerated into a table of contents with an implausible 40+ certified IC metric.
Now write the brief a busy reader will read first, and often only: "brief", a Markdown list of at most eight items, each one sentence or two, covering in this order: what the outcome is (your first choice, the vote, whether the plan beats the initial proposals); what happened in the rounds (genuine convergence or imitation, criticism or copying, who drove it); the two or three problems that matter most; the two or three changes that would help most. Name proposals and steps by number. Nothing new: only what your analysis already says.
[ROUND 0]
{"round_summary": "All three agents converge on the same skeleton — severity matrix → roles → paid on-call → tool consolidation → comms SLAs → blameless postmortems with tracked actions → metrics → phased rollout — but differ sharply in depth and in the on-call model. P2 is the most operationally concrete (baseline evidence pack, SLO-based detection, federated ownership, SOC 2 control mapping plus dry run); P1 is exhaustive but serializes training and drills after full rollout; P3 is the compact version and leaves audit evidence and detection engineering thin.", "shared": ["**Four-tier severity scale as the keystone**: every plan keys paging, comms cadence and postmortem obligation off SEV1–SEV4 (P1 S1, P2 S3, P3 S2).", "Same role set — IC who commands but does not debug, comms lead, scribe, SME responders — with explicit decision rights (P1 S2, P2 S4, P3 S4).", "Collapse the six alerting tools into one platform and attack the 85% noise with actionable-alert standards, dedup and per-team caps (P1 S5–S6, P2 S7/S14, P3 S5).", "Paid on-call plus mandatory blameless postmortems for SEV1/SEV2 with action items tracked in Jira and reviewed by leadership; pilot first, then waves, not a big bang (P1 S4/S12/S13/S19, P2 S9/S12/S13/S18, P3 S3/S8/S9/S11)."], "differences": ["**On-call architecture**: P2 S8 is federated — one owning team per service, \"own your code, own your pager\", minimum rotation of six, 16 teams must build a rotation or transfer ownership, plus a central 24x7 IC roster. P3 S3 does the opposite, merging 12 rotations into one unified pool covering all 28 teams, which directly recreates the \"pager for other teams' code\" grievance it claims to solve. P1 S3 sits in between: dedicated IC pool of 4–6 plus per-team SME on-call.", "**Detection engineering**: only P2 S6 defines SLIs/SLOs for the top 20 customer journeys, symptom-based alerting, external synthetics in three locations, ledger-specific PostgreSQL signals (replication lag, TXID exhaustion) and a separate blast-radius resilience track. P3 S6 has synthetics only; P1 treats detection mostly as alert routing and never defines SLOs.", "**SOC 2 rigour**: P2 devotes S19–S20 to Trust Services Criteria mapping, a named owner and evidence artifact per control, retention rules and a dry run six weeks out. P1 S16 has a control mapping plus a month-6 mock audit. P3 has one bullet in S11 (\"simulate auditor questions\") — far too light for an eight-month Type II window.", "**Ordering flaws differ**: P1 S17 depends on all sixteen prior steps and pushes training (S18), rollout (S19) and the first drill (S20) to the end, so nobody practises before going live. P2 front-loads a 12-month baseline register (S2) that every metric later hangs on. P3 S11 defers comp and role rollout to months 3–4 while alert cleanup starts month 1, and P3 S9 ties action-item completion to performance reviews, which cuts against its own blameless charter (S8)."], "proposals": [{"proposal": 1, "summary": "21 steps covering severity tiers, roles, a hybrid IC-pool/per-team SME rotation, concrete comp numbers ($500–1,000/week stipend, 1.5x callback, comp day), single alert tool with suppression rules, per-severity playbooks, status-page timings (SEV1 in 3 minutes), postmortem and action-item tracking, metrics and governance cadence. Rollout is a single mega-step (S17) depending on all sixteen predecessors, followed by training, waves and drills. Targets are the most aggressive of the three: MTTD <8 min, credits <$100k, 0 findings.", "approach": "Exhaustive checklist with concrete dollar figures and clocks; weak sequencing at the end"}, {"proposal": 2, "summary": "Starts with a funded charter and a named process owner (S1), then a 12-month incident and alert baseline used as both problem statement and audit evidence (S2). Builds severity with a SEV0 for security/regulatory events, federated per-team on-call with a central IC roster, SLO- and synthetic-based detection, a 90-day noise sprint with a two-pages-per-shift budget, comms timing SLAs with pre-approved legal templates and a status page outside the failure domain, evidence-based action closure, then pilot, four gated waves and a SOC 2 dry run. Ends with a standing council and a separate resilience roadmap.", "approach": "Programme-management rigour: baseline, gates, evidence, governance"}, {"proposal": 3, "summary": "Twelve steps: executive charter, severity matrix tied to transaction failure rates, a single mandatory paid on-call pool for all 28 teams, RACI roles plus a Rapid Response Team for the shared PostgreSQL cluster, tool consolidation with runbook-or-no-page hygiene, synthetic transactions and auto-escalation, tight status-page timings (SEV1 first update in 5 minutes), mandatory 5-day postmortems, Jira-integrated action tracking, game days, and a month-by-month rollout to month 7. Metrics target MTTD <5 min and 95% internal detection.", "approach": "Lean, fast-moving version; centralised pool and light on audit evidence"}]}
[ROUND 1]
{"round_summary": "P2's round-0 architecture became the de facto template: P1 and P3 both rebuilt their plans on it step-for-step, while P2 itself deepened its version with genuinely new mechanics (Triage Owner, severity x class, priced opt-out, three-action cap, regulator clock matrix, SOC 2 evidence clock). The round converged strongly on structure; what still separates the plans is depth of incident mechanics, realism of comms timings and where audit work sits in the sequence.", "converging": true, "differences": ["**Command mechanics between alert and declaration.** P2 alone closes the gap that caused the two hour-long ownership failures: Triage Owner rule (step 5), Watch state with a 30-minute timer, \"declaring is free\", the ambiguity rule and the two-simultaneous-SEV-1 rule (step 6). P1 has an escalation ladder (step 9) but no owner-from-first-ack rule; P3 has no escalation or lifecycle step at all and dropped its round-0 5-minute auto-escalation.", "**On-call shape.** P2 builds three rotations including a paid Platform Duty for the shared PostgreSQL/Kubernetes estate (step 8) and prices opting out against a paid pool (step 9). P1 (steps 10–11) and P3 (step 5) stop at federated team rotations plus a central IC roster, leaving shared infrastructure ownership implicit.", "**Customer communication timings and money.** P1 demands a status-page update within 3 minutes and SEV-1 updates every 5 minutes (step 13) — contradicting its own success metric of 30 minutes. P2 uses 30/60 minutes (step 12) plus a regulator clock matrix naming NYDFS Part 500 and a customer-impact ledger driving SLA credit automation (step 13). P3 uses 15/30 minutes (step 9) with no credit process at all.", "**Where audit work sits.** P2 maps controls and defines the \"golden incident file\" in month one (step 3) on the argument that Type II evidence cannot be backfilled. P1 places control mapping at step 21, P3 at step 12 after postmortems. P3 also schedules game days (step 16) and the metrics dashboard (step 17) only after full rollout, so nothing is drilled or measured during the pilot."], "influences": ["P1 and P3 rebuilt on P2's round-0 skeleton almost wholesale: charter (P2 s1), baseline evidence pack (s2), severity trigger matrix (s3), role cards (s4), SLO/synthetic detection (s6), \"no runbook, no page\" and the 90-day noise sprint (s7), federated on-call with a six-engineer floor (s8), paid on-call with compensatory rest (s9), pilot then four waves with readiness gates (s17, s18), control mapping and dry run (s19, s20).", "P1 kept only one structural idea of its own: severity playbooks with decision trees (its round-0 s9, now step 12); everything else in its 23 steps mirrors P2's ordering.", "P2 took P3's ROI framing (P3 s12, credit avoidance vs programme cost) into its step 13, and P1's mobile-pager requirement (P1 s5/s8) into step 11 (\"run a SEV-1 from a phone at 3am\").", "Nobody adopted P1's round-0 per-incident bonus (s4) — P2 explicitly bans pay attached to incident counts (step 9) — and nobody adopted P3's tying of action-item completion to performance reviews (s9), which P2 contradicts with a published amnesty.", "P1 carried over P2's round-0 metric of 40+ certified ICs, which P2 itself cut to 12–16 this round as more realistic for 260 engineers."], "proposals": [{"proposal": 1, "assessment": "improved", "what_changed": "P1 abandoned its generic round-0 framework and adopted P2's structure nearly step-for-step, gaining a charter, baseline evidence pack, SLO-based detection, a paging contract, a federated on-call model and a pilot-then-waves rollout. It kept its own useful severity playbooks step. Residual weaknesses are internal inconsistencies in timings and severity definitions.", "improvements": ["Added step 1 (executive mandate, named program owner, $400–600K budget) and step 2 (12-month incident register, alert census, survey) — the round-0 plan started at severity definitions with no baseline.", "Added step 7 detection strategy: SLOs for 20 customer journeys, external synthetics in three locations, PostgreSQL-specific signals (replication lag, txid exhaustion), detection contract per service — directly attacks the 40% customer-detected figure that round-0 only listed as a metric.", "Step 8 replaces vague noise rules with a paging contract (\"no runbook, no page\"), page budgets, expiry dates on silences and a 90-day burn-down of the top 100 rules.", "Step 16 now requires artifact-based closure of action items and a repeat-incident design review, instead of round-0's self-reported Jira tracking.", "Metrics gained dates and process/health tiers (first-update timeliness, action median age, IC roster coverage)."], "regressions": ["Step 3 announces \"a special SEV0 for security/regulatory events\" then never defines it — a dangling category borrowed from P2 without its content.", "Comms timings contradict the success metrics: step 13 promises a status-page update within 3 minutes on SEV-1, the metric list says 30 minutes on ≥95%.", "Step 8's \"no service may exceed two pages per on-call shift per month\" garbles P2's budget into an unmeasurable unit.", "Target of 40+ certified ICs by month 4 is implausible for a 260-engineer org and inflates the training load in step 18.", "Step 19 (pilot) carries 12 dependencies, so almost nothing can be validated early; SOC 2 control mapping lands at step 21, losing P2's point that Type II evidence must accumulate from month one."], "taken": [{"from_proposal": 2, "steps": [1, 2], "what": "Programme charter with a single accountable owner, steering group, and a 12-month baseline evidence pack.", "why": "Became P1 steps 1–2, giving the plan a funded mandate and defensible before/after numbers it previously lacked."}, {"from_proposal": 2, "steps": [6, 7], "what": "Symptom-based SLO alerting on customer journeys plus the alert-quality standard and noise sprint.", "why": "Adopted as steps 7 and 8, replacing round-0's suppression-rule-only approach to noise."}, {"from_proposal": 2, "steps": [8, 9], "what": "Federated ownership (\"no team paged for code it does not own\"), six-engineer rotation floor, central IC roster, paid on-call with documented rest and opt-out.", "why": "Became steps 10–11, resolving the pager-resistance problem structurally rather than by stipend alone."}, {"from_proposal": 2, "steps": [11, 13, 19, 20], "what": "Status page hosted outside the production failure domain, evidence-based action closure, SOC 2 control mapping and dry run.", "why": "Adopted as steps 14, 16, 21 and 22, adding audit provability and communications survivability."}, {"from_proposal": 3, "steps": [5], "what": "A hard cap on alert volume whose breach triggers a mandatory alert-quality review.", "why": "Wired into step 8 as the enforcement mechanism behind the noise budget."}], "rejected": [{"from_proposal": 2, "steps": [10], "what": "Status-page first update within 30 minutes of a SEV-1.", "why": "P1 step 13 keeps its own 3-minute first update and 5-minute cadence, though its own metric list still cites the 30-minute target."}, {"from_proposal": 3, "steps": [9], "what": "Blocking SEV-1 closure until high-priority action items are done and tying completion to performance reviews.", "why": "P1 step 16 instead uses weekly ageing reviews, VP escalation and facilitator sign-off, keeping incident closure separate from remediation."}]}, {"proposal": 2, "assessment": "improved", "what_changed": "P2 kept its 21-step shape but added several mechanisms that close real gaps rather than restating policy: the SOC 2 evidence clock, severity x class, the Triage Owner rule, three rotations, a priced opt-out, an action-item cap and a regulator clock matrix. It also made its own targets more realistic.", "improvements": ["Step 1's **evidence clock**: names the Type II operating-period problem and starts evidence collection in week one; step 3 moves control mapping and the auto-assembled \"golden incident file\" into month one.", "Step 5's Triage Owner rule — ownership from first acknowledgement until an IC takes over — targets the actual failure mode (the unowned gap before declaration), which round 0 only addressed at declaration.", "Step 4's severity x class matrix with \"class can raise a response, never lower it\" gives a data-integrity SEV-2 a SEV-1 posture, a real improvement for a ledger business.", "Step 6 adds a Watch state with a 30-minute timer, free false declarations tracked as a metric, the two-responder ambiguity rule and a no-IC-runs-two-incidents rule.", "Step 8 adds a paid Platform Duty rotation for the shared PostgreSQL/Kubernetes estate — the one place \"own your code\" breaks down — and caps load at two weeks per quarter by tool configuration.", "Step 9 prices the opt-out (the team buys coverage from a paid pool) and publishes an amnesty so incident records never reach performance reviews.", "Step 15 diagnoses 11-of-64 as an over-generation problem and caps postmortems at three actions with artifact-based definitions of done.", "Step 12 adds a regulator clock matrix naming NYDFS Part 500, money-transmitter, breach and card-network windows; step 13 adds a customer-impact ledger feeding comms, credits, regulatory reporting and ROI.", "Step 17's \"never publish incident count as a team metric\" plus reporting metrics (near-misses, detection gaps, false declarations) guards against hidden incidents.", "IC roster target cut from 40 to 12–16 certified ICs — more credible staffing."], "regressions": ["Step 18 (pilot) still depends on 12 prior steps, so the pilot cannot begin until nearly everything is built; no minimum viable process is defined for the interim weeks.", "Success-metric list has grown to 23 items, several of which (page volume, false-positive rate, no service above two pages) overlap and will be costly to report monthly.", "No costed figure for compensation or tooling despite step 1 promising a funding envelope — P1 at least names $400–600K.", "Alert probation and page budgets are asserted without a fallback if a team simply cannot meet the standard by its wave gate, beyond deferral."], "taken": [{"from_proposal": 3, "steps": [12], "what": "Calculate SLA credit avoidance against programme cost to prove ROI to leadership.", "why": "Folded into step 13 so the funding case is a reconciled number from the customer-impact ledger."}, {"from_proposal": 1, "steps": [5, 8], "what": "The pager and incident tooling must work from a mobile device.", "why": "Added to step 11 as an explicit requirement that an IC can run a SEV-1 from a phone at 3am."}, {"from_proposal": 1, "steps": [9], "what": "Severity-keyed playbooks and pre-written messages.", "why": "Absorbed into step 12 as templates per severity and class, pre-approved by Legal, with next-update time built in."}], "rejected": [{"from_proposal": 1, "steps": [4], "what": "A $50–100 bonus per SEV-1/SEV-2 incident mitigated.", "why": "Step 9 states pay attaches to shifts and callouts, never to incident counts, to prevent gaming."}, {"from_proposal": 3, "steps": [9], "what": "Tie action-item completion rates to team performance reviews.", "why": "Step 9 publishes an amnesty — incident records, near-misses and false declarations are never used in performance reviews; only failure to report is."}, {"from_proposal": 3, "steps": [3], "what": "Consolidate the 12 on-call teams into one unified pool covering all 28 teams, mandatory with no exemptions.", "why": "Step 8 does the opposite: service on-call stays with the owning team, and shared risk gets its own Platform Duty rotation."}]}, {"proposal": 3, "assessment": "improved", "what_changed": "P3 grew from 12 to 20 steps by adopting P2's skeleton — charter, baseline, detection SLOs, compensation, control mapping, dry run, culture — which fills most of the prompt's requirements it previously skipped. It remains the thinnest plan in mechanics, dropped its own escalation automation, and its dependency order pushes drills and metrics past full rollout.", "improvements": ["Added step 1 charter with a Director of Incident Management and the \"Own Your Code, Own Your Pager\" principle, plus step 2 baseline register, alert census and immutable evidence store.", "Added step 8 detection strategy with SLIs for transaction success, settlement lag, external synthetics in both regions and error budgets — round 0 had only synthetic transactions buried in an escalation step.", "Step 5 now separates a federated SME rotation from a central IC rotation and adds the \"Unowned Service\" rule (transfer or decommission), replacing round-0's implausible single unified pool.", "Added step 12 SOC 2 control mapping with \"evidence of operation\" per control and step 19 dry run sampling 10 incidents and interviewing engineers.", "Added step 18 culture and change management: publicly correcting blameful leadership language and rotating engineers off when call-out thresholds break."], "regressions": ["Lost round-0 step 6's auto-escalation (5-minute no-ack to team lead, then IC pool); the round-1 plan has no escalation ladder, no acknowledgement timers and no unresponsive-team path — a direct miss on the prompt's \"escalation paths\" and on the hour-long ownership failures.", "Step 16 (game days, comms drills) depends on step 15 (full rollout), so nothing is rehearsed during the pilot; step 17's metrics dashboard likewise lands after rollout although step 14 claims to measure the pilot against baseline.", "Success metrics carry no dates or checkpoints (\"MTTD < 10 minutes\") and are the shortest list of the three.", "No SLA credit process, no regulator windows (only \"define triggers for NY financial regulations\"), no reserved reliability capacity and no postmortem-action cap — the three levers the peers use to fix $1.3M in credits and the 17% closure rate.", "Step 11 keeps \"SEV-1 cannot be closed until high-priority actions are completed\", which conflates incident closure with remediation and will keep incidents nominally open for weeks."], "taken": [{"from_proposal": 2, "steps": [1, 2], "what": "Named process owner with a charter and steering committee, plus a 12-month baseline register and top-50 noisiest rules.", "why": "Became steps 1–2, adding the mandate and measurement base absent in round 0."}, {"from_proposal": 2, "steps": [7], "what": "Paging contract (symptom-based, runbook-linked) and a 90-day noise sprint over the top 100 rules.", "why": "Reproduced almost verbatim as step 7, replacing round-0's looser \"actionable alerting\" rule."}, {"from_proposal": 2, "steps": [8, 9], "what": "Federated per-team rotation plus a central IC roster, paid stipend, call-out pay and mandatory rest.", "why": "Became step 5, abandoning round-0's \"one unified pool, no exemptions\" model."}, {"from_proposal": 2, "steps": [11, 19, 20], "what": "Status page hosted outside production, control mapping to Trust Services Criteria, and an audit dry run with engineer interviews.", "why": "Adopted in steps 6, 12 and 19, turning audit readiness from a single checklist into a sequence."}, {"from_proposal": 1, "steps": [19, 21], "what": "Change champions, celebration of wins and a published reliability newsletter.", "why": "Became step 18 (culture and change management) and the Reliability Reports bullet in step 17."}], "rejected": [{"from_proposal": 2, "steps": [5], "what": "Escalation, paging and incident lifecycle policy with acknowledgement targets and automatic timers at every layer.", "why": "P3 has no lifecycle or escalation step and also deleted its own round-0 auto-escalation, relying on step 4 role cards alone to fix ownership ambiguity."}, {"from_proposal": 2, "steps": [13], "what": "Reserve a fixed percentage of every team's sprint for reliability work.", "why": "Step 11 relies instead on blocking SEV-1 closure and weekly leadership review, with no protected capacity to do the work."}, {"from_proposal": 1, "steps": [20], "what": "Run the first drill during the pilot (week 8) before wider rollout.", "why": "P3 step 16 places all game days and communication drills after Wave 4, so the process is never rehearsed before it is org-wide."}]}]}
[ROUND 2]
{"round_summary": "P1 absorbed almost the entire distinctive vocabulary of P2's round-1 plan (Triage Owner, severity×class, three rotations, golden incident file, capped action items), producing two near-twin heavyweight plans; P2 added the genuinely new ideas of the round (a two-week minimum viable process, a funded ledger-resilience track, a cost-of-downtime model). P3 went the other way and collapsed into 22 bare titles with no content, losing everything that made it assessable.", "converging": true, "differences": ["**Bridging the gap before tooling exists**: P2 S5 defines ten day-one rules, a manual duty-IC rotation drawn from the 12 teams that already have on-call, and a daily 15-minute stand-up for month one. P1 has no interim process between the charter (S1) and platform selection (S11); P3 has none either.", "**Prevention as a funded track**: P2 S23 is a standalone engineering roadmap with concrete bets (ledger read-only tripwire, connection-pool isolation, PITR restore tests with published timings, rollback on SLO burn). P1 keeps resilience as two bullets inside S9 and S22; P3 does not mention it.", "**Communication clocks**: P1 S13 demands a status-page update within 3 minutes of SEV-1 and internal updates every 5 minutes; P2 S13 sets 15 minutes internal first, 30 minutes to status page, with a mandatory no-news update. P1's 3-minute rule contradicts its own success metric of 30 minutes for 95% of SEV-1s.", "**Level of specification**: P1 and P2 give thresholds, timers, dollar ranges and dates throughout; P3 R2 gives only step titles and a one-line rationale each — no severity definitions, no timings, no compensation mechanics, no dates on any metric."], "influences": ["P1 took nearly all of P2's round-1 signature ideas: Triage Owner (P2 S5→P1 S5/S6), severity×class with \"class can raise, never lower\" (P2 S4→P1 S4), the three-rotation model (P2 S8→P1 S7), priced opt-out and amnesty (P2 S9→P1 S8), golden incident file and month-one control mapping (P2 S1/S3→P1 S1/S3), three capped action items and the repeat-incident design review (P2 S15→P1 S16).", "P2 took P1's weekly synthetic-page testing of escalation ladders (P1 R1 S9→P2 S7) and P1's status-page-component-to-customer-journey mapping plus a named status-page owner (P1 R1 S14→P2 S14).", "P2 took P3's culture and change-management step (P3 R1 S18) and turned it into S24: on-the-spot correction of blame language, pager-fatigue monitoring, public recognition for deleted alerts.", "P3 took P2's \"evidence clock\" framing into its S1 title but nothing else of substance; it adopted no new mechanisms this round.", "Nobody adopted P3's error budgets triggering feature freezes (P3 R1 S8) or its rule that a SEV-1 cannot close until high-priority actions are done (P3 R1 S11) — P1 S16 and P2 S16 instead cap actions at three and track them in a separate reliability backlog."], "proposals": [{"proposal": 1, "assessment": "improved", "what_changed": "P1 rewrote itself around P2's round-1 mechanisms while keeping its own depth. New steps 4, 5, 7, 12, 14, 15, 20, 21 replace vaguer round-1 equivalents, and detection, escalation and postmortem policy are now far more concrete. A truncated step 11 and a few internal contradictions are the cost.", "improvements": ["S5 adds the **Triage Owner rule** with a 15-minute triage decision, closing the unowned gap between page and declaration that caused the two hour-long ambiguity incidents.", "S4 adds response classes orthogonal to severity, with automatic triggers (ledger write failure → SEV-1, payment success <99% for 5 min, missed settlement window) instead of round-1's generic definitions.", "S7 replaces the single \"on-call architecture\" step with three named rotations (Service, Platform Duty, IC roster of 12–16) and a hard gate: build a rotation of ≥6 or transfer ownership by end of month 2.", "S16 caps postmortems at three action items, requires an artifact for closure, and adds the repeat-incident design-review rule — a plausible fix for 11-of-64 rather than more tracking.", "S12 adds mechanical escalation timers, the dual-IC rule, the ambiguity rule and a 30-minute Watch state.", "Metrics now carry per-month deadlines and include credit avoidance, detection contracts on all 180 services, and golden-file completeness."], "regressions": ["**Step 11 is cut off mid-sentence** (\"Define rollback criteria (e.g.,\"), losing the dual-run exit criteria and cutover date.", "Success metric says SEV-1 MTTM <45 minutes by month 9, but S18 sets the target at <60 minutes — two numbers for the same thing.", "S13's 3-minute status-page update for SEV-1 contradicts the stated metric of 30 minutes for ≥95% of SEV-1s, and is barely achievable with legal-cleared templates.", "Alert target tightened to <400/month in the metrics while S10 still says <600 — inconsistent.", "No interim operating model: the plan produces nothing usable until tooling is selected, despite claiming \"working process in month 2\"."], "taken": [{"from_proposal": 2, "steps": [5], "what": "The Triage Owner owns the incident from page acknowledgement until an IC takes over or stand-down.", "why": "Adopted verbatim as a new role in S5/S6 and made mandatory for all incidents."}, {"from_proposal": 2, "steps": [4], "what": "Severity times response class, where class can raise the response but never lower it.", "why": "Became S4, with data-integrity incidents granted SEV-1 posture regardless of scope."}, {"from_proposal": 2, "steps": [8], "what": "Three rotations — Service On-Call, Platform Duty, central IC roster.", "why": "Became S7, framed as making the \"other teams' code\" objection structurally impossible."}, {"from_proposal": 2, "steps": [9], "what": "Opt-out with a price paid by the team, plus an amnesty on incident records in performance reviews.", "why": "Both adopted into S8 to convert a cultural argument into a budget line."}, {"from_proposal": 2, "steps": [1, 3], "what": "Start the SOC 2 evidence clock on day one and map controls in month one, with a golden incident file as the audit unit.", "why": "Added to S1 and S3, replacing round-1's late control-mapping step 21."}, {"from_proposal": 2, "steps": [15], "what": "Cap postmortems at three action items; a repeat contributing factor triggers a design review, not another ticket.", "why": "Adopted in S16 with artifact-based closure sign-off."}, {"from_proposal": 2, "steps": [17], "what": "Never publish incident count as a team metric; publish reporting metrics instead.", "why": "Added verbatim to S18 to stop incentivising hidden incidents."}, {"from_proposal": 2, "steps": [12], "what": "A regulator clock matrix mapping event type to regulator, window and signer.", "why": "Folded into S13, with Compliance owning all filings and pre-cleared templates."}, {"from_proposal": 2, "steps": [7, 10], "what": "Detection-gap ticket when a customer reports first, and two-week ticket-only probation for new alerts.", "why": "Both added to S9 and S10 as automatic, no-judgment mechanisms."}, {"from_proposal": 2, "steps": [19], "what": "Sequence rollout waves by incident density and cost of failure, not by ease.", "why": "S20 now puts the highest-credit teams first so their improvement is the proof."}], "rejected": [{"from_proposal": 2, "steps": [12], "what": "A 15-minute internal first update and a 30-minute status-page clock for SEV-1.", "why": "P1 S13 instead demands a 3-minute internal update and a 3-minute status-page post, keeping its round-1 aggressive cadence."}, {"from_proposal": 3, "steps": [11], "what": "A SEV-1 cannot be marked closed until high-priority action items are completed or deferred with VP approval.", "why": "P1 S5 closes incidents at the Closed state once actions are \"tracked or dismissed\" and tracks completion separately in S16."}]}, {"proposal": 2, "assessment": "improved", "what_changed": "P2 kept its round-1 architecture intact and added the three things it was missing: an immediately usable interim process, a funded prevention track, and an explicit culture step. Compliance sequencing, metrics and communications are essentially unchanged and were already strong.", "improvements": ["New S5 **day-one operating rules**: ten rules needing no procurement, a manual duty-IC rotation drawn from the 12 teams that already have on-call, a register entry within 24 hours, and a daily 15-minute stand-up for month one — the only plan that produces value before tooling lands.", "New S23 turns \"resilience roadmap\" from a bullet into a funded track with testable items: ledger read-only tripwire, connection-pool isolation per domain, measured per-region RTO, scheduled PITR restores with published timings, rollback on SLO burn.", "S2 adds a cost-of-downtime model (dollars per minute per journey) that later drives wave sequencing and the funding case.", "S9 adds a New York labour, overtime and tax review with Legal and Finance before announcing paid on-call — the only plan that notices the legal exposure.", "S14 broadens the credit ledger into a customer-trust workstream: CTO-signed note to the top 50 accounts, quarterly public reliability report, account-manager script.", "New S24 makes incentives explicit: restate amnesty at every wave, correct blame language in the executive review itself, pull engineers off rotation when the intrusion cap breaks."], "regressions": ["24 steps with S5, S14, S23 and S24 overlapping existing steps (S14 duplicates parts of S13; S24 duplicates parts of S17 and S22) — the plan is getting long at the edges.", "S23 depends only on S1 and S17, so a major engineering track sits outside the rollout gates with no stated budget split against the $500–700K envelope.", "Success metrics are almost unchanged from round 1 despite the new steps; nothing measures the day-one process (S5) or the resilience bets (S23)."], "taken": [{"from_proposal": 3, "steps": [18], "what": "A dedicated culture and change-management step covering pager fatigue, recognition and blame correction.", "why": "Became S24, with amnesty restated at every wave and engineers pulled off rotation when intrusion caps break."}, {"from_proposal": 1, "steps": [9], "what": "Weekly synthetic pages to test every escalation path.", "why": "Added to S7 alongside the timing adjustment after the first month of real operation."}, {"from_proposal": 1, "steps": [14], "what": "Status-page components mapped to customer journeys rather than internal services, with a named owner.", "why": "Added to S14 so the page speaks in the customer's terms."}, {"from_proposal": 1, "steps": [7, 23], "what": "Fund a resilience roadmap separate from incident response to reduce shared-database blast radius.", "why": "Promoted from a bullet in its own round-1 S21 to a full step (S23) with named bets and quarterly milestones."}], "rejected": [{"from_proposal": 1, "steps": [13], "what": "Status-page update within 3 minutes of SEV-1 declaration and internal updates every 5 minutes.", "why": "S13 deliberately keeps 15 minutes internal, 30 minutes to the status page, with a mandatory no-new-information update."}, {"from_proposal": 1, "steps": [], "what": "", "why": "A roster of 40+ certified incident commanders.\n"}, {"from_proposal": 1, "steps": [], "what": "A roster of 40+ certified incident commanders as a success metric.", "why": "P2 holds at 12–16 certified ICs with primary and secondary on one-week shifts, arguing depth of two per team group is what matters."}, {"from_proposal": 3, "steps": [11], "what": "Blocking incident closure on completion of high-priority action items.", "why": "S16 instead caps actions at three, requires an artifact for closure, and escalates ageing items to the VP of Engineering."}]}, {"proposal": 3, "assessment": "worsened", "what_changed": "P3 discarded all step content and submitted 22 titles with a single sentence of rationale each. Every operational detail it had in round 1 — severity thresholds, status-page timings, compensation mechanics, wave schedule, training curriculum — is gone, and several metrics were loosened or stripped of dates.", "improvements": ["Step 1 now names the SOC 2 evidence clock, acknowledging that Type II evidence cannot be backfilled.", "Step 3 is retitled \"Severity & Class Taxonomy\", nodding to the class dimension.", "MTTD target tightened from <10 to <5 minutes."], "regressions": ["**All step bullets removed**: no SEV-1 definition, no status-page timing, no stipend structure, no wave schedule, no IC certification content — the plan is now an outline, not a process.", "Metrics loosened: SEV-1 MTTR from <45 to <60 minutes, alert volume from <400 to <600, action completion from >90% to >80%; no metric carries a date or month.", "\"40+ certified Incident Commanders\" is retained with no justification and conflicts with the 12–16 roster both other plans converged on.", "Dropped concrete round-1 content that was its own: the 6-week pilot with named teams (Payments Core, Ledger/API), the Wave 2–5 month schedule, the new-on-call Help Desk, error budgets triggering feature freezes.", "Ordering is now weaker: SOC 2 control mapping sits at step 20 after full rollout (step 19), despite depending only on 1 and 3 — evidence design arrives after eight months of incidents have already been handled.", "Step 5 (tooling) depends only on step 1, so platform selection precedes the severity taxonomy and alert-quality rules it must encode."], "taken": [{"from_proposal": 2, "steps": [1], "what": "The SOC 2 evidence clock starting on day one.", "why": "Merged into its step 1 title, though with no supporting detail on what evidence is captured."}, {"from_proposal": 2, "steps": [4], "what": "Severity paired with a response class.", "why": "Adopted as the title of step 3, but with no definition of the classes or the raise-never-lower rule."}], "rejected": [{"from_proposal": 2, "steps": [8], "what": "An IC roster of 12–16 certified senior engineers.", "why": "P3's metrics instead demand 40+ certified ICs, following P1's round-1 figure."}, {"from_proposal": 2, "steps": [5, 15], "what": "The Triage Owner rule and the three-action-item cap.", "why": "Step 4 remains \"Incident Command Roles & Structure\" and step 15 remains generic action tracking, taking neither mechanism despite both being prominent in round 1."}]}]}
[FINAL]
{"summary": "The final round offers two near-identical heavyweight plans (P1 and P2) built on the same architecture — evidence clock and SOC 2 control mapping in month one, baseline register, severity×class taxonomy, Triage Owner rule, three on-call rotations (Service/Platform/IC), paid on-call with priced opt-out, SLO and synthetic detection, paging contract, capped and verifiable postmortem actions, pilot then four gated waves, dry run, standing council — plus P3, which submitted 22 step titles with no content. P2 is distinguished by a two-week minimum viable process, a funded ledger-resilience track, a cost-of-downtime model and internally consistent metrics; P1 is distinguished by sharper numeric thresholds but carries several self-contradictions and one truncated step. P3 is not executable as written.", "assessments": [{"proposal": 1, "fitness": "strong", "strengths": ["**Sharpest numeric specification of the three**: SEV-1 as >5% transaction failure for >5 min, replication lag >10s triggering escalation review, payment success <99% for 5 min, detection targets of <5 min payment path / <10 min elsewhere (S4, S9).", "Closes the command-ambiguity gap concretely: Triage Owner owns from acknowledgement (S5), escalation timers 5/10/15 min with severity-based tempo, dual-IC rule, Watch state capped at 30 min, ambiguity rule \"if in doubt, declare\" (S12).", "On-call model answers the stated objection structurally: Service rotation for own code only, separate paid Platform Duty for the shared PostgreSQL/Kubernetes estate, 12–16 certified ICs on a central roster, 16 teams must build a rotation or transfer ownership by end of month 2 (S7).", "Compensation is priced and testable: $600–1,000/week stipend, 1.5× callout with a one-hour minimum, documented rest day, 3-page/week intrusion cap, priced opt-out from a paid pool, amnesty clause (S8).", "Alert programme is enforceable, not exhortation: paging contract checked in CI, two-week ticket-only probation for new alerts, expiry dates on all silences, 90-day burn-down of the top 100 rules with deletion celebrated (S10).", "Audit work starts in month 1 (S3: CC7.1–7.5 mapping, named control owner, golden incident file, retention/immutability) and ends with a sampled dry run of 10–15 real incidents plus mock auditor interviews (S21).", "Training is unusually concrete: 4-hour IC bootcamp with 75% written pass plus graded live simulation, 12-month recertification, monthly tabletops from real incidents, game days that drill the process's own failure modes (S17)."], "weaknesses": ["**Internally contradictory clocks**: S13 demands a status-page update within 3 minutes of SEV-1 and internal updates every 5 minutes, while its own success metric is 30 minutes for ≥95% of SEV-1. A 3-minute public clock is not achievable with a legal/compliance-cleared template process and will be missed constantly.", "Other metric mismatches: success metrics say <400 alerts/month while S10 targets <600; metrics say SEV-1 MTTM <45 min while S18 sets the target at <60 min; S10's \"2 pages per on-call shift per month\" is ambiguous wording.", "Step 11 is truncated mid-sentence (\"Define rollback criteria (e.g.,\"), leaving the tool cutover criteria unspecified.", "**No interim process**: nothing runs between the charter (S1) and platform selection (S11). Incidents will keep happening in months 1–3 with the old, broken process and no day-one rules.", "Prevention is only bullets inside S9 and S22, not a funded track; MTTM of 45 min against a single shared ledger cluster depends mostly on architecture that the plan does not schedule.", "Budget looks understated: $500–800K for tooling, training, comp and resilience, yet ~30 concurrent paid rotations at $600–1,000/week alone approach $1M/year before callouts, secondaries and platform licences.", "Rotation arithmetic conflicts: a 6-engineer rotation floor plus one-week shifts is ~8.7 primary weeks per engineer per year, which breaches the \"no more than 2 weeks per quarter\" cap once secondary shifts are added.", "Rollout timing is optimistic: metrics claim all 28 teams by month 5, but pilot plus four waves spaced three weeks apart cannot finish that early after a 10-dependency pilot gate."]}, {"proposal": 2, "fitness": "strong", "strengths": ["**S5 minimum viable process is the single most valuable step in the round**: ten day-one rules with no procurement, a manual duty-IC rotation drawn from the 12 teams that already have on-call, a register entry within 24 hours, and a daily 15-minute stand-up for month one. Something works in week two, not month four.", "S23 funds prevention as a separate engineering track with concrete bets: ledger read-only tripwire, connection-pool isolation per domain, statement timeouts, PITR restore tests with published timings, automatic rollback on SLO burn, change freeze around settlement. This is the only credible path to a 60-minute SEV-1 mitigation on a shared ledger.", "Communications are the most defensible: 15 min internal / 30 min status page for SEV-1, mandatory no-news updates, \"internal leads, external follows\", and a regulator clock matrix (NYDFS Part 500, money-transmitter, breach, card network, public-company disclosure) with the rule that **the clock starts at awareness, not root cause** and Compliance always signs (S13).", "S14 turns $1.3M in credits into a mechanism: one durable customer-impact record per incident used for comms, credit computation, regulatory reporting and the annual review, with end-to-end credit automation and monthly accrued-vs-paid reconciliation.", "Realism details others miss: checking NY labour, overtime and tax treatment with Legal and Finance before announcing paid on-call; pay attaches to shifts and callouts, never to incident counts (S9).", "Metrics are dated, baselined and mutually consistent (MTTD <8 min by month 6, customer-detected <10% by month 9, 28/28 rotations by month 6, credits −50%, every page with a recorded disposition within 10 working days) — and S17 bans publishing incident count as a team metric.", "S2 adds a cost-of-downtime model (dollars per minute per journey) that is then actually used to sequence rollout waves by cost of failure (S20).", "Compliance sequencing is right: control mapping in month 1 (S3) with auditor readiness meeting inside 90 days, dry run six weeks out with re-testing of remediated controls (S21), audit liaison as close as possible to the process owner.", "S24 makes culture an explicit control surface: amnesty restated at every wave, blame language corrected on the spot in executive reviews, pager fatigue acted on as soon as the intrusion cap breaks."], "weaknesses": ["Less numeric than P1 in places: \"replication lag above threshold\", \"a per-shift stipend benchmarked to the New York market\" and \"a maximum number of off-hours pages\" leave concrete figures to be decided later.", "No dedicated per-severity playbook/decision-tree step; playbook content is spread across roles (S6), lifecycle (S7) and communications (S13).", "Same budget realism gap as P1: $500–700K a year will not cover ~30 paid rotations plus platform licensing for 260 engineers plus drill time.", "Same rotation arithmetic tension: a 6-engineer floor with one-week shifts sits uneasily with \"no engineer on call more than two weeks per quarter\" (S8).", "24 steps with some duplication — status page appears in S12 (hosting, mobile), S13 (cadence) and S14 (owner, component mapping).", "S19 pilot depends on eleven prior steps, a heavy gate; only S5 keeps the organisation covered in the meantime.", "\"Customer-impacting incidents fall at least 40% year over year\" depends almost entirely on S23 landing, and S23's milestones are quarterly rather than tied to specific dates."]}, {"proposal": 3, "fitness": "weak", "strengths": ["The 22 step titles cover the right territory and a sensible dependency graph (charter → baseline → taxonomy → roles → tooling → detection → alerts → on-call → comms → postmortems → metrics → pilot → waves → dry run → governance).", "Keeps SOC 2 control mapping (S20) and a dry run (S21) as distinct steps rather than folding them into rollout."], "weaknesses": ["**No content at all**: every step is a title plus one sentence. There are no severity thresholds, no escalation timers, no compensation figures, no comms cadences, no wave schedule, no readiness checklist, no postmortem template, no evidence artifacts. Nothing here can be executed or reviewed.", "Regressed from its own round-1 plan, which had status-page timings, comp mechanics, a month-by-month rollout and a training curriculum.", "Metric \"40+ certified Incident Commanders available 24x7\" is unrealistic for 260 engineers and was already corrected to 12–16 by the other plans; most metrics carry no dates (MTTD <5 min, MTTR <60 min, credits <$100K \"annually\").", "No interim process, no resilience/prevention work, no Triage Owner or ownership-gap rule, no regulator clock, no credit process, no action-item cap or closure-evidence rule — the mechanisms that actually fix the stated failures are absent.", "S20 (control mapping) depends only on S1 and S3 but is numbered after rollout; as written it is unclear whether evidence capture starts early enough for a Type II operating period.", "No risk handling: no rollback for tool cutover, no what-if for unready teams, no failure modes of the process itself."]}], "ranking": [2, 1, 3], "ranking_reasons": "- **P2 over P1** on three things that matter operationally: a two-week minimum viable process (S5) so the organisation is covered during the months P1 leaves uncovered; a funded resilience track (S23) that gives the 60-minute SEV-1 target a mechanism instead of a hope; and internally consistent, dated metrics. P1's 3-minute status-page rule contradicts its own 30-minute metric, its alert target (<400 vs <600) and MTTM target (45 vs 60 min) disagree with its own steps, and step 11 is cut off mid-sentence.\n- P1 beats P2 on numeric thresholds (severity triggers, comp figures, escalation tempo, training pass marks) but those are easier to fill in later than the structural gaps P2 closes.\n- **P1 far above P3**: P1 is executable today; P3 is a table of contents. P3 also carries an unrealistic metric (40+ certified ICs) and lost every operational detail it had in earlier rounds.", "versus_initial": [{"proposal": 1, "verdict": "better", "why": "Initial P1 was an exhaustive checklist with no evidence architecture, no SLO-based detection and a fatal ordering flaw; final P2 is a sequenced programme with mechanisms attached to each named failure.", "how": "- Initial P1 S17 depended on all sixteen prior steps and put training (S18), rollout (S19) and the first drill (S20) after it, so nobody practised before going live; final P2 trains and pilots before waves and starts a manual process in week two (S5).\n- Initial P1 treated detection as alert routing; P2 S10 defines SLIs/SLOs for 20 customer journeys, external synthetics in three locations and PostgreSQL-specific signals, plus a detection-gap ticket whenever a customer reports first.\n- Initial P1 had an audit checklist at S16; P2 maps Trust Services Criteria in month one (S3) with named control owners, a golden incident file and retention rules, because Type II evidence cannot be backfilled.\n- Initial P1 paid a $50–100 bonus per SEV-1/2 mitigated, which incentivises declaring incidents; P2 S9 explicitly attaches pay to shifts and callouts only.\n- P2 adds what initial P1 had no equivalent of: a regulator clock matrix, an SLA credit ledger, a capped three-action rule and a funded ledger-resilience track."}, {"proposal": 2, "verdict": "better", "why": "Same architecture, but the final version adds the mechanisms that make it start fast, stay honest and prevent recurrence.", "how": "- New S5 day-one rules and a manual duty-IC rotation cover the gap before tooling exists — the initial plan had nothing between charter and platform selection.\n- New Triage Owner rule (S6) closes the ownership gap that actually caused the two hour-long command failures; the initial plan only defined roles from declaration onward.\n- New severity×class taxonomy with \"class can raise, never lower\" (S4) gives data-integrity incidents SEV-1 posture.\n- New S23 resilience track (ledger read-only tripwire, connection-pool isolation, PITR restore timings, rollback on SLO burn) turns prevention from a bullet into a funded roadmap.\n- New S14 credit ledger and S13 regulator clock matrix (NYDFS Part 500, clock starts at awareness) replace generic \"regulatory obligations\" bullets.\n- New priced opt-out, amnesty and three-action cap; IC roster target cut from 40+ to a realistic 12–16."}, {"proposal": 3, "verdict": "better", "why": "Initial P3 was the thinnest of the three and got two structural choices wrong; final P2 fixes both and adds the audit and detection depth P3 lacked entirely.", "how": "- Initial P3 S3 merged all 28 teams into one unified on-call pool, recreating the exact \"pager for other teams' code\" grievance; P2 S8 runs per-team Service rotations plus a separate paid Platform Duty for shared infrastructure.\n- Initial P3 S9 tied action-item completion to performance reviews, contradicting its own blameless charter; P2 S9 publishes an amnesty and P2 S16 enforces closure by artifact and sign-off instead.\n- Initial P3's entire SOC 2 preparation was one bullet (\"simulate auditor questions\"); P2 has month-one control mapping, a golden incident file, a gap register and a sampled dry run with re-testing.\n- Initial P3's 5-minute status-page rule was unachievable; P2 uses 30 minutes with mandatory no-news updates and pre-cleared templates.\n- P2 adds detection engineering, a cost-of-downtime model and a resilience track that P3 never had."}], "improved_over_initial": true, "improvement_summary": "- **Gained: an interim process.** The best final plan makes something work in two weeks (day-one rules, manual duty IC, register entry within 24 hours) instead of waiting for procurement.\n- **Gained: the ownership gap is named and closed.** The Triage Owner rule, the Watch state with a 30-minute timer, the ambiguity rule and the two-simultaneous-SEV-1 rule target the specific failure described in the brief.\n- **Gained: compliance sequencing.** Control mapping, the golden incident file and retention rules moved to month one on the argument that Type II evidence covers an operating period and cannot be backfilled.\n- **Gained: prevention as a funded track** (ledger tripwire, connection-pool isolation, PITR restore timings, rollback on SLO burn) — the only credible route from a 3h10 mitigation to under an hour.\n- **Gained: honest economics.** Priced opt-out, a cost-of-downtime model, credit avoidance tracked against programme cost, and pay detached from incident counts.\n- **Gained: metric hygiene.** Incident count banned as a team metric; targets dated and baselined against a 12-month register plus a silent-failure register.\n- **Lost: per-severity playbooks with decision trees** (initial P1 S9) survive only as templates and training in the final plans.\n- **Lost: plan diversity.** Two of the three finals are near-twins, and the third was hollowed out to titles; no alternative on-call architecture or staged scope was argued at the end.\n- **Not gained: a budget that adds up.** Both heavyweight finals keep a $500–800K envelope that ~30 paid rotations alone would exhaust.", "process_evaluation": "", "process_issues": [], "suggestions": []}
[VOTE COMPARISON]
{"agrees": true, "comment": "The vote lands on P2, my first choice, and the two P2 voters cite exactly the structural reasons I did: the two-week minimum viable process (S5) that covers the gap before full rollout, the funded resilience track aimed at the shared PostgreSQL ledger, and the explicit handling of the \"pager for other teams' code\" resistance. The dissenting vote for P1 rests on step count and specificity (22 steps, numeric thresholds, control mapping in step 3) but never tests P1's internal consistency \ncontradictory 3-minute vs 30-minute status-page timing, <400 vs <600 alert target, 45 vs 60 min MTTM, and step 11 truncated mid-sentence \nwhich looks like a length-and-confidence bias. No voter ranked P3 highly, consistent with my view that it degenerated into a table of contents with an implausible 40+ certified IC metric."}
[PROCESS]
{"process_evaluation": "- **Convergence was half-earned, half-imitation.** P2's round-0 plan really was the strongest on the merits (baseline evidence pack, SLO/synthetic detection, federated ownership, TSC mapping plus dry run), so migrating toward it was rational. But the migration happened by wholesale vocabulary transfer — Triage Owner, severity×class, golden incident file, three rotations, priced opt-out, three-action cap, no-runbook-no-page — with no agent ever stating *why* that mechanism beats what it replaced. P1 abandoned its own hybrid IC-pool model and per-incident bonus without a single sentence of rebuttal.\n- **Criticism existed only as silent non-adoption.** The only visible judgements were implicit: P2 banning pay tied to incident counts (killing P1's bonus), nobody adopting P3's action-items-in-performance-reviews (which contradicted its own blameless charter), P2 cutting its own \"40+ certified ICs\" to 12–16. No agent ever named a rival's flaw and argued against it. Consequently obvious defects survived three rounds: **P1's status-page rule of 3 minutes contradicts its own success metric of 30 minutes for 95% of SEV-1s**, and P1's final step 11 is literally truncated mid-sentence (\"Define rollback criteria (e.g.,\").\n- **Premises were barely questioned.** The one real challenge is P2 S23: incident process cannot save a single shared PostgreSQL ledger, so fund a resilience track (read-only tripwire, connection-pool isolation, PITR timings) alongside. Nobody questioned whether MTTD <5 min is attainable, whether a 99.95% SLA is compatible with one shared ledger cluster, or when the SOC 2 Type II *observation window* actually opens relative to the \"audit in eight months\" — P2 invokes an \"evidence clock\" but no agent checked the date.\n- **Nobody audited the numbers, and the numbers do not hold.** 28 service rotations with primary and secondary is roughly 56 engineers on call at any hour; at the proposed $600–1,000/week stipend that alone is ~$1.8–2.9M/year, before Platform Duty and the IC roster — three to five times P1's ($500–800K) and P2's ($500–700K) stated envelopes, and above the $1.3M in credits being used to justify the programme. Similarly, 260 engineers over 28 teams is ~9 per team, so the six-engineer rotation floor plus a two-weeks-per-quarter cap is arithmetically tight; \"merge small teams\" was asserted, never modelled.\n- **What was lost: diversity, then content.** By round 2 there was effectively one plan with two echoes, so the vote adjudicated nothing. P3 destroyed its own work — round 2 is 22 titles with all severity thresholds, comms timings, compensation mechanics, wave schedule and curriculum deleted, and it still carries the stale \"40+ certified ICs\" metric P2 had already retired. Dropped without argument: P3's error budgets triggering feature freezes, P3's rule that a SEV-1 cannot close with open high-priority actions, P1's per-severity playbooks with decision trees (diluted), and any serious alternative to the federated on-call model.\n- **Pressure was purely additive.** P2 grew 21→24 steps, P1 21→22 with far denser bullets; no round asked what to delete. A plan that celebrates deleting alerts never pruned itself, and nobody assessed whether one Director plus four pilot teams can absorb 24 concurrent workstreams inside eight months.", "process_issues": ["No critique mechanism: agents never named a rival's flaw or defended a rejected idea; disagreement was expressed only by silently not copying.", "P1's internal contradiction survived all three rounds — status-page update within 3 minutes of SEV-1 (S13) versus its own success metric of 30 minutes for 95% of SEV-1s — and so did SEV-1 internal updates every 5 minutes, which is unrunnable for a 3h-class incident.", "No output quality gate: P1's final step 11 is cut off mid-sentence, and P3's round-2 submission is 22 bare titles with zero operational content, a strict regression from its round-1 plan.", "Cross-step arithmetic never checked: the on-call compensation budget ($500–800K) is 3–5x too small for ~56 engineers on call at $600–1,000/week, and the six-engineer rotation floor was never reconciled with ~9 engineers per team.", "Metric-to-step date consistency never validated: e.g. P2 promises 100% paid on-call shifts from month 2 while 16 teams still have no rotation and waves run to month 6.", "Premature monoculture: after round 1 all three plans shared one skeleton, so rounds 2's vote had no genuine alternative to choose between — the outcome ratified round-0 seniority, not argued superiority.", "No end-to-end stress test of any plan against a concrete scenario (ledger corruption during settlement, IC unreachable, status page down), even though all plans prescribe exactly such drills for the organisation.", "Unbounded growth: success-metric lists reached 24 bullets (P2) and step counts 24; no round rewarded simplification, feasibility under a single programme owner, or a critical path to the audit date.", "The SOC 2 premise went unexamined: nobody established when the Type II observation period starts, which determines whether \"audit-ready by month 7\" is a real plan or a hope.", "Copy without freshness checks: P3 carried P2's abandoned \"40+ certified ICs\" target into the final round after P2 had corrected it to 12–16."], "suggestions": ["Require an explicit objections block each round: every agent names at least two concrete defects in each rival plan (quoting the step) plus the fix, and must justify in one sentence any mechanism it adopts from another plan.", "Assign a numbers-auditor role for one round: recompute compensation cost from rotation headcount, rotation size against 260/28 engineers, IC roster hours, and check every success-metric date against the step that delivers it.", "Force a scenario walkthrough: each plan must narrate minute-by-minute a SEV-1 ledger corruption at 02:00 on a settlement day with the IC unreachable and the status page down — gaps like the missing pre-declaration owner surface immediately.", "Add a pruning round with a hard budget (e.g. max 15 steps, max 12 success metrics): agents must delete and justify what survives, countering additive drift.", "Validate submissions mechanically: reject steps with no acceptance criteria, owner or date, and reject any round whose content is shorter than the agent's previous round unless the deletion is argued.", "Mandate divergence for at least one round: assign one agent to defend a centralised responder pool, one to defend a minimal 90-day version aimed only at the audit and the noise, so the final vote compares real alternatives rather than near-twins.", "Require one premise challenge per agent per round (e.g. is MTTD <5 min reachable on a shared ledger; is the SLA renegotiable; has the Type II window already opened), scored as part of the plan's merit.", "Run an automated consistency check on each plan between success metrics and step text (timings, targets, headcounts) and publish the discrepancies to all agents before the next round."]}
[BRIEF]
{"brief": "- **Outcome: P2 wins, and I agree.** The vote landed on P2 (two votes, one dissent for P1), citing the same structural reasons I did: the two-week minimum viable process (S5), the funded ledger-resilience track (S23) and the explicit handling of \"pager for other teams' code\".\n- **The final plan beats every round-0 proposal.** New gains: an interim process usable in two weeks, the ownership gap closed by name (Triage Owner, Watch state with 30-minute timer, ambiguity rule, two-simultaneous-SEV-1 rule), SOC 2 control mapping and the golden incident file moved to month one, prevention funded as a track, and honest economics (priced opt-out, credit avoidance vs programme cost, pay detached from incident counts).\n- **Convergence was real on the merits but executed as imitation.** P2's round-0 architecture (baseline evidence pack, SLO/synthetic detection, federated ownership, TSC mapping plus dry run) was genuinely strongest, but P1 and P3 rebuilt on it step-for-step and P1 abandoned its hybrid IC-pool model and per-incident bonus without a sentence of argument.\n- **Criticism existed only as silent non-adoption.** No agent ever named a rival's flaw; disagreement showed up as P2 banning incident-count-based pay, nobody copying P3's action-items-in-performance-reviews, and P2 quietly cutting its own \"40+ certified ICs\" to 12–16 — which P3 then carried into the final round anyway.\n- **P3 regressed to a table of contents.** Round 2 was 22 bare titles: severity thresholds, status-page timings, compensation mechanics, wave schedule and curriculum all deleted, making it unexecutable and correctly ranked last by everyone.\n- **Biggest problem: nobody audited the arithmetic.** ~28 service rotations with primary and secondary is ~56 engineers on call at $600–1,000/week, roughly $1.8–2.9M/year against stated envelopes of $500–800K — above the $1.3M in credits justifying the programme — and the six-engineer rotation floor was never reconciled with ~9 engineers per team.\n- **Second problem: no quality gate and unchecked internal consistency.** P1's 3-minute status-page rule contradicts its own 30-minute metric for 95% of SEV-1s, its alert (<400 vs <600) and MTTM (45 vs 60 min) targets disagree with its steps, and its step 11 is truncated mid-sentence; the SOC 2 Type II observation-window start date was never established.\n- **Fixes that would help most:** (1) a mandatory objections block — two quoted defects per rival plan plus a one-sentence justification for every borrowed mechanism; (2) a numbers-auditor round recomputing on-call cost, rotation headcount and every metric-to-step date; (3) a forced scenario walkthrough (SEV-1 ledger corruption at 02:00 on settlement day, IC unreachable, status page down) plus a pruning round with a hard cap of ~15 steps to counter additive drift."}The final round, ranked blind
Why this order
- P2 over P1 on three things that matter operationally: a two-week minimum viable process (S5) so the organisation is covered during the months P1 leaves uncovered; a funded resilience track (S23) that gives the 60-minute SEV-1 target a mechanism instead of a hope; and internally consistent, dated metrics. P1's 3-minute status-page rule contradicts its own 30-minute metric, its alert target (<400 vs <600) and MTTM target (45 vs 60 min) disagree with its own steps, and step 11 is cut off mid-sentence.
- P1 beats P2 on numeric thresholds (severity triggers, comp figures, escalation tempo, training pass marks) but those are easier to fill in later than the structural gaps P2 closes.
- P1 far above P3: P1 is executable today; P3 is a table of contents. P3 also carries an unrealistic metric (40+ certified ICs) and lost every operational detail it had in earlier rounds.
The ranking
- Proposal 2 strong selected by the vote Strengths
- S5 minimum viable process is the single most valuable step in the round: ten day-one rules with no procurement, a manual duty-IC rotation drawn from the 12 teams that already have on-call, a register entry within 24 hours, and a daily 15-minute stand-up for month one. Something works in week two, not month four.
- S23 funds prevention as a separate engineering track with concrete bets: ledger read-only tripwire, connection-pool isolation per domain, statement timeouts, PITR restore tests with published timings, automatic rollback on SLO burn, change freeze around settlement. This is the only credible path to a 60-minute SEV-1 mitigation on a shared ledger.
- Communications are the most defensible: 15 min internal / 30 min status page for SEV-1, mandatory no-news updates, "internal leads, external follows", and a regulator clock matrix (NYDFS Part 500, money-transmitter, breach, card network, public-company disclosure) with the rule that the clock starts at awareness, not root cause and Compliance always signs (S13).
- S14 turns $1.3M in credits into a mechanism: one durable customer-impact record per incident used for comms, credit computation, regulatory reporting and the annual review, with end-to-end credit automation and monthly accrued-vs-paid reconciliation.
- Realism details others miss: checking NY labour, overtime and tax treatment with Legal and Finance before announcing paid on-call; pay attaches to shifts and callouts, never to incident counts (S9).
- Metrics are dated, baselined and mutually consistent (MTTD <8 min by month 6, customer-detected <10% by month 9, 28/28 rotations by month 6, credits −50%, every page with a recorded disposition within 10 working days) — and S17 bans publishing incident count as a team metric.
- S2 adds a cost-of-downtime model (dollars per minute per journey) that is then actually used to sequence rollout waves by cost of failure (S20).
- Compliance sequencing is right: control mapping in month 1 (S3) with auditor readiness meeting inside 90 days, dry run six weeks out with re-testing of remediated controls (S21), audit liaison as close as possible to the process owner.
- S24 makes culture an explicit control surface: amnesty restated at every wave, blame language corrected on the spot in executive reviews, pager fatigue acted on as soon as the intrusion cap breaks.
Weaknesses- Less numeric than P1 in places: "replication lag above threshold", "a per-shift stipend benchmarked to the New York market" and "a maximum number of off-hours pages" leave concrete figures to be decided later.
- No dedicated per-severity playbook/decision-tree step; playbook content is spread across roles (S6), lifecycle (S7) and communications (S13).
- Same budget realism gap as P1: $500–700K a year will not cover ~30 paid rotations plus platform licensing for 260 engineers plus drill time.
- Same rotation arithmetic tension: a 6-engineer floor with one-week shifts sits uneasily with "no engineer on call more than two weeks per quarter" (S8).
- 24 steps with some duplication — status page appears in S12 (hosting, mobile), S13 (cadence) and S14 (owner, component mapping).
- S19 pilot depends on eleven prior steps, a heavy gate; only S5 keeps the organisation covered in the meantime.
- "Customer-impacting incidents fall at least 40% year over year" depends almost entirely on S23 landing, and S23's milestones are quarterly rather than tied to specific dates.
- Proposal 1 strong Strengths
- Sharpest numeric specification of the three: SEV-1 as >5% transaction failure for >5 min, replication lag >10s triggering escalation review, payment success <99% for 5 min, detection targets of <5 min payment path / <10 min elsewhere (S4, S9).
- Closes the command-ambiguity gap concretely: Triage Owner owns from acknowledgement (S5), escalation timers 5/10/15 min with severity-based tempo, dual-IC rule, Watch state capped at 30 min, ambiguity rule "if in doubt, declare" (S12).
- On-call model answers the stated objection structurally: Service rotation for own code only, separate paid Platform Duty for the shared PostgreSQL/Kubernetes estate, 12–16 certified ICs on a central roster, 16 teams must build a rotation or transfer ownership by end of month 2 (S7).
- Compensation is priced and testable: $600–1,000/week stipend, 1.5× callout with a one-hour minimum, documented rest day, 3-page/week intrusion cap, priced opt-out from a paid pool, amnesty clause (S8).
- Alert programme is enforceable, not exhortation: paging contract checked in CI, two-week ticket-only probation for new alerts, expiry dates on all silences, 90-day burn-down of the top 100 rules with deletion celebrated (S10).
- Audit work starts in month 1 (S3: CC7.1–7.5 mapping, named control owner, golden incident file, retention/immutability) and ends with a sampled dry run of 10–15 real incidents plus mock auditor interviews (S21).
- Training is unusually concrete: 4-hour IC bootcamp with 75% written pass plus graded live simulation, 12-month recertification, monthly tabletops from real incidents, game days that drill the process's own failure modes (S17).
Weaknesses- Internally contradictory clocks: S13 demands a status-page update within 3 minutes of SEV-1 and internal updates every 5 minutes, while its own success metric is 30 minutes for ≥95% of SEV-1. A 3-minute public clock is not achievable with a legal/compliance-cleared template process and will be missed constantly.
- Other metric mismatches: success metrics say <400 alerts/month while S10 targets <600; metrics say SEV-1 MTTM <45 min while S18 sets the target at <60 min; S10's "2 pages per on-call shift per month" is ambiguous wording.
- Step 11 is truncated mid-sentence ("Define rollback criteria (e.g.,"), leaving the tool cutover criteria unspecified.
- No interim process: nothing runs between the charter (S1) and platform selection (S11). Incidents will keep happening in months 1–3 with the old, broken process and no day-one rules.
- Prevention is only bullets inside S9 and S22, not a funded track; MTTM of 45 min against a single shared ledger cluster depends mostly on architecture that the plan does not schedule.
- Budget looks understated: $500–800K for tooling, training, comp and resilience, yet ~30 concurrent paid rotations at $600–1,000/week alone approach $1M/year before callouts, secondaries and platform licences.
- Rotation arithmetic conflicts: a 6-engineer rotation floor plus one-week shifts is ~8.7 primary weeks per engineer per year, which breaches the "no more than 2 weeks per quarter" cap once secondary shifts are added.
- Rollout timing is optimistic: metrics claim all 28 teams by month 5, but pilot plus four waves spaced three weeks apart cannot finish that early after a 10-dependency pilot gate.
- Proposal 3 weak Strengths
- The 22 step titles cover the right territory and a sensible dependency graph (charter → baseline → taxonomy → roles → tooling → detection → alerts → on-call → comms → postmortems → metrics → pilot → waves → dry run → governance).
- Keeps SOC 2 control mapping (S20) and a dry run (S21) as distinct steps rather than folding them into rollout.
Weaknesses- No content at all: every step is a title plus one sentence. There are no severity thresholds, no escalation timers, no compensation figures, no comms cadences, no wave schedule, no readiness checklist, no postmortem template, no evidence artifacts. Nothing here can be executed or reviewed.
- Regressed from its own round-1 plan, which had status-page timings, comp mechanics, a month-by-month rollout and a training curriculum.
- Metric "40+ certified Incident Commanders available 24x7" is unrealistic for 260 engineers and was already corrected to 12–16 by the other plans; most metrics carry no dates (MTTD <5 min, MTTR <60 min, credits <$100K "annually").
- No interim process, no resilience/prevention work, no Triage Owner or ownership-gap rule, no regulator clock, no credit process, no action-item cap or closure-evidence rule — the mechanisms that actually fix the stated failures are absent.
- S20 (control mapping) depends only on S1 and S3 but is numbered after rollout; as written it is unclear whether evidence capture starts early enough for a Type II operating period.
- No risk handling: no rollback for tool cutover, no what-if for unready teams, no failure modes of the process itself.
The vote, confronted with the analyst
the analyst agrees with the vote
The vote lands on P2, my first choice, and the two P2 voters cite exactly the structural reasons I did: the two-week minimum viable process (S5) that covers the gap before full rollout, the funded resilience track aimed at the shared PostgreSQL ledger, and the explicit handling of the "pager for other teams' code" resistance. The dissenting vote for P1 rests on step count and specificity (22 steps, numeric thresholds, control mapping in step 3) but never tests P1's internal consistency contradictory 3-minute vs 30-minute status-page timing, <400 vs <600 alert target, 45 vs 60 min MTTM, and step 11 truncated mid-sentence which looks like a length-and-confidence bias. No voter ranked P3 highly, consistent with my view that it degenerated into a table of contents with an implausible 40+ certified IC metric.
Is the analyst's first choice better than the initial proposals? better than every initial proposal
- Gained: an interim process. The best final plan makes something work in two weeks (day-one rules, manual duty IC, register entry within 24 hours) instead of waiting for procurement.
- Gained: the ownership gap is named and closed. The Triage Owner rule, the Watch state with a 30-minute timer, the ambiguity rule and the two-simultaneous-SEV-1 rule target the specific failure described in the brief.
- Gained: compliance sequencing. Control mapping, the golden incident file and retention rules moved to month one on the argument that Type II evidence covers an operating period and cannot be backfilled.
- Gained: prevention as a funded track (ledger tripwire, connection-pool isolation, PITR restore timings, rollback on SLO burn) — the only credible route from a 3h10 mitigation to under an hour.
- Gained: honest economics. Priced opt-out, a cost-of-downtime model, credit avoidance tracked against programme cost, and pay detached from incident counts.
- Gained: metric hygiene. Incident count banned as a team metric; targets dated and baselined against a 12-month register plus a silent-failure register.
- Lost: per-severity playbooks with decision trees (initial P1 S9) survive only as templates and training in the final plans.
- Lost: plan diversity. Two of the three finals are near-twins, and the third was hollowed out to titles; no alternative on-call architecture or staged scope was argued at the end.
- Not gained: a budget that adds up. Both heavyweight finals keep a $500–800K envelope that ~30 paid rotations alone would exhaust.
| Initial proposal | Verdict | Why | How |
|---|---|---|---|
| Proposal 1 |
better | Initial P1 was an exhaustive checklist with no evidence architecture, no SLO-based detection and a fatal ordering flaw; final P2 is a sequenced programme with mechanisms attached to each named failure. |
|
| Proposal 2 |
better | Same architecture, but the final version adds the mechanisms that make it start fast, stay honest and prevent recurrence. |
|
| Proposal 3 |
better | Initial P3 was the thinnest of the three and got two structural choices wrong; final P2 fixes both and adds the audit and detection depth P3 lacked entirely. |
|
Evaluation of the process
- Convergence was half-earned, half-imitation. P2's round-0 plan really was the strongest on the merits (baseline evidence pack, SLO/synthetic detection, federated ownership, TSC mapping plus dry run), so migrating toward it was rational. But the migration happened by wholesale vocabulary transfer — Triage Owner, severity×class, golden incident file, three rotations, priced opt-out, three-action cap, no-runbook-no-page — with no agent ever stating why that mechanism beats what it replaced. P1 abandoned its own hybrid IC-pool model and per-incident bonus without a single sentence of rebuttal.
- Criticism existed only as silent non-adoption. The only visible judgements were implicit: P2 banning pay tied to incident counts (killing P1's bonus), nobody adopting P3's action-items-in-performance-reviews (which contradicted its own blameless charter), P2 cutting its own "40+ certified ICs" to 12–16. No agent ever named a rival's flaw and argued against it. Consequently obvious defects survived three rounds: P1's status-page rule of 3 minutes contradicts its own success metric of 30 minutes for 95% of SEV-1s, and P1's final step 11 is literally truncated mid-sentence ("Define rollback criteria (e.g.,").
- Premises were barely questioned. The one real challenge is P2 S23: incident process cannot save a single shared PostgreSQL ledger, so fund a resilience track (read-only tripwire, connection-pool isolation, PITR timings) alongside. Nobody questioned whether MTTD <5 min is attainable, whether a 99.95% SLA is compatible with one shared ledger cluster, or when the SOC 2 Type II observation window actually opens relative to the "audit in eight months" — P2 invokes an "evidence clock" but no agent checked the date.
- Nobody audited the numbers, and the numbers do not hold. 28 service rotations with primary and secondary is roughly 56 engineers on call at any hour; at the proposed $600–1,000/week stipend that alone is ~$1.8–2.9M/year, before Platform Duty and the IC roster — three to five times P1's ($500–800K) and P2's ($500–700K) stated envelopes, and above the $1.3M in credits being used to justify the programme. Similarly, 260 engineers over 28 teams is ~9 per team, so the six-engineer rotation floor plus a two-weeks-per-quarter cap is arithmetically tight; "merge small teams" was asserted, never modelled.
- What was lost: diversity, then content. By round 2 there was effectively one plan with two echoes, so the vote adjudicated nothing. P3 destroyed its own work — round 2 is 22 titles with all severity thresholds, comms timings, compensation mechanics, wave schedule and curriculum deleted, and it still carries the stale "40+ certified ICs" metric P2 had already retired. Dropped without argument: P3's error budgets triggering feature freezes, P3's rule that a SEV-1 cannot close with open high-priority actions, P1's per-severity playbooks with decision trees (diluted), and any serious alternative to the federated on-call model.
- Pressure was purely additive. P2 grew 21→24 steps, P1 21→22 with far denser bullets; no round asked what to delete. A plan that celebrates deleting alerts never pruned itself, and nobody assessed whether one Director plus four pilot teams can absorb 24 concurrent workstreams inside eight months.
- No critique mechanism: agents never named a rival's flaw or defended a rejected idea; disagreement was expressed only by silently not copying.
- P1's internal contradiction survived all three rounds — status-page update within 3 minutes of SEV-1 (S13) versus its own success metric of 30 minutes for 95% of SEV-1s — and so did SEV-1 internal updates every 5 minutes, which is unrunnable for a 3h-class incident.
- No output quality gate: P1's final step 11 is cut off mid-sentence, and P3's round-2 submission is 22 bare titles with zero operational content, a strict regression from its round-1 plan.
- Cross-step arithmetic never checked: the on-call compensation budget ($500–800K) is 3–5x too small for ~56 engineers on call at $600–1,000/week, and the six-engineer rotation floor was never reconciled with ~9 engineers per team.
- Metric-to-step date consistency never validated: e.g. P2 promises 100% paid on-call shifts from month 2 while 16 teams still have no rotation and waves run to month 6.
- Premature monoculture: after round 1 all three plans shared one skeleton, so rounds 2's vote had no genuine alternative to choose between — the outcome ratified round-0 seniority, not argued superiority.
- No end-to-end stress test of any plan against a concrete scenario (ledger corruption during settlement, IC unreachable, status page down), even though all plans prescribe exactly such drills for the organisation.
- Unbounded growth: success-metric lists reached 24 bullets (P2) and step counts 24; no round rewarded simplification, feasibility under a single programme owner, or a critical path to the audit date.
- The SOC 2 premise went unexamined: nobody established when the Type II observation period starts, which determines whether "audit-ready by month 7" is a real plan or a hope.
- Copy without freshness checks: P3 carried P2's abandoned "40+ certified ICs" target into the final round after P2 had corrected it to 12–16.
- Require an explicit objections block each round: every agent names at least two concrete defects in each rival plan (quoting the step) plus the fix, and must justify in one sentence any mechanism it adopts from another plan.
- Assign a numbers-auditor role for one round: recompute compensation cost from rotation headcount, rotation size against 260/28 engineers, IC roster hours, and check every success-metric date against the step that delivers it.
- Force a scenario walkthrough: each plan must narrate minute-by-minute a SEV-1 ledger corruption at 02:00 on a settlement day with the IC unreachable and the status page down — gaps like the missing pre-declaration owner surface immediately.
- Add a pruning round with a hard budget (e.g. max 15 steps, max 12 success metrics): agents must delete and justify what survives, countering additive drift.
- Validate submissions mechanically: reject steps with no acceptance criteria, owner or date, and reject any round whose content is shorter than the agent's previous round unless the deletion is argued.
- Mandate divergence for at least one round: assign one agent to defend a centralised responder pool, one to defend a minimal 90-day version aimed only at the audit and the noise, so the final vote compares real alternatives rather than near-twins.
- Require one premise challenge per agent per round (e.g. is MTTD <5 min reachable on a shared ledger; is the SLA renegotiable; has the Type II window already opened), scored as part of the plan's merit.
- Run an automated consistency check on each plan between success metrics and step text (timings, targets, headcounts) and publish the discrepancies to all agents before the next round.
Convergence: steps changed per round
- claudeHaiku4.5_refine_1 claudeHaiku4.5 · anthropic/claude-haiku-4-5
- deepseek-flash_refine_2 deepseek-flash · deepseek/deepseek-flash
- qwen3.8-flash_refine_3 qwen3.8-flash · alibaba/qwen3.8-flash
- mean of the agents
| Steps kept, added and removed | Round 1 | Round 2 |
|---|---|---|
| claudeHaiku4.5_refine_1 |
1 | 15 |
| deepseek-flash_refine_2 |
6 | 18 |
| qwen3.8-flash_refine_3 |
3 | 12 |
Contributions of each agent
- claudeHaiku4.5_refine_1 claudeHaiku4.5 · anthropic/claude-haiku-4-5 · extended thinking, 16.0k tokens · temp 1
- deepseek-flash_refine_2 deepseek-flash · deepseek/deepseek-flash · thinking on, effort high
- qwen3.8-flash_refine_3 qwen3.8-flash · alibaba/qwen3.8-flash · thinking on, budget 16.0k tokens
| Agent | Steps of the selected plan it introduced | Steps introduced | Copied by others | Survived to the final round | Ideas taken from it | Ideas rejected | Declared adopted | Declared rejected | Votes received |
|---|---|---|---|---|---|---|---|---|---|
| deepseek-flash_refine_2 selected plan |
24 | 40 | 37 | 34 | 20 | 6 | 2 | ||
| claudeHaiku4.5_refine_1 |
0 | 26 | 12 | 6 | 6 | 5 | 1 | ||
| qwen3.8-flash_refine_3 |
0 | 21 | 3 | 6 | 3 | 5 | 0 |
Efficiency of each agent
Cost per step
Steps per minute of model time
Timeline
Costs
Two separate things were paid for in this run:
- The planning process itself ("Deliberation" below): the 12 LLM calls that produced the plan — 3 agents drafting and refining over 3 rounds, then 3 voters. This is the cost of obtaining the plan. Running
slow-thinkercosts exactly this. - The optional evaluation ("Analysis" below, on yellow like everything the analysis adds): the 7 calls made afterwards by
slow-thinker-report --analyzeso that a reviewing model explains how the proposals evolved and judges the outcome. It does not change the plan and is only paid if you ask for it.
0.31
2.26
2.57
Deliberation (the planning process), by model
| Model | Thinking | Calls | Input tokens | Output tokens | Reasoning | Model time | Cost |
|---|---|---|---|---|---|---|---|
| claudeHaiku4.5 · anthropic/claude-haiku-4-5 | extended thinking, 16.0k tokens · temp 1 | 4 | 57.7k | 38.2k | 13.3k | 7 min 45 s | 0.25 |
| deepseek-flash · deepseek/deepseek-flash | thinking on, effort high | 4 | 51.5k | 59.0k | 37.0k | 4 min 13 s | 0.043 |
| qwen3.8-flash · alibaba/qwen3.8-flash | thinking on, budget 16.0k tokens | 4 | 53.6k | 14.9k | 7.0k | 3 min 53 s | 0.015 |
| Total | 12 | 162.9k | 112.1k | 57.3k | 15 min 51 s | 0.31 |
Deliberation (the planning process), by call
| Phase | Call | Model | Input tokens | Output tokens | Reasoning | Time | Cost |
|---|---|---|---|---|---|---|---|
| Round 0 | claudeHaiku4.5_initial_1 | claudeHaiku4.5 · anthropic/claude-haiku-4-5 | 1.3k | 9.9k | 4.0k | 1 min 41 s | 0.051 |
| Round 0 | deepseek-flash_initial_2 | deepseek-flash · deepseek/deepseek-flash | 1.1k | 16.7k | 10.5k | 1 min 19 s | 0.010 |
| Round 0 | qwen3.8-flash_initial_3 | qwen3.8-flash · alibaba/qwen3.8-flash | 1.1k | 3.8k | 1.3k | 49 s | 0.002 |
| Round 1 | claudeHaiku4.5_refine_1 | claudeHaiku4.5 · anthropic/claude-haiku-4-5 | 15.4k | 10.8k | 3.4k | 2 min 12 s | 0.069 |
| Round 1 | deepseek-flash_refine_2 | deepseek-flash · deepseek/deepseek-flash | 13.7k | 22.4k | 14.7k | 1 min 34 s | 0.015 |
| Round 1 | qwen3.8-flash_refine_3 | qwen3.8-flash · alibaba/qwen3.8-flash | 14.3k | 5.3k | 1.8k | 1 min 38 s | 0.005 |
| Round 2 | claudeHaiku4.5_refine_1 | claudeHaiku4.5 · anthropic/claude-haiku-4-5 | 19.6k | 13.4k | 1.9k | 2 min 50 s | 0.086 |
| Round 2 | deepseek-flash_refine_2 | deepseek-flash · deepseek/deepseek-flash | 17.5k | 18.6k | 10.6k | 1 min 13 s | 0.014 |
| Round 2 | qwen3.8-flash_refine_3 | qwen3.8-flash · alibaba/qwen3.8-flash | 18.3k | 3.2k | 1.4k | 52 s | 0.004 |
| Voting | claudeHaiku4.5_voter_1 | claudeHaiku4.5 · anthropic/claude-haiku-4-5 | 21.3k | 4.1k | 3.9k | 1 min 2 s | 0.042 |
| Voting | deepseek-flash_voter_2 | deepseek-flash · deepseek/deepseek-flash | 19.2k | 1.3k | 1.1k | 7 s | 0.004 |
| Voting | qwen3.8-flash_voter_3 | qwen3.8-flash · alibaba/qwen3.8-flash | 20.0k | 2.6k | 2.5k | 35 s | 0.004 |
Analysis (optional evaluation, not part of the process), by call
| Call | Model | Input tokens | Output tokens | Reasoning | Time | Cost |
|---|---|---|---|---|---|---|
| analysis of round 0 | anthropic/claude-opus-5 | 21.2k | 2.8k | 911 | 42 s | 0.18 |
| analysis of round 1 | anthropic/claude-opus-5 | 49.8k | 12.7k | 6.8k | 2 min 35 s | 0.57 |
| analysis of round 2 | anthropic/claude-opus-5 | 58.3k | 8.7k | 3.5k | 1 min 47 s | 0.51 |
| analysis of final | anthropic/claude-opus-5 | 55.7k | 8.8k | 3.2k | 1 min 52 s | 0.50 |
| analysis of vote comparison | anthropic/claude-opus-5 | 2.5k | 313 | 0 | 6 s | 0.020 |
| analysis of final: process (added afterwards) | anthropic/claude-opus-5 | 54.5k | 5.6k | 3.1k | 1 min 26 s | 0.41 |
| analysis of brief (added afterwards) | anthropic/claude-opus-5 | 10.3k | 1.1k | 0 | 13 s | 0.078 |
| Total | 252.3k | 40.0k | 17.5k | 8 min 40 s | 2.26 |